data unification

From Fragmented Data to Faster Decisions Across 64 Lines of Business for a Major Insurance Carrier

From Fragmented Data to Faster Decisions Across 64 Lines of Business for a Major Insurance Carrier

A major U.S. insurance carrier operating across 64 lines of business had outgrown a fragmented data environment comprising SQL Server OLTP systems, flat files, and Synapse workloads. Each operating unit managed policy, claims, billing, and commission data independently, making enterprise-wide questions about underwriting performance, loss ratios, combined ratios, and profitability slow and difficult to answer. Bitwise modernized the estate through Helix, an enterprise data warehouse framework implemented on a governed, scalable Databricks Lakehouse, creating a single enterprise standard rather than another line-of-business rollout.

Client Challenges and Requirements

  • Fragmented legacy estate: SQL Server OLTP systems, flat files, and Synapse workloads coexisted without a common enterprise architecture.
  • No unified data model: SNF, dimensional, and mirrored structures spread inconsistent definitions across silos.
  • Limited enterprise visibility: Leadership could not consistently compare loss ratios, combined ratios, underwriting performance, and profitability across lines of business.
  • Manual reconciliation: Teams spent significant time aligning policy counts, premiums, losses, and expenses, creating delays and increasing error risk.
  • Weak metadata, lineage, and security controls: Poor discoverability and limited row-level access and masking made governed data use difficult.
  • Inconsistent data quality: The absence of standardized validation allowed bad records to flow into downstream aggregates and undermined confidence in analytics and statutory reporting.
  • Slow, non-reusable onboarding: Each of the 64 operating units ingested data differently, limiting reuse and extending delivery timelines.

Bitwise Solution

  • Enterprise Helix framework: Established a reusable modernization standard across policy, claims, billing, and commission domains.
  • Databricks medallion architecture: Built Raw, Refined, and Modelled layers on the Databricks Lakehouse.
  • Centralized Databricks platform: Consolidated data from multiple core systems into one enterprise warehousing and reporting layer while gradually retiring Synapse workloads.
  • Unity Catalog governance: Implemented a functional-area catalog structure with access control, lineage, and row-level security.
  • Delta Lake reliability: Used ACID transactions and time travel on open Parquet to improve reliability and auditability without lock-in.
  • Embedded reconciliation: Automated the alignment of policy counts, premiums, losses, and expenses so loss and combined ratios were calculated consistently.
  • Governance accelerators: Applied automated data-quality checks, metadata tagging, and row-level masking using FulkrumAI.
  • Reusable ingestion patterns: Standardized PySpark, Fivetran, and Azure Data Factory pipelines to accelerate onboarding across operating units.

Tools & Technologies We Used

Databricks Lakehouse

Delta Lake

Unity Catalog

Databricks Notebooks

Lakeflow Jobs

Lakeflow Connect

Fivetran

Azure Data Factory

Azure SQL

Synapse Dedicated Pool

Azure Repos

Azure DevOps Pipelines

Azure DevOps Pipelines

Microsoft Entra ID

Azure Key Vault

Azure Monitor

Power BI

PySpark

Python

SQL

Data.World

FulkrumAI

Databricks Model Serving

Databricks Agents

APIs

GitHub

Key Results

Established Helix on Databricks as the enterprise standard for policy, claims, billing, and commission data.

Enabled consistent underwriting and actuarial reporting across all 64 lines of business.

Reduced manual reconciliation effort by 70% through centralized, automated controls.

Cut loss ratio and combined ratio reporting time from days to hours—an 85% improvement.

Reduced statutory filing preparation time by 60% across all lines of business.

Accelerated time to market for new lines of business by 30% through reusable frameworks and accelerators.

Enabled predictive analytics for loss reserving and pricing optimization.

Created a unified customer and household view, increasing cross-selling opportunities by 25%.

Solution Architecture

The target architecture moves data from on-premises and operational sources through standardized cloud ingestion into a governed Databricks Lakehouse. Data progresses through Bronze (Raw), Silver (Refined), and Gold (Modelled) layers before being served through data marts, direct query, APIs, AI/ML services, dashboards, and enterprise reporting.

  • Source systems: Guidewire, business data warehouses, financial reporting systems, file servers and documents, and APIs feed the platform alongside Azure SQL and Synapse dedicated pools.
  • Ingestion and orchestration: Fivetran, Lakeflow Connect, Azure Data Factory, and standardized PySpark pipelines support repeatable onboarding; Azure Repos, DevOps pipelines, Terraform modules, and Lakeflow Jobs provide orchestration and CI/CD.
  • Storage and processing: Bronze (Raw) data is profiled and validated, Silver (Refined) data is cleansed and standardized, and Gold (Modelled) data is organized for business-ready reporting.
  • Semantic and consumption layer: Data marts, direct query, API and application integration, AI/ML services, Power BI dashboards, and enterprise reports provide governed access to trusted data.

Cross-cutting governance and security: Unity Catalog, Microsoft Entra ID, Azure Key Vault, Azure Monitor, metadata management, data classification, lineage, masking, archival, and reconciliation controls operate across the platform.

Modernization Architecture

Share

Download Case Study

Let's Engineer Your AI Advantage

Unifying Insurance Data Across 64 Lines of Business | Bitwise