GenAI

Automated Policy Data Extraction for a Fortune 500 Insurer

Automated Policy Data Extraction for a Fortune 500 Insurer

A leading insurance carrier needed to extract key policy attributes—including policy numbers, coverages, limits, deductibles, and effective and expiration dates—from diverse structured XML and unstructured email and PDF documents. Its traditional rules-based process was slow, error-prone, and difficult to scale across DRS and OARS data feeding claims platforms.

Client Challenges and Requirements

  • Manual, rules-based extraction: Processing unstructured emails and PDFs was slow and prone to errors.
  • Inconsistent data quality: Structured XML and unstructured email policy data varied in quality and format.
  • No scalable extraction method: Key policy attributes could not be extracted consistently across diverse document formats.
  • High operational cost: Manual data entry, validation, and reconciliation required significant effort.
  • Slow downstream processing: Claims platforms had to wait for manually extracted data

Bitwise Solution

  • Deployed Llama 3 8B on Databricks to understand context across structured and unstructured policy documents.
  • Built a Medallion architecture that ingests DRS and OARS Parquet files into Bronze, extracts policy data through a prompt-driven process in Silver, and promotes validated outputs to Gold.
  • Used UDF-based batch inference to run the Llama 3 Instruct model at scale across Databricks clusters.
  • Created prompt-driven extraction to query policy data directly from XML and email content.
  • Added built-in validation checks for hallucination and contextual accuracy before data was promoted to the Gold layer.
  • Used PySpark and Hugging Face to support ingestion, transformation, and model integration end to end.
  • DRS/OARS Parquet files flow through a Databricks Workflow into Bronze, Silver, and Gold layers, with the LLAMA3 Instruct Model (batch inference) queried twice — once to extract XML policy data, once to validate results for hallucination and accuracy.

Tools & Technologies We Used

Databricks

DRS/OARS Parquet

Llama 3 8B Instruct

Databricks Workflows

Databricks Clusters

PySpark

Hugging Face

UDF Batch Inference

Bronze/Silver/Gold Medallion Architecture

Hallucination and Accuracy Validation

Claims Platform Integration

Key Results

Achieved up to 99% extraction accuracy on structured XML fields such as effective dates and coverages.

Achieved up to 97% extraction accuracy on key unstructured email fields such as limits and deductibles.

Increased efficiency by automating extraction and reducing manual effort and time to insight.

Scaled processing across large volumes of structured and unstructured policy data.

Reduced operational costs by minimizing manual data entry and validation.

Share

Download Case Study

Let's Engineer Your AI Advantage

AI-Driven Policy Data Extraction for a Fortune 500 Insurer