Build an AI-Powered Health Insurance EOB Decoder
## What We Built KeepMD EOB Decoder is an AI-powered tool that turns confusing health-insurance Explanation of Benefits (EOB) documents into clear, understandable information. EOBs contain important details—charges, insurance payments, adjustments, deductibles, copays, coinsurance, and patient responsibility—but they are often difficult for patients and healthcare teams to interpret. We built the decoder to bridge that gap. A user can upload an EOB, and the system: Step-by-step: 1. Reads the document — extracts text and structured data from PDFs or scanned EOBs using OCR when necessary. 2. Identifies key fields — patient responsibility, provider charges, allowed amounts, insurer payments, adjustments, deductible, copay, coinsurance, claim status, and service dates. 3. Interprets insurance terminology — converts codes and billing language into plain English. 4. Reconstructs the payment calculation — explains how the final amount owed was determined. 5. Highlights important findings — such as denied services, remaining deductible, unusual adjustments, or amounts that may need further investigation. 6. Presents the result clearly — instead of forcing users to decipher a dense insurance statement line by line. The goal isn't simply to extract text from an EOB. The goal is to make the financial meaning of an EOB understandable and actionable. ## How to Build It A practical architecture has five main layers: ### 1. Document ingestion Accept EOBs as PDFs or images. * Upload the document securely. * Detect whether it contains machine-readable text. * Use PDF text extraction when possible. * Fall back to OCR for scanned documents. * Preserve page and line locations so extracted information can be traced back to the source. ### 2. EOB structure extraction Convert the raw document into a consistent schema. For example: ```json { "claim": { "claim_number": "...", "service_date": "...", "provider": "...", "status": "processed" }, "financials": { "billed_amount": 1000.00, "allowed_amount": 650.00, "insurance_paid": 500.00, "deductible": 100.00, "copay": 25.00, "coinsurance": 25.00, "patient_responsibility": 150.00 }, "line_items": [] } ``` Use deterministic extraction and validation wherever possible rather than relying entirely on an LLM. ### 3. Insurance reasoning layer This is the most important part. The system should reason about the relationships between the amounts rather than simply repeating what the document says. For example: Billed amount → allowed amount → contractual adjustment → insurance responsibility → patient responsibility The decoder should be able to explain that chain in ordinary language: > “Your provider billed $1,000. Your insurance plan considers $650 to be the allowed amount. Insurance paid $500, and $150 was assigned to you through your deductible/coinsurance.” Every explanation should be grounded in the extracted EOB data. ### 4. AI explanation layer Use an LLM for tasks where language understanding adds value: * Explaining insurance terminology. * Summarizing the claim. * Answering questions about the EOB. * Explaining why the patient owes a particular amount. * Identifying potentially confusing or contradictory information. The LLM should not invent missing numbers or make unsupported billing conclusions. A strong pattern is: Document → structured facts → validated calculations → LLM explanation rather than: Document → LLM → answer That separation makes the system substantially more reliable. ### 5. User interface The UI should prioritize comprehension. A useful result screen could show: You may owe: $150 Then: * Provider charged: $1,000 * Insurance allowed: $650 * Insurance paid: $500 * Your deductible: $100 * Your coinsurance: $50 * Patient responsibility: $150 Followed by a plain-English explanation and the original EOB evidence supporting each number. ## Reliability and Safety Because EOBs contain sensitive health and financial information, privacy and security need to be part of the architecture from the beginning. Important safeguards include: * Encrypt documents in transit and at rest. * Minimize retention of uploaded EOBs. * Restrict access to patient data. * Log access and processing events appropriately. * Avoid sending unnecessary PHI to external services. * Clearly distinguish what the EOB says from what the system infers. * Provide source references for extracted values. * Validate financial calculations programmatically. * Never fabricate missing information. The product should also make clear that an EOB is not necessarily a bill. The decoder can explain what the insurer says the patient responsibility is, but that doesn't automatically mean the provider's bill is correct.
0 comments