What problem were we solving?
Manual document processing was a bottleneck affecting patient care timelines. Staff spent 70% of their time on paperwork instead of patient-facing activities.
Key Constraints
HIPAA compliance required for all data handling
Documents arrived in varied formats (PDF, fax, email, scanned images)
Accuracy requirements above 99.5% for medical coding
Integration with existing Epic EHR system
What paths were on the table?
OCR with rule-based extraction
Large language model with fine-tuning
Hybrid approach: OCR + ML classification + LLM extraction
Outsource to a BPO vendor
The Decision
Built a hybrid AI pipeline combining OCR for text extraction, a fine-tuned classifier for document type identification, and an LLM for structured data extraction with human-in-the-loop validation.
Measurable business outcomes
Processing time per document dropped from 45 minutes to 2.5 minutes. Accuracy reached 99.7% for structured extraction. Staff reallocated 60% of their time to patient care.
Have a similar challenge?
Tell us about your situation. We can share what worked and what did not from this project.

