Faisal RehmanAI ENGINEER
← All selected work

CareCloud Inc. (formerly MTBC) · Case study 06

Clinical documents into usable data.

Production document intelligence and billing automation at CareCloud Inc. (formerly MTBC).

My role
Data Scientist / Data Science Lead (prior: Jr. Data Scientist)
Period
October 2019–April 2021
Technology
Python · OCR · Clinical NLP · scikit-learn · XGBoost · SQL Server · SSIS · Power BI
60K+Healthcare documents per day
24h → 6hClaim cycle time
60% lowerManual work

The problem

Scanned and handwritten medical records, superbills, and claims created substantial manual work. Information had to move from inconsistent documents into structured systems before billing workflows could proceed.

The challenge was an operational pipeline: capture the information, structure it, and make it usable inside existing healthcare workflows.

What I owned

I also built a separate end-to-end patient–provider EHR assistant, spanning 3D symptom capture, illness summaries, diagnosis and CPT codes, and medication and lab-order suggestions.

I built document classification and extraction pipelines, worked on billing automation, and led junior data scientists. Related work covered clinical coding and claim-denial prediction.

The document pipeline ran in a healthcare environment with PHI-aware logging. Data handling was part of the system design.

The engineering decisions

Scanned records→OCR & classification→Structured extraction→Billing workflow

Design for inconsistent input. The pipeline processes scanned and handwritten records rather than assuming clean digital text.

Connect prediction to the workflow. Classification and extraction feed structured data into the billing process. The useful outcome is less manual work and faster processing.

Build the data foundation. SQL Server, SSIS, reporting models, and embedded analytics made the outputs accessible to operational teams.

The operating principle

A model metric is useful when it connects to the process around it. Classification accuracy and claim-cycle time measure different parts of the same delivery problem.

How the work was measured

Document classification achieved 98%+ accuracy. Operational metrics tracked manual work and claim-cycle time.

Related clinical NLP work reached micro-F1 0.81 on ICD-10 coding. A claim-denial model delivered 2× lift over rules. Those figures describe separate models, not the document classifier.

What changed

The production pipeline processed 60K+ healthcare documents per day. Document and billing automation reduced manual work 60%, while claim-cycle time fell from 24 hours to 6.

The work reinforced a principle that carries into my agent engineering: the model is one component in a larger operational system.

A public engineering summary. Results are scoped to the project and period described; implementation details are summarized at a high level.

NEXT CASE STUDY

Making a voice agent reliable enough to order.

Continue reading
07 WHAT’S NEXT

Hard problems.
Meaningful work.

I’m interested in teams bringing capable AI into the real world. Let’s talk about AI engineering, agent reliability, and products worth building.