Evaluating Clinical AI Coding Performance
Clinical-documentation technology company
The question
How accurately was the coding feature performing, where did it fail, and what was its financial value to a practice? A clinical-documentation technology company needed a repeatable basis for evaluating changes to E&M coding and understanding claims and denial patterns.
The work
We recreated the production coding workflow locally and evaluated prompts, models, and terminology approaches against a 713-case clinician-labeled reference set. The analysis examined overcoding and undercoding separately, including differences across specialties.
The development work separated clinical-note interpretation from parts of the coding decision. An LLM extracted facts; an experimental structured-MDM implementation applied explicit AMA Table-1 rules, including the two-of-three rule, data counting, and risk criteria. A separate ICD-10-to-SNOMED crosswalk constrained code selection to clinically relevant candidates.
We also analyzed payer denial notes and developed a financial-value model from a multi-practice study of 6,509 encounters.
Results and financial value
- Coding evaluations included a 69.6% production baseline and approximately 95% weighted accuracy for an evaluated configuration. These are distinct evaluation results; they do not establish a like-for-like improvement attributable to the rules implementation.
- In the structured-MDM experiment, explicit rules addressed 29% of the baseline rule-application errors. This describes that error subset, not a 29-point increase in overall accuracy.
- The terminology crosswalk won 83.8% of cases in a head-to-head comparison and reduced invalid-code selection from a 47.5% baseline.
- Denial-note classification covered 100% of notes, compared with 68% using keywords, at approximately $6.62 per 1,000 notes. Coverage measures whether a category was assigned; accuracy requires a separate review.
- The financial model estimated approximately $2,283 in value per 1,000 encounters, or $1.4M to $1.7M annualized at the large-practice volumes modeled. These estimates depend on encounter volume, payer mix, and reimbursement assumptions; they are not collected revenue.
What the client gained
A controlled evaluation workflow, a clearer account of the coding errors worth addressing, and a financial model for product and practice discussions. The work provided a way to test proposed changes before interpreting them as production improvements.