Back to Blog
July 10, 2026ExplainerClinical AI hub5 min read

How to Test an AI Chart Summary for Dangerous Omissions

A proposed hospital protocol for finding clinically important omissions in AI chart summaries with source maps, bidirectional review, consequence labels, and blocking gates.

Meddies Research

Clinical AI research at Meddies

How to Test an AI Chart Summary for Dangerous Omissions

A chart summary can support every sentence it contains and still fail the clinician who reads it. The missing fact may be the one that changes urgency, medication choice, or follow-up. An output-to-source review will not find that failure because there is no output sentence to inspect.

Testing for dangerous omissions therefore starts with a coverage contract, not a model score.

Match the benchmark to the clinical job

An encounter note and a longitudinal chart summary do different work. The encounter note records one consultation: the presenting concern, relevant history, assessment, and plan. A longitudinal summary has to reconcile information across visits, medication changes, investigations, procedures, and unfinished follow-up.

That distinction changes the reference standard. A candidate coverage set for an encounter might include the presenting problem, time course, relevant positives and negatives, medications and allergies that affect the visit, assessment, plan, and safety-netting. A longitudinal review may instead prioritise active problems, changes in status, current medication and recent changes, allergies and reactions, important results and trends, major admissions or procedures, pending actions, and known gaps in the record.

Neither list is universal. Before evaluation, the hospital has to name the intended reader, decision, specialty, source window, and required fields. Emergency medicine and oncology should not inherit the same template because a vendor wants one benchmark.

A 2025 npj Digital Medicine framework evaluated consultation transcript-to-note generation by checking both hallucinations and omissions, recording the experiment setup, and grading errors by clinical consequence. The method transfers to longitudinal review; its content checklist does not.

Build a reference map and audit it twice

Two clinicians who understand the intended use should identify the smallest clinically meaningful facts required from each source chart. Every fact needs a source location, time, status, and criticality label. This reference must exist before the generated summary is reviewed, or the output quietly defines its own exam.

The map also records what reached the model. A 2026 npj Health Systems evaluation of EHR-integrated chart review found physician feedback about missing information and omissions associated with input token limits. If an important note never enters the model context, the failure belongs to retrieval or input construction. It should not disappear inside a summary-quality average.

The first audit starts from the source map. For every required fact, the reviewer checks whether the summary preserves its meaning. A topic match is insufficient. Negation, dose, timing, severity, uncertainty, and current status can determine whether a fact survived.

The second audit reverses direction. Every summary claim must point to the correct evidence. This catches unsupported statements, contradictions, cross-patient contamination, and facts joined across incompatible time points.

Keep the results separate. The proportion of written claims with evidence measures support. The proportion of required source facts represented in the summary measures coverage. A citation beside a sentence helps a clinician inspect that sentence. It cannot reveal a required fact that was never written.

Let consequences control the score

Counting every omission equally produces a tidy metric with little safety meaning. The 2025 framework called an error major when leaving it uncorrected could affect diagnosis or management. A hospital can adapt that rule to the decision under test and specify whether the error could change urgency, treatment selection, contraindication handling, monitoring, or the next required action.

A minor omission is not expected to change those decisions, although it may still reduce usability or force more chart review. The label depends on clinical judgment, so reviewers must record why they assigned it.

Use two clinicians to review every source-summary pair independently, followed by adjudication when they disagree. The 2025 study used at least two clinician annotators and senior-clinician consolidation. That is a strong default, not the only defensible staffing model. Preserve agreement, adjudication count, and the reasoning behind every confirmed major error in the report.

Results also need clinical slices. Report omissions by required field, specialty, chart length, source type, and difficult case group. One aggregate rate can hide a system that preserves laboratory results while dropping allergies, or works on short encounters but fails across years of notes.

Decide what failure means before results arrive

There is no universal safe omission percentage. The hospital's clinical and safety reviewers must set thresholds before seeing the results, based on the proposed use and the amount of verification left to the clinician.

For a conservative pilot, we propose five gates.

  1. A confirmed major omission stops that configuration. Fix the workflow, regenerate, and retest before it advances.
  2. A missing required section or incomplete promised source window is blocking even when the average score remains high.
  3. A broken or incorrect evidence pointer fails traceability and must be repaired before further review.
  4. Unadjudicated disagreement, thin coverage of a clinical slice, or an untested challenge group leaves the evidence incomplete.
  5. Minor-error and coverage thresholds must be stated for each required field, not only as a portfolio average.

Pin the model, prompt, generation settings, retrieval rules, source truncation, and run date. A material change creates a new configuration and requires another held-out evaluation. Passing that test is still not the end. The system needs a monitored pilot in the department, chart type, and workflow where it will be used.

What remains unproven

This is a proposed evaluation protocol, not a Meddies benchmark result. We have not published an omission rate for Meddies chart summaries, and this protocol has not been shown to improve patient outcomes.

Its value is narrower and practical. It makes the hospital name what must be covered, where the evidence lives, who reviewed each case, how disagreement was resolved, and which failure blocks deployment. Without that record, a high score cannot answer whether the summary is safe for the decision being tested.

Review the intended workflow

Review the intended workflow and one synthetic medication-safety example, with the evidence boundary kept visible.

Book a demo