Back to Blog
July 10, 2026ExplainerClinical AI hub8 min read

A Relevance Benchmark for Prescription Alerts Clinicians Will Read

A proposed protocol for testing whether prescription alerts are correct, relevant, actionable, and worth interrupting a clinician. No Meddies benchmark result is claimed.

Meddies Research

Clinical AI research at Meddies

A Relevance Benchmark for Prescription Alerts Clinicians Will Read

A prescription alert can be factually correct and still be the wrong interruption.

The drug pair may appear in a trusted reference, yet the message may arrive for the wrong route, the wrong patient, or at a point where the prescriber cannot act. A different alert may occur rarely but carry a consequence that makes one missed case unacceptable. Counting matched rules treats these failures as equivalent. Clinically, they are not.

This article sets out a proposed benchmark for alert relevance. It is a protocol, not a completed evaluation. There is no finished case set, pass threshold, reviewer result, or Meddies performance claim behind it.

Start with the decision the alert is allowed to change

Before constructing cases, the benchmark owner must declare how the output will be used. Is it a reference shown without interrupting the order? A warning the prescriber can dismiss? A hard stop that requires a documented action?

Those products carry different error costs. A low-value message in a reference panel is inconvenient. The same message as a hard stop can delay treatment and teach users to dismiss future warnings. A missed high-harm contraindication has a different cost again.

The intended use therefore determines which errors matter most and which threshold could be acceptable. It must be written before the test set is scored. The FDA's January 2026 clinical decision support guidance is a United States regulatory document, not Vietnamese law, but its separation of software functions by intended purpose is a useful design discipline. A team cannot validate an advisory tool and silently deploy the same output as a directive.

The ONC SAFER guide for computerized order entry and decision support makes the operational point more directly. Interruptive alerts should be limited to the most critical warnings, while alerts in general should be patient-specific, actionable, and delivered at the relevant point in workflow. The guide does not supply a universal cutoff. It tells each organization what its own cutoff must protect.

Build cases that expose dependence on context

A reproducible case needs more than medications and an expected answer. Its record should include the knowledge source and version, date of verification, prescription under review, patient data available to the system, information deliberately withheld, and workflow point where the alert would fire.

Vietnam's contraindicated drug-interaction list issued under Decision 5948/QD-BYT can anchor one family of cases. It cannot define the whole benchmark. The list focuses on contraindicated interactions, while prescription safety also includes allergies, dose limits, renal and hepatic function, pregnancy, therapeutic duplication, and risks that require monitoring rather than avoidance.

Each declared risk family should contain cases where an alert is warranted, cases where it is not, and close counterfactuals. A counterfactual pair changes one clinically meaningful element while holding the rest of the record stable. If a dose warning remains unchanged after the relevant laboratory value changes, the failure is visible. If a route-specific interaction fires regardless of route, the benchmark reveals that the system matched drug names without understanding the order.

Cases used to tune rules or prompts must remain separate from the held-out cases used for the final report. Every change to a reference source, terminology mapping, drug-name resolver, patient-data schema, or alert rule creates a new benchmark version. Without that manifest, a rerun after an upgrade cannot explain whether the system improved or the test moved.

A defensible label needs two professions

A binary answer to "is the interaction present" is only the first label. Reviewers must also judge what the interface should do in the declared context.

A practical rubric records the potential severity of harm, time available to intervene, applicability to this patient and order, availability of a useful action, strength and scope of the supporting source, and missing data that could reverse the judgment. The final label then describes the appropriate response, such as interrupt, present without interruption, suppress, or request missing information.

The exact names belong to the product contract. Their definitions do not. Each label needs a written inclusion rule, clinical rationale, source, and description of the harm caused by assigning it incorrectly.

Collapsing every dimension into one score hides the asymmetry the benchmark is meant to preserve. Correctly suppressing many minor messages cannot compensate for missing a case that the review panel classified in advance as an unacceptable omission. Results should be reported by alert class, with errors weighted by their clinical consequence rather than reduced to an overall average.

Creating those labels requires more than one clinical viewpoint.

Pharmacists and physicians inspect different parts of the same prescribing decision. A pharmacist may identify an interaction, dose concern, monitoring need, or product detail. A physician must integrate that evidence with the diagnosis, treatment goal, urgency, and available alternatives. Both perspectives are required.

Each case should receive independent pharmacist and physician review before either reviewer sees the system output. The benchmark should retain the initial label, cited evidence, and rationale from each reviewer. A separate adjudicator or panel can resolve the production label under a predefined process, but adjudication must not erase the disagreement that preceded it.

Disagreement is diagnostic evidence. It can reveal an ambiguous source, a rubric that combines two decisions, a missing specialty perspective, or a case whose answer depends on information not represented in the data. Reporting only final consensus makes a fragile benchmark look cleaner than it is.

The analysis should therefore break disagreement down by risk family, proposed intervention, patient subgroup, and missing-data pattern. Reviewers who wrote the rule or constructed the case should be separated from final judging where feasible. Any remaining overlap belongs in the limitations.

Passing is not an average

No single metric can validate an alert system. An interruptive configuration needs high detection of predeclared high-harm cases, but it also needs enough precision to avoid spending clinician attention on low-value interruptions. A passive configuration has a different burden profile and may be judged more heavily on timing, clarity, and actionability.

The report should include sensitivity within each harm class, precision for interruptive alerts, alert burden per prescribing opportunity, harm-weighted false negatives, subgroup performance, reviewer agreement, and uncertainty around the estimates. Pass thresholds must be locked before the held-out run and justified against intended use.

A 2025 prospective single-arm study evaluated diagnostic recommendations embedded in medication alerts across 23 specialties at one Taiwanese hospital. Acceptance varied markedly by specialty. The uncontrolled, single-site design cannot establish improved clinical outcomes, and diagnostic recommendations are not the same as interaction warnings. Its useful lesson is narrower: a pooled acceptance metric can hide a workflow mismatch concentrated in one specialty. The benchmark must define specialty strata before analysis rather than discover them after an average looks poor.

Some conditions should fail the configuration regardless of its average. These include a miss in a predeclared unacceptable class, treating missing patient data as a negative finding without disclosure, material degradation in a clinically important subgroup, reference or terminology drift that changes the answer without a version change, and inability to reproduce a run from the published manifest.

The benchmark also needs regression cases. SAFER recommends testing new and existing decision-support rules after changes and major EHR upgrades. A release that fixes one alert while silently breaking an older rule has not passed.

The live signal begins after the benchmark

Offline cases tell us whether the system can produce the expected response under controlled conditions. They do not tell us whether a real alert is useful at the point of prescribing. That requires a later simulation, shadow deployment, or governed live evaluation.

When a clinician overrides an alert, the reason should be captured in a structured form that fits the clinical context, with room for additional explanation when it matters. The AHRQ override taxonomy implementation report shows why. Override reasons can expose logic defects, unsuitable inputs, poor timing, evidence disputes, patient preferences, and feasibility constraints. They are useful only when the organization reviews them and changes the system when warranted.

An override rate alone is not a relevance score. A justified override can show that the clinician supplied context the rule lacked. An accepted alert does not prove the resulting decision improved care. Evaluation has to connect the reason, subsequent action, available patient context, and independent review.

Requiring a reason also creates burden. If no team owns the analysis or no improvement process follows, a mandatory dropdown becomes paperwork and the data deteriorate. The benchmark protocol should state which alerts require an override rationale, who reviews it, and what evidence can trigger a rule change.

Publication should make the result reproducible

A reviewable benchmark release should include the source manifest, case-selection criteria, data schema, labeling guide, adjudication method, initial disagreements, locked thresholds, metrics by subgroup, complete error taxonomy, and change log. Where clinical data or licensed references cannot be released, the team should still publish the schema, selection or construction method, evaluation code, and a shareable sample set.

ONC describes clinical decision support as timely, person-specific information that fits the care workflow. An offline benchmark can test parts of that contract. It cannot establish improved patient outcomes, safe adoption at a particular hospital, or suitability across specialties.

That limitation is not a weakness to hide. It tells us what the next evaluation must do. The benchmark earns trust when another team can reconstruct why an alert interrupted, why a miss counted more than a nuisance warning, and what evidence would force the configuration to fail.

Review the intended workflow

Review the intended workflow and one synthetic medication-safety example, with the evidence boundary kept visible.

Book a demo

References

  1. Clinical Decision Support SoftwareU.S. Food and Drug Administration (2026)
  2. SAFER Guide 3 on computerized provider order entry and decision supportASTP / Office of the National Coordinator for Health IT (2025)
  3. Contraindicated drug interactions issued under Decision 5948/QD-BYTVietnam National DI and ADR Center (2021)
  4. Evaluation of Diagnostic Recommendations Embedded in Medication AlertsJournal of Medical Internet Research (2025)
  5. Override Reason Taxonomy ImplementationAgency for Healthcare Research and Quality (2025)
  6. Clinical Decision SupportASTP / Office of the National Coordinator for Health IT