Back to Blog
July 10, 2026ExplainerClinical AI hub8 min read

How to Evaluate Clinical AI Before a Vietnamese Hospital Pilot

A hospital pilot needs a prespecified decision contract covering intended use, local cases, reviewers, failure thresholds, data paths, human control, stopping rules, and rollback.

Meddies Research

Clinical AI research at Meddies

How to Evaluate Clinical AI Before a Vietnamese Hospital Pilot

A hospital can complete a clinical AI pilot and still learn very little. This happens when the protocol is written around the product rather than the decision the hospital must make. A vendor supplies a convenient set of cases, average accuracy becomes the headline, clinicians try the interface for a few weeks, and the project ends without a defensible answer about safety or clinical use.

A useful pilot starts with a decision contract. Before the first case, the hospital records what the system may do, where it may be used, which failures matter, who will judge them, what data may move, how staff can overrule the output, and what evidence will lead to stopping, revising, or expanding the work. Another team should be able to repeat the evaluation from that record. It begins with Vietnam's current allocation of responsibility, not with the metric a vendor wants to headline.

Establish the legal actors before the evaluation

The Law on Artificial Intelligence 134/2025/QH15 has applied since 1 March 2026. It preserves human control and intervention and does not transfer human authority or responsibility to AI. For high-risk systems, providers and deployers have distinct duties around classification, monitoring, safety, intervention, and incidents. A hospital that controls professional use will usually act as the deployer, although the exact legal position depends on the system and deployment.

Clinical responsibility remains separate. Under the Law on Medical Examination and Treatment 15/2023/QH15, practitioners remain responsible for examination, treatment, and their professional decisions. AI may inform a decision. The clinician must still be able to review and reject the output and remains responsible for the clinical decision.

This does not make every clinical decision-support system high risk. Decree 142/2026/NĐ-CP, effective since 1 May 2026, requires classification before use and reconsideration when integration, function, intended use, modification, or deployment changes the risk. Decision 33/2026/QĐ-TTg has been issued but does not take effect until 15 August 2026. Its final healthcare appendix visibly lists two categories connected to surgery and robotic surgery rather than every diagnostic or treatment-recommendation system. Whether a particular tool is high risk still depends on its actual function and classification.

Write the intended use as an operational boundary

The intended-use statement should name the patient group, user, care setting, moment in the workflow, supported decision, and output. It should also name excluded uses. An evaluation of software that summarizes a chart before a visit does not support using it to recommend a diagnosis. A drug-interaction tool evaluated in one department has not earned a hospital-wide scope.

This boundary affects regulation as well as evaluation. Decree 98/2021/NĐ-CP includes software within the medical-device definition when the owner's intended purpose concerns diagnosis, prevention, monitoring, or treatment. That does not make every hospital AI product a medical device. It means the provider and hospital need a classification decision for the actual product and claim before the pilot, not after it.

DECIDE-AI is a consensus reporting guideline for small-scale early clinical evaluation of AI decision support. It asks investigators to report intended use, users, setting, workflow, inputs, outcomes, significant errors, safety risks, human factors, subgroup results, and system changes. It is not Vietnamese law and it is not a complete pilot protocol. Once the actors and legal path are clear, it is a strong check against measuring a model while ignoring the system around it.

Keep the cases and reference standard under hospital control

Start with retrospective data or shadow mode, where AI output cannot alter care. Sample cases from the workflow in which the system is meant to operate. Preserve ordinary cases, incomplete charts, conflicting notes, local abbreviations, out-of-scope requests, uncommon high-consequence cases, and patient groups that a convenient average could hide. Record the sampling rule and the size of each group.

The vendor may help prepare data, but it should not be the only party selecting cases or grading output. The hospital should own the locked case set and reference process. Specialty clinicians judge clinical correctness and severity. Pharmacists review medication tasks. Coding or BHYT staff review outputs that touch coding and payment data. Frontline users judge timing, workload, and whether the interface changes how decisions are made.

Disagreement between reviewers is evidence, not noise to delete. The protocol should say how disagreements are adjudicated and where more than one answer is clinically reasonable. If the reference standard is unstable, the report must show that uncertainty rather than converting it into a clean accuracy score.

Set thresholds from failure consequences

There is no universal accuracy threshold for clinical AI. A retrieval tool and a contraindication alert do not carry the same cost when they miss. For each intended use, the hospital should select metrics that expose the relevant harm, set minimum values before examining pilot results, and record who approved them.

One aggregate score is not enough. Report correct outputs separately from omissions, unsupported claims, wrong-patient data, inappropriate non-abstention, and errors within prespecified case groups. Where the system presents evidence, record source and citation failures separately. Workflow measures belong beside model measures, including when the tool was available, whether users followed the intended process, what they overruled, and how long review took.

The WHO regulatory considerations for AI in health emphasize intended use, data quality, validation, transparency, accountability, and risk management across the lifecycle. They do not supply one cutoff for every hospital. The threshold is a hospital governance decision tied to a defined use and the consequence of being wrong.

Make data handling and human review observable

Vietnam's Law on Personal Data Protection 91/2025/QH15 and Decree 356/2025/NĐ-CP have applied since 1 January 2026. The decree treats health status as sensitive personal data and establishes access-control, security, and impact-assessment obligations.

The pilot record should therefore map every data path. It should show which fields leave the EMR, what reaches the vendor, which subprocessors or cloud services participate, who can access the records, where logs are stored, how long each copy remains, when it is deleted, and whether any transfer leaves Vietnam. The parties must resolve their controller and processor roles, legal basis, processing purpose, and applicable assessment duties. A broad confidentiality clause is not a data-flow map.

Human oversight also needs an operating design. The reviewer must know when review is required, have access to the underlying chart and evidence, have enough time to check the output, be able to reject it, and know where to escalate an incident. A 2025 CORE-MD consensus paper argues that a named human reviewer does not by itself make an AI medical device low risk. Oversight depends on the user's expertise and actual ability to inspect inputs and recognize poor advice. The paper addresses European medical-device evaluation, not Vietnamese law, but it changes the pilot method. Test whether reviewers catch unsafe output instead of recording only that a reviewer was present.

Circular 32/2023/TT-BYT requires medical records to be accurate, truthful, and complete. An AI output does not move that responsibility away from the practitioner and the healthcare facility.

After shadow mode, any live phase should limit the users, department, patient group, and decisions the system can affect. Each use should preserve the system version, input, output, displayed evidence, final human decision, override, and reason when one can be recorded. A model or prompt change during the pilot creates a new evaluated version. It should not disappear inside the old results.

Define stopping, rollback, and exit before launch

The protocol should separate single-event stop conditions from rate-based thresholds. Wrong-patient output, unauthorized data disclosure, a safety-critical action taken without required confirmation, missing audit records, or a failed rollback deserve explicit stop rules. Performance below a prespecified threshold may lead to a narrower scope, a return to shadow mode, or termination, depending on the approved plan.

A workable rollback does more than hide the AI panel. It restores the previous workflow, blocks system calls, preserves evidence for investigation, notifies affected users, and identifies records that need review. A serious AI incident and a personal-data breach follow separate reporting routes. The AI Law divides urgent remediation, suspension or recall, recording, and reporting duties among providers, deployers, and users. The Personal Data Protection Law gives qualifying harmful breaches a 72-hour notification route. The stop plan should assign both routes rather than collapse them into one incident report.

The final packet should preserve the evaluated version, intended use, case-selection method, user characteristics, reference process, thresholds, results by failure type and subgroup, incidents, workflow changes, and every modification made during the pilot. From that record, the hospital can stop, revise and repeat within a narrower boundary, or proceed to a larger controlled evaluation. Passing does not prove improved patient outcomes. It shows that one version met a prespecified local evaluation contract under defined conditions, which Meddies treats as the minimum evidence a hospital should have before considering broader clinical use.

Review the intended workflow

Review the intended workflow and one synthetic medication-safety example, with the evidence boundary kept visible.

Book a demo