Before scoring the translations, we found four errors in our own reference cards.
Three cards had pulled facts from an earlier turn in the consultation. Another interpreted a Vietnamese phrase more strongly than the selected utterance allowed. We corrected the cards, retained the raw model output, and recorded why each label changed.
That failure belongs in the audit. A polished reference answer is not necessarily a trustworthy one.
A small audit designed to find errors
We drew the Vietnamese source text from the vietnamese config of Meddies Consultant, Meddies' synthetic consultation dataset. These are not records from real patients. The dataset card warns against assuming that its Vietnamese and English configs are balanced. This audit did not treat them as aligned translations.
We fixed the sampling frame to the first 100 rows returned by the public dataset API. From those rows, we purposefully selected 24 patient turns from 24 separate consultations. The cases probe details that translation can disturb: negation, onset, progression, severity, anatomy, laterality, radiation, triggers, relief, medication context, uncertainty, and the effect of a symptom on daily life.
This is an error-seeking sample, not a representative one. It cannot estimate a model's error rate or support a claim about how Vietnamese patients generally describe symptoms.
Each utterance was translated on its own. The model did not receive the target disease or earlier turns, because a translator should not recover missing meaning from hidden context. The run used gemini-3.1-pro-preview on 10 July 2026 with temperature 0 and low reasoning. We retained the serialized request, prompt version, run time, and hashes of the prompt and source cohort. This run did not set a seed, so another run is not expected to reproduce the output character for character.
The easy check passed
All 24 translations preserved the numeric tokens in their source text. Durations such as “2-3 weeks,” frequencies such as “3-4 times a week,” and pain scores such as “6/10” remained present. The mechanical result was 24 out of 24.
It was also incomplete.
One Vietnamese source said that lower-back pain spread to “mông và đùi phải.” The final modifier can plausibly attach only to the thigh or to the whole coordinated phrase. The English output said “the right buttock and thigh” and silently chose the second reading. The word “right” was still there. The ambiguity was not.
Our editorial screen marked six translations. We applied conservative post-edits to five: three wording repairs, one edit that removes an unsupported interpretation, and one naturalization that preserves the source ambiguity. The anatomical scope in the sixth source remains unresolved because a single unannotated English sentence would have to choose one reading. That case now requires clarification or an explicit translator note. These are editorial decisions, not clinically adjudicated findings.
Fluency and fidelity fail differently
Several candidates sound obviously translated. A respectful kinship term became the literal vocative “child.” A statement that congestion had worsened became “it is heavier.” A small amount of blood on tissue became “only spots a little on the paper.” The clinical quantities survived, but the patient voice did not.
Two candidates needed different repairs. “Nuốt vướng” describes difficulty or obstruction while swallowing, so we replaced “feel a lump when swallowing” with “swallowing feels obstructed.” For “người nhẹ đi một chút,” we used “I feel a little lighter than before.” The sentence now sounds natural without turning the patient's perception into measured weight loss. Both decisions still need independent clinical review.
That distinction is why a fluency score cannot stand in for a clinical review. Smooth wording can carry a scoped fact incorrectly. Literal wording can preserve a fact while making the patient sound unnatural. The two problems need separate labels.
The VLSP 2025 English-Vietnamese medical machine-translation shared task combined SacreBLEU with human judgment. It establishes a current Vietnamese evaluation context, but its corpus-level task does not answer whether a particular symptom detail survived in one patient utterance.
A 2025 npj Digital Medicine study of discharge-instruction translation used linguists, clinicians, and family caregivers to assess six languages. It found that performance varied by language and that human post-editing changed the result. Vietnamese was not one of the languages, so those findings cannot be transferred to our pairs. The relevant lesson is methodological: language-specific evaluation needs more than one kind of reviewer.
The review form has two directions
For every pair, the reviewer first reads from Vietnamese to English.
- Is every source fact present?
- Did the translation add a fact?
- Are negation, time, severity, laterality, and uncertainty attached to the same concept?
- Does the English sound like a patient utterance rather than a word-by-word reconstruction?
The reviewer then works backwards from the English output. Every claim in the translation must point to words in the source. This second direction catches additions that a source-completeness checklist can miss.
Errors should be recorded with the exact span, the changed fact, and a consequence category tied to the intended use. Awkward English is not automatically dangerous. Silently resolving an anatomical ambiguity may matter if a downstream system extracts laterality or a clinician relies on the translation during triage.
A 2025 implementation framework for machine-assisted translation recommends separating accuracy, fluency, terminology, local appropriateness, and error severity. A 2026 prospective English-Spanish evaluation used blinded bilingual clinicians who still received scenario context. Both studies sit outside the Vietnamese setting. We use them to shape the review process, not to imply equivalent performance.
What remains unverified
The raw output and the five post-edits still need independent bilingual clinical review. The reviewer should score the raw pairs without seeing our flags or repairs during the first pass. A second reviewer should score the same pairs before disagreement is resolved. Only then can we report which findings were confirmed, how severe they were, and whether the initial cards missed anything else.
Even after that review, this remains an exploratory audit of one run from one model. It does not compare translation systems, establish a failure rate, or show that the model is safe for clinical communication. A deployment evaluation would need a larger prespecified sample, independent reviewers, defined escalation rules, and cases drawn from the workflow where the translation will be used.
The current result is narrower and more useful. Matching every number did not preserve every ambiguity. Reading the model output was not enough either, because the reference layer had its own errors. Clinical translation needs both checks, with neither side trusted by default.
