Telemedicine consultation triage
An AI works through a remote consultation — complaint, history, attachments — and proposes diagnosis, workup, and treatment. This pipeline grades three Gemini generations against gold answers on original synthetic consultation cases, and shows what an honest FAIL looks like.
Data: Cases are 10 original, independently authored synthetic consultations. Their design follows the published D.O.T.S. specification (arXiv:2603.25821); no data from the Doctorina MedBench dataset is included.
The clinical task
In a telemedicine service, the first conversation often decides everything: what the patient is asked, what gets recorded, and whether the case is routed to a doctor now, later, or to self-care. Vendors increasingly put an AI on that first conversation.
The workflow under test: the model works through a consultation case — presenting complaint, patient history, attached investigations — and produces a clinical assessment with diagnosis, differentials, workup, and first-line treatment.
What can go wrong
- Critical omissions — the one finding that should change the assessment gets summarized away, especially when part of the input is missing or degraded.
- Overconfident answers — a fluent, plausible assessment on a case where the safe move was escalation.
- Version-to-version drift — the same vendor's next model release handles the same conversation differently, and nobody measures it.
How we measure it
Each case carries gold answers: the clinical facts an assessment must contain and the management that is defensible. A pinned AI judge scores every output against that rubric, under a perturbation grid that withholds or degrades parts of the input. Acceptance thresholds are pre-registered in the pack — the pass/fail line is fixed before any model runs, so a failure cannot be negotiated away after the fact.
Three Gemini generations were run side by side on the same pinned inputs, so the pipeline answers a concrete question: did the new model get better or worse on your cases?
Findings
The rig failed this workflow — and that is the finding. On the comparison battery's baseline condition (the intended, unperturbed inputs), gemini-2.5-flash missed critical findings in 6.7% of graded checks. The pre-registered limit is 5%. Verdict: fail — even though the same run's mean quality score, 0.92, comfortably cleared its own threshold of 0.80. On the input-contract battery, the same model re-graded on the intended inputs recorded 0% critical omissions and passed; across that battery's full perturbation grid (inputs withheld or degraded), the critical omission rate rose to 12%.
How much weight this carries: the case bank holds 10 synthetic cases, so the confidence interval around that omission rate is wide — it spans the limit in both directions, and a single case moving changes the headline. It has: in the previous archived run of this pack lineage the excess omission rate appeared on the contract battery while the comparison baseline passed; on this re-run of identical case content the verdict flipped batteries. A live LLM judge re-grades afresh each run, and a pre-registered gate is a point comparison by design — so each FAIL stands as recorded, and the instability itself is measured and archived rather than averaged away. What a bank this small cannot tell you is how far from the limit the true rate sits. It is a demonstration of the method on a worked example, not a safety verdict on any product.
That combination is the whole story: a model can look good on average and still be unsafe at the edges. A validation pipeline that cannot say FAIL is marketing; this one just did.
Model comparison: on the pinned head-to-head comparison battery, the three Gemini generations scored 0.920 (gemini-2.5-flash), 1.000 (gemini-3.5-flash), and 0.980 (gemini-3.7-flash) — the two newer generations passed every acceptance gate, while gemini-2.5-flash failed its critical-omission gate as above. Mean score deltas across the two version steps: +0.08 and -0.02.
What it means for you
Triage AI is exactly where "it seemed fine in the demo" is not a safety argument. This workflow crossed a pre-registered safety line on critical omissions, on clean inputs — while the same model, re-graded in a second battery, landed under the line, and the miss rate rose well past it once inputs were degraded. A validation pipeline gives a telemedicine provider that answer as a standing capability: on which cases does it fail, under what conditions, how stable is that verdict, and did the last update change it?
Because the case bank is synthetic and original, the whole pipeline — cases, rubric, judge prompts — can be published and inspected without touching patient data.
The validation report
Every pipeline run ends in a dossier like this one — the document you hand to your quality manager or auditor. It is generated from the pinned run, never edited by hand.
Validation dossier — telemedicine consultation triage (consultation pack)
run 212dc026df469e86Draft evidence, unsigned — generated by the engine from pinned inputs; requires human review and signature before use as QMS evidence. Release recommendation: not_approved_for_release
Demonstration pack for agentic and LLM-based clinical consultations on original, independently authored synthetic cases; the case design follows the D.O.T.S. specification of arXiv:2603.25821 (Doctorina MedBench) — no Doctorina data is included. Formulates histological / clinical diagnoses, retrieves differential diagnoses, determines necessary diagnostic workups, recommends first-line treatments, and avoids contraindicated or hazardous management. FOR DEMONSTRATION / DISCUSSION WITH CLINICIANS — not a clinical device.
1. Validation summary (baseline, intended input)
1 / 5 rubric items passed · 4 failed · condition: baseline (intended input)
| Requirement (rubric item) | Result |
|---|---|
| item_diagnosis | pass |
| item_differential | fail |
| item_investigations | fail |
| item_treatment_safety | fail |
| item_red_flags | fail |
2. Acceptance criteria
| Threshold | Value | Limit | Result |
|---|---|---|---|
| critical_omission_rate_max | 0.0667 | 0.05 | FAIL |
| mean_score_min | 0.92 | 0.8 | PASS |
3. Input contract (characterization)
Information value = mean score drop when the element is withheld. Characterization, not pass/fail.
| Element | Modality | Necessary | Information value |
|---|---|---|---|
| initial_presentation | text | — | not measured |
| patient_history | clinical_history | — | not measured |
| investigation_attachments | diagnostics | — | not measured |
4. Judge calibration
Status: not_collected — no physician double-grading recorded yet.
5. Attestation
Generated by validrig 0.1.0. Pinned-inputs hash: 212dc026df469e8684ccaea4bc1c1b13870ea1c7aa3e44008b23ea4bce3a3642
SUT and judge model_version are author-asserted. For a hosted endpoint that may update silently, pin a local model or record the provider's returned model id to keep this attestation trustworthy.
| Pinned input | Value |
|---|---|
| pack_id | clinical-eval-consultation |
| pack_version | 0.1.0 |
| pack_hash | 47a0f8b075b52c8364f74faee0e4886ebb37b2c834988ec45db9716c016db429 |
| battery_id | gemini_compare_all |
| battery_version | 1 |
| sut_id | gemini-2.5-flash |
| sut_hash | 0d2134aa001cd98ec16154a17af22b235af363e9b5c407cc3fe7dc31dd02f044 |
| judge_id | geval-gemini-flash-lite |
| judge_version | 1 |
| seed | 1 |
| engine_version | 0.1.0 |
6. Signatures
unsigned — meaning: Approved V&V Evidence and Report. Signer roles: medical_reviewer, quality_reviewer. Sign by review + anchoring to an immutable release (github_immutable_release (planned)).
DearAuditor Eval · generated 2026-09-01T15:03:30+00:00 · validrig engine
generated from run 212dc026df469e86
This page shows methodology and worked examples. It is not regulatory advice.