Get involved
We are looking for two things: clinical experts who can grade cases, and a real-world AI workflow that needs validation. If either fits you, we want to hear from you.
For clinicians: grade cases in your specialty
Every validation pipeline depends on gold answers: what a correct output must contain, decided by a clinician before any model runs. Right now, our two public examples (tumor board preparation on TCGA lung cancer, telemedicine triage on synthetic consultations) have author-set reference answers. No clinician has graded them yet. Every dossier on this site honestly reports that calibration status as not_collected.
That is where you come in. If you are a clinician (oncologist, internist, emergency medicine, family medicine, or another specialty), we need you to:
- Review 10-30 cases against a rubric. You read the clinical input, you write the gold answer, and you decide what the AI must get right.
- Grade a sample of AI outputs so we can check whether the AI judge agrees with you. This is the calibration step that makes the judge trustworthy.
- Time commitment: roughly 60-100 physician-hours for a full pack. You can start with 5 cases and see if the work fits you.
Your contribution is credited by name in the published dossier and on the site. The packs are open source (AGPL-3.0); your adjudication becomes part of the pack's gold standard.
For clinics and institutions: bring us a workflow
We have a working engine and two demonstration pipelines. What we need next is a real clinical AI workflow to validate with a partner. The ideal candidate:
- Text in, text out. The workflow takes clinical text and produces text a clinician can grade against a rubric.
- Already in use, or newly integrated. A live workflow with a real model is what the instrument is built for.
- Has a data route. De-identified retrospective cases, via your ethics committee or equivalent. We bring the pipeline, the rubric, and the engine. You bring the cases and the clinical expertise to adjudicate them.
Tumor board preparation is the natural first candidate (ground truth is generated for free by every board that meets). But any text-based clinical AI workflow fits: a scribe, a discharge-letter drafter, a referral summarizer, a triage agent.
What a partnership looks like
We bring
- The validrig engine, pack authoring, stress grid, AI judge, regression diff, QMS dossier
- A draft rubric for your workflow. We write it; your clinicians refine it.
- Two worked examples as templates
You bring
- One AI workflow you want validated
- A clinician (or small panel) to adjudicate 10-30 retrospective cases
- A data route for de-identified cases
Everyone gets
- A validated workflow with clinician-adjudicated ground truth
- A regression suite that re-runs on every model change
- A judge calibrated against your clinicians
- A co-authored public example on eval.dearauditor.ch
Get in touch
Tell us who you are and what fits. We reply within a week.
Prefer email? Write to aliaksei@dearauditor.ch directly.