Skip to content
BenchEval

Benchmark · In development

EM-ARC: adversarial reasoning cases from emergency medicine.

Our first benchmark — working title EM-ARC (Emergency Medicine: Adversarial Reasoning Cases) — tests frontier models where they fail most consequentially: hard, ambiguous, multi-pathology emergency presentations.

Thesis

Saturated QA sets measure recall. Medicine is decided under ambiguity.

Frontier models score impressively on public medical QA — and still fail on the presentations that fill an emergency department: ambiguous symptoms, more than one active pathology, incomplete histories, and decisions that can't wait for certainty. EM-ARC is built from exactly those cases: physician-authored adversarial presentations, each independently reviewed by a second clinician and graded against a rubric that scores the reasoning — workup, escalation, uncertainty handling — not just the final diagnosis.

What makes it hard

Four failure modes, deliberately provoked

Each case targets at least one of the ways confident models — and tired clinicians — get emergency medicine wrong.

01

Anchoring

Cases are engineered to punish premature closure: an early, plausible diagnosis that fits the first half of the vignette and fails the second. Built so that pattern-matching to the obvious answer fails the way anchoring fails clinicians.

02

Medication interactions

Polypharmacy the way emergency departments actually receive it: anticoagulants plus new prescriptions, renal dosing on incomplete information, the interaction that only matters because of what the patient didn't mention.

03

Atypical presentations

The MI without chest pain, the sepsis that looks like a fall, the paediatric presentation textbooks describe in adults. Written from clinical experience, precisely because training data under-represents them.

04

Guideline recency

Recommendations that changed recently enough that a model trained on stale guidance answers confidently and wrongly. Each case records which guideline version it tests, so scores stay interpretable as guidance moves.

Method

Authored, reviewed, rubric-graded

01

Physician-authored

Every case is written from scratch by a practising physician — original, unpublished, and absent from any training set.

02

Dual review

A second, independent physician reviews each case blind: clinical accuracy, a defensible single best pathway, and no unintended giveaways.

03

Rubric grading

Scoring goes beyond the final answer: workup sequence, escalation timing, uncertainty handling, and the safety of the reasoning that got there.

Reporting format · Illustrative

What results will look like

A layout preview of the eventual report card. No evaluation runs exist yet — every cell below is deliberately empty.

Illustrative layout — no results yet. Benchmark in development.

Illustrative, empty results table showing the planned report format. The benchmark is in development and no results exist.
ModelDiagnostic accuracyEscalation safetyUncertainty handlingOverall
Frontier model ANo dataNo dataNo dataNo data
Frontier model BNo dataNo dataNo dataNo data
Frontier model CNo dataNo dataNo dataNo data

Early access

Labs and health-AI teams will run EM-ARC before it's public.

Early access means evaluating against the benchmark as it hardens — and helping decide what a defensible clinical reasoning score should measure. Physicians who want to author cases can apply to the network.