Skip to content
BenchEval

About

Started in the emergency department. Aimed at the frontier.

BenchEval exists because the person evaluating a medical AI should be qualified to disagree with it.

Founder

A clinician on both sides of the exam

BenchEval's founder is a practising emergency physician who has authored benchmark tasks for frontier AI labs — writing the cases that probe what models genuinely understand, and watching where confident answers quietly go wrong.

Emergency medicine turns out to be ideal training for this work. It is the specialty of undifferentiated problems, incomplete information and decisions that can't wait — precisely the conditions under which AI systems fail most interestingly, and most dangerously.

The gap was obvious from inside: labs and health-AI teams need expert evaluation at frontier quality, and physicians who can produce it have no serious front door into the work. BenchEval is that front door, built from both sides.

Mission

Make expert judgment the standard AI is measured against

Starting with medicine, because medicine is where the cost of a wrong answer is clearest — and expanding to other high-stakes domains as the approach proves out.

Adversarial, not adversary

We try to break models because that's how they get safe. The point of finding a failure is that a patient never does.

No inflated numbers

You won't find invented headcounts, logos or testimonials here. When we cite a figure, it will be real, and we'll show where it came from. Until then, we'd rather say 'early' than pretend otherwise.

Experts are the product

Verification, specialty-matching and fair rates aren't overheads — they're the entire reason the work is worth buying. We build for the physicians first.

Publish what survives scrutiny

Benchmarks and findings should hold up when someone hostile reads them. We write everything as if a sceptical clinician and a sceptical researcher will both check it — because they will.

Where we are

Early, and honest about it

BenchEval is founder-led and early: we are building our physician network, developing our first benchmark, and taking on scoped engagements where senior attention is a feature rather than a bottleneck.

If you're a lab or health-AI team that values depth over headcount — or a physician who wants in at the beginning — this is the right moment to talk to us.

For AI labs & healthcare AI

Put your model in front of physicians whose job is to break it.

Tell us what failure would cost you, and we'll scope an evaluation against it. Founder-led, so you talk to the person who designs the work.

Talk to us

For physicians

Your clinical judgment, applied where it shapes frontier AI.

Paid, remote, specialty-matched work — writing cases, grading model outputs and red-teaming systems that will meet patients.

Apply as an Expert