GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
← Back to feed

OpenAI o1 Outperforms ER Doctors at 67.1% Diagnostic Accuracy in Harvard Science Study

OpenAI’s o1-preview reasoning model achieved 67.1% exact or very-close diagnostic accuracy on 76 real emergency department cases at Boston’s Beth Israel Deaconess Medical Center, outperforming two expert attending physicians who scored 55.3% and 50.0%. The study, led by Arjun Manrai from Harvard Medical School and Adam Rodman from Beth Israel, was published April 30 in the journal Science.

Blinded physician reviewers could not tell AI-generated assessments from human ones. The model was evaluated at three clinical decision points: initial triage on arrival, first physician contact, and admission to the medical floor or ICU.

Key Numbers

AssessorAccuracy
o1-preview67.1%
Expert Physician A55.3%
Expert Physician B50.0%

On a separate dataset of 143 complex clinical vignettes published in the New England Journal of Medicine, o1-preview included the correct diagnosis in its differential in 78.3% of cases and suggested a helpful diagnosis in 97.9% of cases. GPT-4 achieved 72.9% on the same vignette set. A broader comparison of 302 clinical vignettes found standard medical resources and textbooks returned 44.5% accuracy.

The AI was strongest at the initial triage stage, when information was most limited — precisely the decision point where speed and accuracy matter most and physician bandwidth is most constrained.

What This Is Not

o1-preview is two generations behind the current OpenAI frontier. No equivalent study has been published for o3, GPT-5.4, or GPT-5.5 in live clinical settings. The trial was designed to test diagnostic reasoning, not treatment recommendation or patient interaction — areas where AI performance gaps have been documented separately.

The researchers explicitly framed the findings as pointing toward collaborative care models, not physician replacement. That framing aside, the study is notable because it used real cases, not synthetic vignettes, and because the benchmark physicians were sourced from elite university medical institutions rather than a general pool.

Why It Matters Now

Prior benchmarks of AI in clinical settings — including HealthBench Pro, where GPT-5.4 scored 59.0 against a 42.9% physician baseline — focused on knowledge retrieval and recommendation quality, not live triage reasoning. This study establishes that the diagnostic accuracy gap exists in production ER conditions, not just curated test sets.

It also anchors a comparison point for any future clinical study. The next lab to publish a clinical AI trial will need to beat 67.1% on real cases to claim progress.