GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

OpenAI LifeSciBench: Best Frontier Model Solves 36% of Real Biology Tasks — Domain Expert Beats General AI by 10 Points

OpenAI published LifeSciBench on June 17, a 750-task benchmark designed to measure AI performance on the kind of biology research a PhD scientist actually does: multi-step, figure-laden, imperfect-evidence tasks with no clean single answer. Five frontier models were evaluated. None broke 37%.

The Benchmark

LifeSciBench was authored by 173 PhD-level scientists in biotechnology and pharmaceuticals, then validated by 453 additional expert reviewers. Each task required 90% expert consensus to qualify. The result is 750 tasks covering six workflow categories — evidence handling, analysis, design and optimisation, scientific reasoning, validation and operations, and translation and communication — across seven biological domains including genomics, medicinal chemistry, and clinical research.

Each task carries a rubric with an average of 25 atomic grading criteria, totalling 19,020 criteria across the full benchmark. Scoring is based on normalised rubric credit, not binary pass/fail. The pass rate column counts tasks where a model exceeded a minimum rubric threshold.

OpenAI ran all five models in a single-turn setting with unrestricted internet access.

The Results

ModelNormalised ScorePass Rate
GPT-Rosalind0.57636.1%
GPT-5.50.51925.7%
Gemini 3.1 Pro0.51523.6%
GPT-5.40.47920.7%
Grok 4.30.39913.0%

GPT-Rosalind is OpenAI’s domain-specialised life sciences model, first disclosed in April 2026. It leads on 386 of 750 tasks by per-task mean score. Gemini 3.1 Pro uniquely outperformed it on 214 tasks — the aggregate ranking obscures real task-specific variation.

The specialisation premium is real but bounded. GPT-Rosalind clears 36.1% to GPT-5.5’s 25.7%, a 10-point margin. On LabWorkBench, a proprietary wet-lab sub-evaluation, Rosalind scores 63.2% to GPT-5.5’s 55.8% while using 5.3% fewer tokens. On GeneBench (long-horizon genomics tasks), Rosalind scores 21.6% to GPT-5.5’s 20.4% with 31% fewer tokens — a marginal accuracy gain but a meaningful efficiency win. On MedChemBench (medicinal chemistry): 27.5% Rosalind vs 25.1% GPT-5.5.

Where Every Model Fails

The benchmark surfaces two structural weaknesses that general scaling has not closed.

Artifact tasks. When tasks require interpreting figures, PDFs, sequences, or structures alongside text, performance drops sharply. GPT-Rosalind falls from 45.1% on text-only tasks to 28.1% on artifact tasks. GPT-5.5 falls from 29.9% to 21.9%. Multi-modal biological reasoning is still an open problem.

Sequence and structure generation. Exact-output tasks — generate a DNA sequence, produce a protein structure — had the widest spread across models, ranging from 46.9% to 18.0% success depending on model and task. GPT-Rosalind’s gain over GPT-5.5 on construct/generate items was +0.001 normalised score. Domain specialisation does not help here.

The hard floor: 171 tasks (22.8% of the benchmark) passed by no model at all. An additional 261 tasks (34.8%) had a best-model pass rate below 20%. These are not edge cases. They represent the majority of the benchmark by any reasonable definition of difficulty.

The Conflict of Interest Problem

OpenAI designed LifeSciBench, validated it internally, and evaluated it — primarily against its own models. GPT-Rosalind leads the benchmark that OpenAI built to measure GPT-Rosalind. Gemini 3.1 Pro and Grok 4.3 are included as baselines, but Claude Opus 4.8 and Claude Fable 5 — the two highest-ranked models on the Artificial Analysis Intelligence Index — are absent. No explanation was given.

LifeSciBench is nonetheless a more demanding and thoughtful benchmark design than most in the biological domain, where the dominant alternative is multiple-choice question-answering. Tasks were externally authored and independently validated. The methodology is credible. The leaderboard composition is not.

What This Means

The ceiling on general frontier performance for biology research tasks is somewhere around 25-26% pass rate. Domain specialisation can push that to ~36%. Both numbers leave most real scientific work out of reach in a single-turn setting.

LifeSciBench was built as single-turn only. Real research is iterative. The benchmark measures a floor, not a ceiling. OpenAI is using it as the primary evaluation driver for future GPT-Rosalind development — the implication is that the 36% number will move, and that iterative agentic settings will close the gap faster than architecture changes alone.