GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.4 Runs a Drug Discovery Lab: 52% Yield Jump on Chan-Lam Coupling in Autonomous Chemistry Trial

OpenAI and Molecule.one ran GPT-5.4 inside a high-throughput automated chemistry platform called Maria and gave it an open-ended task: improve one of several important reaction classes. No target substrate was specified. The model identified the problem independently, proposed a solution, and the solution worked — confirmed by human chemists at bench scale.

Key Numbers

  • Mean reaction yield: 16.6% → 25.2% (+52% relative improvement)
  • Share of reactions above 30% yield: 15.6% → 37.5%
  • 88% of boronic acids tested showed improvement
  • 83% of sulfonamides tested showed improvement
  • Bench-scale validation: 11 of 14 substrate pairs improved, most >2x

What Actually Happened

GPT-5.4 was connected to Maria, Molecule.one’s agentic chemistry platform with access to an automated high-throughput laboratory. The model generated research proposals, designed experiments, analyzed experimental data, and proposed follow-up runs. Humans set the steering and grading prompts, selected which proposals to test, and made limited corrections to experimental plans. They did not specify a substrate class.

The task was Chan-Lam coupling — a reaction chemists use to form carbon-nitrogen bonds, useful across a wide range of drug scaffolds. GPT-5.4 independently identified primary sulfonamides as a difficult but high-value target class, then proposed that mild oxidants, including TEMPO, could improve yields on that specific substrate.

Two microliter-scale experimental cycles produced the improvement. Human chemists then ran representative reactions at bench scale. Those experiments confirmed the result held under practical lab conditions, with more than twofold increases in most substrate pairs tested.

Why It Matters

The sulfonamide group appears in dozens of medicines across antibiotics, diuretics, and anticonvulsants. Chan-Lam coupling on sulfonamides is a known bottleneck — the reaction is useful but finicky, and medicinal chemists need it to work reliably, not just in controlled screening conditions.

The model was not told where to look. It selected the substrate class, generated the mechanistic hypothesis (mild oxidant improving the coupling), and the hypothesis produced a measurable result that survived transfer to bench scale. That last step — bench-scale validation with real noise and practical volumes — is what separates this from computational prediction work.

This is OpenAI’s fourth published experimental result outside pure mathematics and computer science benchmarks. Previous cases: the unit distance problem (mathematics), gluon amplitudes in theoretical physics, and GPT-5’s contribution to lowering the cost of cell-free protein synthesis. The chemistry work is distinctive because it required physical validation in a wet laboratory, not proof or simulation alone.

The Agentic Chemistry Stack

Maria is built around a feedback loop: propose, run, analyze, revise. GPT-5.4 operated on two-cycle iterations within that loop. The scaffolding handles the lab hardware; the model handled the scientific reasoning. Human oversight remained at the proposal selection and steering level throughout.

OpenAI describes the system as “near-autonomous” — a phrase they are careful about. The humans in the loop were not passive observers. They were doing real scientific steering work. The model was generating the hypotheses.

The gap between “AI helps scientists” and “AI does science” is narrowing in specific, measurable ways. A 52% yield improvement on a drug-discovery coupling reaction, confirmed in triplicate at bench scale, is not a benchmark score. It is a chemistry result.