GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

AI Hovers Near Chance at Predicting Scientific Discoveries — Paper Tests 4,760 Events and Finds a Hard Ceiling on Foresight

A paper titled “Forecasting Scientific Progress with AI” (arXiv 2605.22681), published this week, tests frontier AI models across 4,760 scientific events and finds a structural gap between two things that look similar but are not: recognizing a plausible research path, and predicting what will actually happen.

The models are strong at the first. They fall near chance at the second.

Recognition vs Prediction

The distinction matters. When a question is presented in multiple-choice form — given a scientific problem, which of these approaches is plausible — models perform well. The answer is nearby in the form of options, and the model can reason toward the most coherent one.

But when asked whether a specific discovery will be realized, when it will happen, and what method will produce it, models drop toward random performance. The signal is not just weak — it is near chance.

Models also show a systematic bias when predicting timing: they push predicted dates too far into the future. Given historical context, they improved modestly, but did not become reliable forecasters of when scientific progress would arrive.

Why This Matters

The common framing for AI in science is a version of “AI has read all the papers, so it can find new connections.” That framing is not wrong — models do surface plausible research directions, and they do this better than most search-based approaches. AlphaFold, AlphaEvolve, and similar systems have demonstrated real scientific output.

But those systems are not doing open-ended prediction. They are running structured search within a defined problem space. The new paper is testing something different: can models forecast the open frontier of science — the what, when, and how of discoveries that have not yet been made?

The answer, across 4,760 events, is largely no.

What AI Can and Cannot Do in Science

The paper draws a clean line. Models are good at:

  • Recognizing whether a proposed research path is scientifically coherent
  • Selecting among plausible options when options are provided
  • Summarizing the current state of a field

Models are weak at:

  • Predicting whether a specific discovery will be made in the next N years
  • Predicting when progress will arrive
  • Identifying the correct method before anyone has tried it

The authors gave models additional historical context and found modest improvement but no fundamental change. More data about past science does not translate into reliable foresight about future science.

This is consistent with how frontier models handle other domains requiring genuine forecasting rather than pattern matching over known distributions: financial markets, geopolitical events, long-range planning. Models that perform well on benchmarks built from historical data do not necessarily generalize to predicting events outside that distribution.

For AI-in-science claims — and there are many — the distinction between plausible exploration and genuine prediction is the one worth tracking.