GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

AI Wins the Metaculus Cup for the First Time, Beating Rivals That Spent $15M Combined

On September 5th, a bot built by Jeffrey Liang won the seasonal Metaculus Cup, the first time an AI has taken the top spot in the competition. Liang says he spent fewer than 150 hours on the project and approximately $2,000 in compute. His bot beat four venture-backed AI forecasting startups that have collectively raised more than $15 million. AI systems also placed second and fifth in the competition, with humans finishing third and fourth.

The Metaculus Cup runs over a four-month period. Participants predict outcomes across a wide range of questions on geopolitics, science, markets, and current events, scored on the distance between their prediction and the actual result over time. Liang’s bot is also currently leading a separate AI-only Metaculus competition with a $50,000 prize pool.

ForecastBench: frontier AI now at superforecaster parity

The Metaculus result sits within a broader shift documented by the Forecasting Research Institute’s ForecastBench, the primary benchmark comparing AI forecasting systems against human superforecasters on a rolling question set.

FRI’s July 2026 analysis found that AI systems had reached parity with superforecasters. On the tournament leaderboard, Cassi AI holds the top position and is the first system to outrank superforecasters on market questions, a category considered a harder test than dataset questions because it requires judgment on novel, one-off events rather than base-rate analysis or data lookup. xAI and Google DeepMind have also submitted systems that rank at superforecaster parity on the main tournament leaderboard. Across FRI’s preliminary leaderboard, 17 submissions now rank above the superforecaster median.

The cost gap

The economic gap between human and AI forecasting is widening sharply. A professional human superforecaster forecast can cost more than $10,000 and require a week or more of work. FutureSearch, one of the commercial AI forecasting services active in the space, operates at a few dollars and ten minutes per forecast.

The returns are starting to track. FutureSearch has seen a 6% return since June on a $100,000 Kalshi portfolio. A developer behind Preseen, one of the competition’s entrants, turned an initial $35 into $1.94 million over seven months on Kalshi, the sixth-best return in the platform’s history.

Where humans still have an edge

Forecasting over years rather than months remains an area where AI has not been tested sufficiently long to establish clear results. Human superforecaster comparisons on ForecastBench rely on predictions last collected in 2024, making extrapolation less reliable over time. Many AI system confidence intervals on ForecastBench still overlap substantially with superforecasters, placing the frontier genuinely close rather than conclusively ahead.

Human forecasters increasingly treat AI as a collaborative tool. Yann Rivière, a professional forecaster at Mantic and a Metaculus Pro Forecaster, described working with AI systems as he would with another professional forecaster, using AI output to improve his own predictions rather than replace them.