GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

KellyBench: Every Frontier AI Model Loses Money Betting the Premier League

General Reasoning released KellyBench on April 11 — an evaluation environment that exposes a gap standard benchmarks miss entirely: can an LLM maintain coherent strategy, adapt to new data, and manage risk over an extended timeline?

What KellyBench Tests

Agents are given a £100,000 starting bankroll and must navigate a 100–150 matchday Premier League season. Each round they predict outcomes, size bets using the Kelly criterion, and protect their capital from ruin. The test is explicitly adversarial to confident overconfidence — the same quality that lets models ace reasoning benchmarks can cause them to overbid and bust.

Results

ModelFinal ROI
Claude Opus 4.6-11.0%
GPT-5.4-13.6%
Several others-100% (bankrupt)

No model finished in profit. Claude Opus 4.6 was the most conservative — it lost money, but survived. GPT-5.4 followed at -13.6%. Multiple models went completely bankrupt, taking the full £100,000 to zero.

Why This Matters

Standard benchmarks measure isolated competency: can the model solve this math problem, complete this code task, answer this question. KellyBench measures something harder — sustained coherence under changing conditions, with compounding consequences for errors.

The failure mode isn’t stupidity. Models that score 89%+ on AIME 2026 still can’t beat a random walk on Premier League betting over a full season. They lack temporal consistency: a model that sizes bets well in week 3 may abandon that logic by week 17 when context has shifted.

The benchmark is available at gr.inc/releases/introducing-kellybench.

What It Says About Frontier Capability

Autonomous agents are increasingly deployed in finance, operations, and planning workflows where decisions compound over time. KellyBench is a controlled version of exactly those environments. The results suggest the limiting factor isn’t raw intelligence — it’s stability across an extended task horizon.

Claude Opus 4.6 at -11% ROI is meaningfully different from a model going bankrupt. That gap matters for any real-world deployment where “losing some money” is recoverable and “losing all money” is not.