GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Arena Overhauls Agent Leaderboard: Per-Task Cost and Code/Chat/Work Categories, Drawn from 1.7M Sessions

Arena’s Agent Arena leaderboard got its most significant structural update yet on August 14, adding two features that change how the rankings should be read: per-task cost and category-level filtering across Code, Chat, and Work.

The changes reflect a scaling problem. With 1.7+ million user sessions analyzed in Agent Mode, Arena had the data to ask a more useful question than “which model scored highest?” — namely, “which model costs least to get a task done at each capability tier?”

Price Per Task, Not Per Token

Per-token pricing has always been a poor proxy for agent economics. Agents vary enormously in how many tokens they consume per task: a model that’s cheaper per million tokens but loops inefficiently can cost more per completed job than a pricier but leaner alternative.

Arena’s new cost view uses price per task as the primary measure. A Pareto frontier visualization surfaces the models that are cost-optimal at every performance level — the cheapest way to hit any given capability target. Models that sit off the frontier charge more than a comparably capable competitor.

This is the number enterprises actually care about. A coding agent run at scale isn’t budgeted in tokens. It’s budgeted in tasks, workflows, and deployments.

Three Task Categories from Real Usage

The Code, Chat, and Work split was derived from the session data, not from taxonomy assigned in advance. After clustering 1.7M actual user delegations, Arena identified the three dominant modes:

  • Code: Software development tasks, debugging, code review, implementation
  • Chat: Information retrieval, summarization, Q&A, general reasoning
  • Work: Document drafting, research, multi-step productivity workflows

Rankings within each category differ meaningfully from the aggregate leaderboard. A model that’s strong at Chat may rank fourth in Code. The category filter exposes that gap directly.

Why This Matters

The original Agent Arena launched in June 2026 with a single aggregate ELO and a causal tracing methodology. It was a significant methodological leap over simple benchmark scores, but it collapsed all task types into one number.

The new release acknowledges that agents are not a single task type. The model you’d use to review a pull request is not necessarily the same one you’d use to research a contract or answer a customer support query — and paying frontier prices for a commodity Chat task is provable waste.

The Pareto frontier + category filtering combination gives procurement-level information to a leaderboard that previously served mostly research-level comparison.