GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Three Labs, Three Arenas: GPT-6 Astra, Muse Spark 1.3, and DeepSeek V4.1 Flash All Enter Leaderboards in Four Days

Arena.ai’s changelog logged three distinct leaderboard additions in four days, spanning September 8-11. Each entry represents a model’s first appearance in a new evaluation category, opening it to blind human comparison against established competitors.

GPT-6 Astra: From Agent to Text Arena

GPT-6 Astra max joined Agent Arena on September 8, then entered Text Arena on September 11. The Text Arena addition is the more significant milestone. Text Arena is Arena’s largest and oldest benchmark — blind human voting across open-ended conversation quality — and GPT-6 Astra had not previously been eligible for it.

GPT-6 Astra arrives in Text Arena with a documented ceiling. Its ARC-AGI-3 score is 99.9%, the highest of any model tracked. On Terminal-Bench 4.0 it hit 57.9%, 2.1 points ahead of Fable 5.1. Its Artificial Analysis Intelligence Index sits at 53, tied with Claude Fable 5.1 for the top slot on the AA leaderboard. Human chat preferences may diverge significantly from benchmark rankings — that gap is what the Text Arena ELO will eventually resolve.

Muse Spark 1.3: Meta’s Fastest Frontier Model Gets Agent Eval

Muse Spark 1.3 entered Agent Arena on September 11, a few weeks after it debuted at #4 on LiveBench with an 81.6 overall score. What distinguishes it from the other Agent Arena entrants is inference speed: Artificial Analysis records it at 239 tokens per second in max mode, making it the fastest model in the top tier by a factor of four over GPT-6 Astra (60 tps) and Claude Fable 5.1 (67 tps).

Agent tasks are latency-sensitive. At 239 tps, Muse Spark 1.3 can complete multi-step reasoning sequences in wall-clock time that rivals take four times as long to process. Whether that speed advantage translates to better agent task scores depends on whether the bottleneck in Agent Arena tasks is model reasoning or inference throughput. The benchmark will answer that question.

Muse Spark 1.3 is priced at $1.60 per million input tokens in max mode — cheaper than Claude Fable 5.1 max ($7.63 per task) and GPT-6 Astra max ($3.26 per task) on the AA cost-per-task metric.

DeepSeek V4.1 Flash: Open-Weight Model Enters Code Arena WebDev

DeepSeek V4.1 Flash joined Code Arena WebDev on September 10, bringing open-weight competition to a leaderboard previously dominated by closed commercial models. The model recorded 77.3% on LiveBench Agentic Coding in the September run — the highest score on that metric across all models tested, outpacing Claude Fable 5.1 (66.1%) and GPT-6 Astra (57.3%).

The pricing difference is structural. DeepSeek V4.1 Flash is priced at $0.03 per million input tokens on its standard tier. At this price, running the same agentic coding workloads that cost $7.63 per task with Fable 5.1 max costs less than one-tenth as much with DeepSeek V4.1 Flash. If Code Arena WebDev results confirm the LiveBench Agentic Coding advantage, the cost-performance argument for open-weight deployment becomes hard to dismiss.

Why These Three Matter Together

The three additions in four days are not coincidental. They reflect each lab’s interest in establishing baseline positions across Arena’s evaluation taxonomy before ELO data accumulates. Arena ELO is slow to stabilize — it requires thousands of human votes — so entering early captures votes at a formative stage when rank volatility is highest.

The September window sees Anthropic and Meta competing in Agent Arena simultaneously for the first time, OpenAI extending GPT-6 Astra into Text Arena while it already holds a position in Agent Arena, and an open-weight model entering Code Arena at cost parameters that challenge the business case for closed-model web development deployments. The Arena data that accumulates over the next 30-60 days will be the first apples-to-apples human-preference comparison for this specific cohort.