GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Cursor Composer 2.5 Takes #3 on AA Coding Agent Index at $0.07 Per Task — 60x Cheaper Than Codex

Artificial Analysis has published its independent evaluation of Cursor’s Composer 2.5 on the AA Coding Agent Index, placing it third globally with a score of 62 — a 14-point gain over Composer 2 (48). The result puts Cursor ahead of every single-model agent that costs less than a dollar per task, and within striking distance of the frontier at a fraction of the price.

The Leaderboard Position

AgentModelIndex ScoreCost Per Task
Claude CodeClaude Opus 4.7 (max)66$4.10
CodexGPT-5.5 (xhigh)65$4.82
CursorComposer 2.562$0.07

The agents above Composer 2.5 cost 10x to 60x more per task depending on which Composer variant you use. At $0.07 (standard) and $0.44 (Fast), Composer 2.5 is the only coding agent scoring above 60 that sits on the cost-quality Pareto frontier.

Benchmark Breakdown

The 14-point index gain over Composer 2 is concentrated in the hardest problems:

  • SWE-Bench-Pro-Hard-AA: +35 points (12% → 47%). At 47%, Composer 2.5 matches Claude Opus 4.7 (max) in Claude Code on this subset.
  • Terminal-Bench v2: +2 points (64% → 66%)
  • SWE-Atlas-QnA: +3 points (69% → 72%)

Wall time averages 6.7 minutes per task in Fast mode, third-fastest among tested agents — behind Claude Opus 4.7 (medium) in Claude Code at 5.8 minutes and GPT-5.5 (medium) in Cursor CLI at 6.2 minutes.

What’s Under the Hood

Composer 2.5 continues training on Moonshot AI’s open-weights Kimi K2.5 base, with Cursor reporting approximately 85% of total compute coming from its own additional training and reinforcement learning on top of that base. The model is not exposed via any public API — it runs exclusively inside Cursor IDE and Cursor CLI.

Pricing is structured in two tiers: standard at $0.50/$2.50 per million input/output tokens, and Fast at $3.00/$15.00. The Fast variant runs roughly 30% faster (6.7 vs 9.3 minutes per task) but costs approximately 6x more per task ($0.44 vs $0.07).

The Practical Implication

For engineering teams running coding agents at volume, Composer 2.5 makes the cost-performance trade-off much sharper. At the frontier, you are paying $4-5 per task for a 3-4 point score advantage over a $0.07 agent. Whether that delta is worth 60x the cost depends on what you are automating.

The result also validates the strategy of building on strong open-weights foundations rather than training from scratch. Kimi K2.5 + Cursor’s RL investment has produced a model that is independently competitive — not just a cheap wrapper around an Anthropic or OpenAI API.