Cursor Composer 2.5 Takes #3 on AA Coding Agent Index at $0.07 Per Task — 60x Cheaper Than Codex
Artificial Analysis has published its independent evaluation of Cursor’s Composer 2.5 on the AA Coding Agent Index, placing it third globally with a score of 62 — a 14-point gain over Composer 2 (48). The result puts Cursor ahead of every single-model agent that costs less than a dollar per task, and within striking distance of the frontier at a fraction of the price.
The Leaderboard Position
| Agent | Model | Index Score | Cost Per Task |
|---|---|---|---|
| Claude Code | Claude Opus 4.7 (max) | 66 | $4.10 |
| Codex | GPT-5.5 (xhigh) | 65 | $4.82 |
| Cursor | Composer 2.5 | 62 | $0.07 |
The agents above Composer 2.5 cost 10x to 60x more per task depending on which Composer variant you use. At $0.07 (standard) and $0.44 (Fast), Composer 2.5 is the only coding agent scoring above 60 that sits on the cost-quality Pareto frontier.
Benchmark Breakdown
The 14-point index gain over Composer 2 is concentrated in the hardest problems:
- SWE-Bench-Pro-Hard-AA: +35 points (12% → 47%). At 47%, Composer 2.5 matches Claude Opus 4.7 (max) in Claude Code on this subset.
- Terminal-Bench v2: +2 points (64% → 66%)
- SWE-Atlas-QnA: +3 points (69% → 72%)
Wall time averages 6.7 minutes per task in Fast mode, third-fastest among tested agents — behind Claude Opus 4.7 (medium) in Claude Code at 5.8 minutes and GPT-5.5 (medium) in Cursor CLI at 6.2 minutes.
What’s Under the Hood
Composer 2.5 continues training on Moonshot AI’s open-weights Kimi K2.5 base, with Cursor reporting approximately 85% of total compute coming from its own additional training and reinforcement learning on top of that base. The model is not exposed via any public API — it runs exclusively inside Cursor IDE and Cursor CLI.
Pricing is structured in two tiers: standard at $0.50/$2.50 per million input/output tokens, and Fast at $3.00/$15.00. The Fast variant runs roughly 30% faster (6.7 vs 9.3 minutes per task) but costs approximately 6x more per task ($0.44 vs $0.07).
The Practical Implication
For engineering teams running coding agents at volume, Composer 2.5 makes the cost-performance trade-off much sharper. At the frontier, you are paying $4-5 per task for a 3-4 point score advantage over a $0.07 agent. Whether that delta is worth 60x the cost depends on what you are automating.
The result also validates the strategy of building on strong open-weights foundations rather than training from scratch. Kimi K2.5 + Cursor’s RL investment has produced a model that is independently competitive — not just a cheap wrapper around an Anthropic or OpenAI API.