Qwen3.8 Max Takes AA Agentic Index #1 at Intelligence 56 — but Costs $1.14 Per Task, Twice Its Predecessor
Alibaba’s Qwen3.8 Max has reached the top of Artificial Analysis’s agentic index, scoring 56 on the AA Intelligence Index — a position that puts it ahead of every US lab except Anthropic and OpenAI. It is the first Chinese model to claim the #1 position on AA’s agentic ranking.
The result has a catch: getting there cost more than twice as much per task.
The Benchmark Position
On Artificial Analysis’s agentic methodology, Qwen3.8 Max now holds the leading position at Intelligence Index 56. That places it above Kimi K3 (Moonshot AI, Intelligence Index rank #4 on AA overall), GLM-5.2, and the rest of the non-Anthropic/OpenAI field globally. Kimi K3 carries more total parameters at 2.8 trillion versus Qwen3.8 Max’s 2.4 trillion, and has a more complete independent benchmark record, including the AA front-end coding leaderboard.
On LiveBench, Qwen3.8 Max scores 78.5 overall with 64.6% on the agentic coding subcategory — results already reported when it launched. The AA agentic index result reflects a different evaluation methodology: AA’s framework measures multi-step task completion across its proprietary τ³-benchmark suite, which includes banking, legal, and research domains. The outcomes can diverge substantially from raw coding benchmarks.
The Cost Problem
Agentic index results measure the ceiling. Cost figures measure the floor. On the AA agentic framework, Qwen3.8 Max costs $1.14 per completed task. Its predecessor Qwen3.7 Max cost $0.53 per task on the same framework. Kimi K3 Max, which is competitive on intelligence metrics, costs $0.86 per task.
That cost structure is not an accident of the model architecture — it reflects how Qwen3.8 Max achieved the benchmark result. The model’s internal capability score (on Alibaba’s own multi-benchmark composite) rose from 0.474 to 0.725 generation-over-generation. But it appears to be taking more turns to complete multi-step agentic tasks than its predecessor, driving the per-task token count — and therefore the cost — up disproportionately.
The practical consequence: enterprise customers evaluating Qwen3.8 Max against Qwen3.7 Max or Kimi K3 for agentic workloads face a 2.1x to 1.3x cost premium for the extra performance. Whether that’s worth it depends entirely on what the incremental intelligence buys on the specific task.
The Competitive Context
The AA agentic index ranking is meaningful because it uses AA’s own controlled evaluation rather than relying on vendor-submitted results. The methodology applies consistently across models, so the relative positions reflect real differences in multi-step task performance.
The broader picture: Chinese open-weight and proprietary models now occupy multiple positions across the top tier of every major benchmark. Qwen3.8 Max at #1 on the AA agentic index, Kimi K3 at #4 on the overall intelligence index, GLM-5.2 competitive with GPT-5.5 on real-world agent tasks — these are not cherry-picked comparisons. China’s AI benchmark performance has converged with the US frontier on intelligence metrics. The remaining gap is in safety evaluation depth, independent third-party verification, and compute infrastructure. On raw benchmark numbers, the gap has closed to single digits on most standard evaluations.