GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Qwen3.8 Max Takes AA Agentic Index #1 at Intelligence 56 — but Costs $1.14 Per Task, Twice Its Predecessor

Alibaba’s Qwen3.8 Max has reached the top of Artificial Analysis’s agentic index, scoring 56 on the AA Intelligence Index — a position that puts it ahead of every US lab except Anthropic and OpenAI. It is the first Chinese model to claim the #1 position on AA’s agentic ranking.

The result has a catch: getting there cost more than twice as much per task.

The Benchmark Position

On Artificial Analysis’s agentic methodology, Qwen3.8 Max now holds the leading position at Intelligence Index 56. That places it above Kimi K3 (Moonshot AI, Intelligence Index rank #4 on AA overall), GLM-5.2, and the rest of the non-Anthropic/OpenAI field globally. Kimi K3 carries more total parameters at 2.8 trillion versus Qwen3.8 Max’s 2.4 trillion, and has a more complete independent benchmark record, including the AA front-end coding leaderboard.

On LiveBench, Qwen3.8 Max scores 78.5 overall with 64.6% on the agentic coding subcategory — results already reported when it launched. The AA agentic index result reflects a different evaluation methodology: AA’s framework measures multi-step task completion across its proprietary τ³-benchmark suite, which includes banking, legal, and research domains. The outcomes can diverge substantially from raw coding benchmarks.

The Cost Problem

Agentic index results measure the ceiling. Cost figures measure the floor. On the AA agentic framework, Qwen3.8 Max costs $1.14 per completed task. Its predecessor Qwen3.7 Max cost $0.53 per task on the same framework. Kimi K3 Max, which is competitive on intelligence metrics, costs $0.86 per task.

That cost structure is not an accident of the model architecture — it reflects how Qwen3.8 Max achieved the benchmark result. The model’s internal capability score (on Alibaba’s own multi-benchmark composite) rose from 0.474 to 0.725 generation-over-generation. But it appears to be taking more turns to complete multi-step agentic tasks than its predecessor, driving the per-task token count — and therefore the cost — up disproportionately.

The practical consequence: enterprise customers evaluating Qwen3.8 Max against Qwen3.7 Max or Kimi K3 for agentic workloads face a 2.1x to 1.3x cost premium for the extra performance. Whether that’s worth it depends entirely on what the incremental intelligence buys on the specific task.

The Competitive Context

The AA agentic index ranking is meaningful because it uses AA’s own controlled evaluation rather than relying on vendor-submitted results. The methodology applies consistently across models, so the relative positions reflect real differences in multi-step task performance.

The broader picture: Chinese open-weight and proprietary models now occupy multiple positions across the top tier of every major benchmark. Qwen3.8 Max at #1 on the AA agentic index, Kimi K3 at #4 on the overall intelligence index, GLM-5.2 competitive with GPT-5.5 on real-world agent tasks — these are not cherry-picked comparisons. China’s AI benchmark performance has converged with the US frontier on intelligence metrics. The remaining gap is in safety evaluation depth, independent third-party verification, and compute infrastructure. On raw benchmark numbers, the gap has closed to single digits on most standard evaluations.