GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Qwen 3.8 Max Leads LiveBench Agentic Coding at 64.6 — Beats Fable 5 Max Effort at $0.28 Per Task

Qwen 3.8 Max from Alibaba’s Tongyi Lab has entered LiveBench’s global top 10 with a result that will draw attention: 64.6 on Agentic Coding, the highest score in that category outside the Max Effort tier of the full frontier runs.

The Numbers

LiveBench measures seven capability dimensions. Qwen 3.8 Max scores:

CategoryScore
Overall78.5
Reasoning88.2
Coding72.9
Agentic Coding64.6
Mathematics91.3
Data Analysis78.4
Language79.7
Instruction Following74.1
Cost per successful task$0.275

The agentic coding score is the headline. At 64.6, Qwen 3.8 Max beats Claude Fable 5 Max Effort (62.2), Kimi K3 (62.2), and GPT-5.5 Thinking xHigh Effort (54.0) on this single dimension. The only scores above it are GPT-5.6 Terra (68.0) and GPT-5.6 Sol Max Effort (56.2 — notably lower, reflecting Sol’s known weakness on agentic tasks versus absolute intelligence).

What Agentic Coding Measures

LiveBench’s Agentic Coding category tests multi-step software engineering work: tasks that require planning, tool use, state management, and iterative correction across a codebase. It is harder to saturate than standard coding benchmarks because partial credit is limited — either the agent ships working code or it doesn’t.

A score of 64.6 on this dimension, from a model costing $0.275 per successful task, is the kind of cost-efficiency result that changes how teams build routing layers.

Positioning

Qwen 3.8 Max is also now on Arena’s Image-to-WebDev leaderboard as of August 4. Pricing details beyond the $0.275 per LiveBench task are confirmed at competitive mid-tier rates.

Overall at 78.5, Qwen 3.8 Max is just below Kimi K3 (79.2) and significantly below the full Max Effort frontier runs (Claude Fable 5 Max Effort at 83.0, GPT-5.6 Sol Max Effort at 81.0). But overall rankings compress models that are strong across all categories. For agentic coding specifically, Qwen 3.8 Max’s position at the top of the non-Max Effort tier makes it relevant for production routing decisions where cost per successful agentic task is the operative metric.

Mathematics at 91.3 is notably strong and suggests the model’s reasoning backbone remains competitive with the frontier tier on structured problem-solving.

The Qwen Trajectory

This follows Qwen3.6-Plus and Qwen3.7-Max, which each held top positions on the AA Intelligence Index for their respective tiers. Qwen 3.8 Max continues the pattern: Alibaba releases a model that doesn’t top the overall leaderboard but dominates in the category most relevant to agentic workloads.

External SWE-bench Verified data for Qwen 3.8 Max has not yet been published independently.