GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Qwen3.8 Max Ships: 2.4T Parameters, 86.6% Terminal-Bench, #4 Frontend Code Arena on Day One

Alibaba released Qwen3.8-Max on August 3, 2026. It is the largest model in the Qwen series: 2.4 trillion total parameters, 95 billion active per forward pass, multimodal input covering text, image, and video, and a 1 million token context window. The model is live on Alibaba Cloud Model Studio. Open weights are scheduled for the following week.

Artificial Analysis placed it 16th of 186 models at an Intelligence Index score of 53 — squarely in the upper tier, five points below Kimi K3 and ten below the Anthropic-led frontier. Pricing is $2.00 per million input tokens, $6.00 per million output, with cache hits dropping to $0.25.

Benchmark Breakdown

Alibaba published a benchmark table against Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol. The results are vendor-run; independent evaluations had not appeared by launch day.

BenchmarkQwen3.8 MaxOpus 4.8Fable 5GPT-5.6 Sol
Terminal-Bench 2.186.6%84.6%84.6%88.8%
SWE-bench Pro67.7%69.2%80.0%64.6%
PaperBench93.0%80.3%88.8%90.5%
GPQA Diamond92.6%92.0%92.6%94.1%
HLE43.6%45.7%53.3%47.2%
IFBench82.8%62.2%63.5%72.7%

The Terminal-Bench lead over Fable 5 and Opus 4.8 is the headline number from Alibaba’s table, but it comes with a caveat: GPT-5.6 Sol still leads at 88.8%. On the harder software-engineering benchmark, SWE-bench Pro, Qwen3.8 Max posts 67.7% — above GPT-5.6 Sol, behind both Anthropic models. PaperBench is a clear win: 93.0% against Fable 5’s 88.8% and GPT-5.6 Sol’s 90.5%. Instruction-following benchmark IFBench shows the largest gap, with Qwen3.8 Max at 82.8% against a field clustered in the 60s.

The weaker numbers are on Humanity’s Last Exam (43.6%, last among the four compared) and deepSWE (56.6%, 16 points behind GPT-5.6 Sol’s 72.7%). Broad knowledge and the hardest agentic tasks remain more expensive to acquire.

Arena Results

On Arena.ai leaderboards:

  • Frontend Code Arena: #4 at 1,668 ELO, behind Claude Opus 5 max effort (1,705), Kimi K3 max effort (1,676), and Claude Opus 5 high effort (1,669). Ahead of Claude Fable 5 high effort (1,630), GPT-5.6 Sol highest (1,620), and GLM-5.2 (1,586).
  • Vision Arena: #2 at 1,305, behind Claude Fable 5 high effort (1,318). Ahead of every Claude Opus variant, Gemini 3 Pro, GPT-5.5, and Grok 4.5.
  • Text Arena: #5.

Alibaba reports the model ranks consistently across Frontend Code category breakdowns — second in Consumer Product, third in Brand & Marketing, Reference-based design, Gaming, and Content Creation Tools — rather than winning one narrow area and dragging down on the rest.

Architecture and Long-Horizon Claims

2.4T total parameters at 95B active is a large activation ratio for a MoE model: roughly 4% of parameters fire per token. Qwen reports a 1M token context window with multimodal inputs.

The release post introduces RecreationBench, a long-horizon benchmark where the model reconstructs running applications from scratch in a black-box environment — no source code, no internet, only interaction and visual feedback. Alibaba also cites a 16-day autonomous software engineering run in internal testing, during which the model produced an open-sourced self-evolving agent framework called “oh-my-cli.”

These results are self-reported and have not been reproduced by third parties.

What to Watch

Independent evaluations on SWE-bench Verified and vals.ai should land within days. That number will reveal where Qwen3.8 Max actually sits in the agentic capability stack alongside Kimi K3 (93.4% Verified) and GPT-5.6 Sol (96.2% Verified). The open-weight release, expected this week, will determine whether the hardware demands match the parameter count, how the weights quantize, and whether community deployments confirm Alibaba’s benchmark numbers.

At $2/$6 per million with a 1M context window and multimodal inputs, Qwen3.8 Max enters at a price point below Kimi K3 ($3/$15) and well below Fable 5.