GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Qwen3.7-Max-Preview Enters Arena at Rank 13 — Alibaba Reaches 6th Lab Globally

Alibaba’s Qwen3.7-Max-Preview entered the Arena text and Vision leaderboards on May 14. Results published May 19 show it at rank 13 globally on the text leaderboard, slotting between GPT-5.5 and Grok 4.2. It is the top-ranked Chinese model on the platform.

The entry lifts Alibaba to rank 6 among labs globally on Arena — the first time it has held that position.

Sub-Leaderboard Breakdown

The aggregated rank obscures where the model is genuinely competitive:

Sub-leaderboardRank
Mathematics7
Expert prompts9
Software and IT9
Coding10
Text (overall)13

Top-7 in math and sub-10 in expert tasks puts Qwen3.7-Max at a tier above its previous generation on the benchmarks that matter most for research and engineering workflows.

The Plus-Preview variant, which anchors the vision leaderboard, entered at rank 16, placing Alibaba at rank 5 globally in multimodal evaluation — also the top Chinese lab position.

The Acceleration Story

Qwen3.7’s arrival is as notable for its timing as its ranking. Alibaba has shipped three full model generations since February 2026:

  • Qwen3.5 — February 2026, open-weight 397B MoE, Apache 2.0
  • Qwen3.6 — April 2026, closed flagship (Max-Preview) ranked 3rd on Artificial Analysis at index score 52; open-weight 27B and 35B variants under Apache 2.0
  • Qwen3.7 — May 2026, preview on Arena, ranked 13th globally in text

By comparison, in all of 2025, Alibaba shipped two primary Qwen versions. The shift to monthly preview cadences — where community testing runs in parallel with stable-release preparation — mirrors the pace Anthropic adopted with Claude’s rapid 4.x succession.

Where the Gap Remains

Rank 13 is a genuine frontier entry, but the distance to the top is measurable. GPT-5.5 and Grok 4.2, which Qwen3.7 sandwiches, are themselves a tier below GPT-5.5-high, Claude Opus 4.7, and Gemini 3.1 Pro, which hold the top positions. On SWE-bench Verified, Qwen3.6-27B (77.2%) remains roughly 10 points behind Claude Opus 4.7 (87.6%). Closing that gap on agentic coding tasks is where the next generation wins or loses.

The stable release of Qwen3.7 has not been dated. Based on the Qwen3.6 preview-to-production cycle of roughly six weeks, a June window is plausible.