GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Qwen3.7-Max Hit Intelligence Index #5 Globally by Answering Fewer Questions — Abstention Is Now a Frontier Strategy

Alibaba’s Qwen3.7-Max scores 56.6 on the Artificial Analysis Intelligence Index v4.0, placing fifth globally and becoming the highest-ranked Chinese model on the leaderboard. It leads Gemini 3.5 Flash (55.3) and sits 0.7 points behind Claude Opus 4.7 (57.3) and Gemini 3.1 Pro Preview (57.2). GPT-5.5 leads the Index at 60.2.

The number is real. But how it was achieved is worth understanding.

What Actually Improved

Most of the 4.8-point gain over predecessor Qwen3.6 Max Preview (51.8) is concentrated in three areas:

  • CritPt: +9.7 percentage points (3.7% to 13.4%)
  • Humanity’s Last Exam: +9.2 points (28.9% to 38.1%)
  • Terminal-Bench Hard: +6.9 points (43.9% to 50.8%)
  • GDPval-AA: +42 Elo points (1504 to 1546)

These are genuine capability improvements. HLE gains at this magnitude are not trivial. Terminal-Bench Hard improvement is consistent with the model’s documented 35-hour autonomous kernel optimization results.

The Abstention Problem

AA-Omniscience is the benchmark that makes this complicated.

On that evaluation, Qwen3.7-Max’s raw accuracy fell 7.6 percentage points — from 37.7% to 30.1%. That is a worse outcome on the underlying task. At the same time, its hallucination rate fell 21.3 points, from 44.2% to 22.9%.

The mechanism: the model is now refusing to answer more questions. Its attempt rate dropped from 67.3% to 48.0%, the lowest among all frontier models tested. More than half of AA-Omniscience questions got no answer.

AA-Omniscience’s scoring structure rewards correct answers and penalises hallucinations. It has no penalty for abstaining. A model that says “I don’t know” to 52% of questions will post a strong hallucination number without retrieving more facts. Qwen3.7-Max currently holds the lowest hallucination rate at the frontier — but the rate reflects a participation strategy as much as a knowledge improvement.

What This Means for Index Rankings

The AA Intelligence Index v4.0 aggregates ten evaluations including GDPval-AA, Terminal-Bench Hard, SciCode, AA-Omniscience, HLE, and GPQA Diamond. Abstention-driven improvement on one benchmark component propagates into the composite. The model’s overall Index position is legitimate under the current ruleset, but the AA-Omniscience component overstates knowledge quality.

Frontier lab model teams read these benchmark structures carefully. Qwen3.7-Max appears to be the first model to implement abstention-tuning at a measurable scale as a deliberate strategy. If the approach holds, expect others to follow — until benchmark operators add a refusal penalty.

Other Specifications

  • Context window: 1 million tokens (up from 256K on Qwen3.6 Max Preview)
  • Input/output format: text only
  • Pricing: not yet announced (Qwen3.6 Max Preview was priced at $1.30/$7.80 per million on Alibaba Cloud)
  • Weights: closed
  • Arena ELO: 1,475 (#13 overall, #7 mathematics, #10 coding)

The leading open-weights Qwen on the Intelligence Index remains Qwen3.6 27B Reasoning at 45.8.