GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Sub-32B Open Weights Hit GPT-5 Intelligence: Qwen3.5 27B Scores 42, Gemma 4 31B Scores 39

The sub-32B weight class cleared a threshold this week. Alibaba’s Qwen3.5 27B (Reasoning) and Google DeepMind’s Gemma 4 31B (Reasoning) both land in GPT-5 territory on the Artificial Analysis Intelligence Index — 42 and 39 respectively, matching GPT-5 (medium) and GPT-5 (low).

That comparison comes with a sharp asterisk.

Where the Gap Closed

On agentic and reasoning tasks, the open models are genuinely competitive:

  • Qwen3.5 27B Agentic Index: 55 — ahead of GPT-5 (medium) at 46
  • Gemma 4 31B on TerminalBench Hard: 36% vs GPT-5 (low) at 27%
  • Gemma 4 31B on HLE: 23% vs GPT-5 (low) at 18%

Both families include reasoning and non-reasoning variants, multiple sizes, and native multimodal input — feature parity with proprietary mid-tier models.

Where the Gap Didn’t Close

Factual knowledge and hallucination avoidance remain frontier-gated. AA-Omniscience scores tell the story:

ModelAA-Omniscience
Qwen3.5 27B-42
Gemma 4 31B-45
GPT-5 (medium)-10
GPT-5 (low)-10

A 32-35 point gap on factual recall means these models code well and reason structurally, but still confabulate at rates that matter in production deployments that rely on accurate knowledge retrieval.

What It Means

Open-weights intelligence is no longer a lagging indicator of proprietary capability — it’s a concurrent track. Sub-32B models that would have been mid-tier one year ago now clear the GPT-5 floor on task performance benchmarks.

The practical consequence: for agentic workflows where factual precision is bounded (code generation, structured reasoning, tool use), Qwen3.5 27B and Gemma 4 31B are now viable cost-zero alternatives to mid-range API calls. Both are Apache 2.0 licensed and run on consumer hardware.

The remaining moat for frontier proprietary models is knowledge depth, not reasoning architecture.