GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Arena Adds Gemini 3.7 Flash High and Grok 4.6 High Across Three Leaderboards in 48 Hours

Six models entered Arena evaluation across four days this week. Two of them — Google’s Gemini 3.7 Flash High and xAI’s Grok 4.6 High — represent the most cost-aggressive frontier-tier offerings currently available. Both now have live human preference voting data, which is the benchmark type automated evaluations still can’t replicate.

The August 11–13 Additions

August 11: NVIDIA’s Nemotron 3.5 Lightning 30B-A3B (NVfp4 quantisation) entered the text leaderboard. At 30B active parameters in a mixture-of-experts configuration, it’s the first NVIDIA-branded model to appear on Arena since the Nemotron series launched.

August 12: Grok 4.6 High joined Code Arena and Text Arena. The model carries 95.6% on SWE-bench Verified — matching Claude Fable 5 within half a point and making it one of two open-market models above the 95% threshold. Intelligence Index sits at 61, matching GPT-5.6 Sol, at $2 per million output tokens. That’s roughly 25x cheaper than Fable 5 on a per-token basis.

August 13: Gemini 3.7 Flash High entered Agent Arena, Text Arena, and Code Arena — three leaderboards on launch day, the same day Google published its release notes. The model posts 65.3% on DeepSWE, a 16-point improvement over Gemini 3.6 Flash. Pricing came in at roughly half of 3.6 Flash’s launch price.

Also on August 13: Black Forest Labs’ Flux 3 Video joined the Image-to-Video leaderboard, and MiniMax H3 — a 33B open-weight multimodal generation model — was added to the Video Edit leaderboard.

Why the Cadence Matters

Arena’s coverage question has always been whether it keeps pace with release velocity. This week it did: Gemini 3.7 Flash High entered Arena on the same calendar day it shipped, meaning developers have preference signal available alongside the initial benchmark data rather than weeks behind it.

The more interesting data will emerge from the Agent Arena results for Gemini 3.7 Flash High. Grok 4.6 High’s absence from that leaderboard — it entered Code and Text but not Agent — is a gap worth watching. The model’s 95.6% SWE-bench score suggests it belongs there; xAI may be holding back on the submission or still running evaluations.

The Cost-Intelligence Axis

Two data points define the current frontier:

  • Top-end SWE-bench (95%+): Grok 4.6 High and Claude Fable 5. Cost per million output: $2 vs. $50.
  • Flash-tier agentic coding (65%+ DeepSWE): Gemini 3.7 Flash High. Cost: substantially below $3/M output.

Human preference results from Arena won’t change these numbers, but they’ll answer the question automated benchmarks leave open: whether users experience the cost-competitive models as actually equivalent on the tasks that matter to them.

Both leaderboards are live. Grok 4.6 High and Gemini 3.7 Flash High enter the voting pool with existing baselines from Claude Opus 5, Fable 5, and GPT-5.6 Sol already established. The rank positions should begin to stabilise within a few days of voting volume.