GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Qwen3.7 Plus Hits ScreenSpot Pro 79.0: Alibaba's GUI+CLI Hybrid Agent Enters Frontier-Tier Computer Use

Alibaba quietly shipped Qwen3.7 Plus on June 1, eleven days after the text-only Qwen3.7 Max. The model adds native vision input on top of Max’s language backbone and is built specifically for the intersection that defines modern agent work: see a UI, operate it; read a terminal, act on it.

Key Numbers

  • ScreenSpot Pro: 79.0 (GUI grounding, frontier-tier threshold: 75+)
  • Terminal-Bench 2.0: 70.3 (agentic coding in sandboxed shell)
  • Vision Arena: #16 overall at launch
  • Pricing via OpenRouter: $0.40/M input, $1.60/M output (≤256K tokens); $1.20/M input, $4.80/M output (>256K)
  • Context: 1M tokens
  • Autonomous run ceiling: 35 hours, 1,000+ sequential tool calls

The Benchmark That Matters

ScreenSpot Pro is the bottleneck metric for GUI automation in 2026. The test: feed the model a screenshot of real software — an admin dashboard, Photoshop, an enterprise SaaS UI — along with a natural-language instruction, and score whether it can identify the exact pixel coordinates of the element to interact with. Claude Computer Use, OpenAI Operator, and every browser automation stack live or die on this capability. State-of-the-art sits in the 75–82 range. Qwen3.7 Plus at 79.0 places Alibaba inside that band.

For comparison: Qwen3.5-27B scored 70.3 on ScreenSpot Pro at its release. Plus adds 8.7 points, and does it while carrying full agentic coding capability in the same model.

The Hybrid Claim

What makes Plus distinct from earlier Qwen multimodal work is that the vision capability doesn’t trade away language depth. Terminal-Bench 70.3 puts Plus marginally ahead of the text-only Qwen3.7 Max (69.7 on Terminal-Bench 2.0-Terminus) on coding tasks. Most multimodal models sacrifice language performance for visual parameters; Plus appears to hold both.

The architecture integrates GUI and CLI surfaces inside a single agentic loop. A task that requires navigating a web interface to find a configuration value and then running a shell command to apply it can run without model switching or external orchestration — the perception and execution both live in Plus.

Pricing Context

The OpenRouter price of $0.40/M input places Qwen3.7 Plus significantly below both Claude Opus 4.8 ($15/M input) and GPT-5.5 ($5/M input) for comparable computer-use capability. At $1.60/M output it undercuts Gemini 3.1 Pro ($2/M input, $12/M output) substantially.

The practical catch: Plus carries its multimodal stack on every request. Independent testing showed roughly 14% latency overhead versus Max on text-only tasks, and 7% higher token cost on agentic workflows that never invoke vision. For workloads where visual inputs appear less than 20% of the time, Max remains the leaner option.

What This Does to the Computer Use Market

Three weeks ago, the frontier-tier GUI grounding options were Claude Computer Use and OpenAI Operator — both expensive, both from US labs. Qwen3.7 Plus at ScreenSpot Pro 79.0 and $0.40/M changes the calculus for teams building browser automation, screenshot-to-code pipelines, and GUI testing infrastructure. The moat that Western labs held on computer use capability has a credible challenger.

The model doesn’t generate images or video. It reads them. The generation side stays in separate Alibaba model families. For pure computer use and visual agent work, Plus is the most cost-effective frontier option in the market as of this week.