Qwen3.7 Plus Hits ScreenSpot Pro 79.0: Alibaba's GUI+CLI Hybrid Agent Enters Frontier-Tier Computer Use
Alibaba quietly shipped Qwen3.7 Plus on June 1, eleven days after the text-only Qwen3.7 Max. The model adds native vision input on top of Max’s language backbone and is built specifically for the intersection that defines modern agent work: see a UI, operate it; read a terminal, act on it.
Key Numbers
- ScreenSpot Pro: 79.0 (GUI grounding, frontier-tier threshold: 75+)
- Terminal-Bench 2.0: 70.3 (agentic coding in sandboxed shell)
- Vision Arena: #16 overall at launch
- Pricing via OpenRouter: $0.40/M input, $1.60/M output (≤256K tokens); $1.20/M input, $4.80/M output (>256K)
- Context: 1M tokens
- Autonomous run ceiling: 35 hours, 1,000+ sequential tool calls
The Benchmark That Matters
ScreenSpot Pro is the bottleneck metric for GUI automation in 2026. The test: feed the model a screenshot of real software — an admin dashboard, Photoshop, an enterprise SaaS UI — along with a natural-language instruction, and score whether it can identify the exact pixel coordinates of the element to interact with. Claude Computer Use, OpenAI Operator, and every browser automation stack live or die on this capability. State-of-the-art sits in the 75–82 range. Qwen3.7 Plus at 79.0 places Alibaba inside that band.
For comparison: Qwen3.5-27B scored 70.3 on ScreenSpot Pro at its release. Plus adds 8.7 points, and does it while carrying full agentic coding capability in the same model.
The Hybrid Claim
What makes Plus distinct from earlier Qwen multimodal work is that the vision capability doesn’t trade away language depth. Terminal-Bench 70.3 puts Plus marginally ahead of the text-only Qwen3.7 Max (69.7 on Terminal-Bench 2.0-Terminus) on coding tasks. Most multimodal models sacrifice language performance for visual parameters; Plus appears to hold both.
The architecture integrates GUI and CLI surfaces inside a single agentic loop. A task that requires navigating a web interface to find a configuration value and then running a shell command to apply it can run without model switching or external orchestration — the perception and execution both live in Plus.
Pricing Context
The OpenRouter price of $0.40/M input places Qwen3.7 Plus significantly below both Claude Opus 4.8 ($15/M input) and GPT-5.5 ($5/M input) for comparable computer-use capability. At $1.60/M output it undercuts Gemini 3.1 Pro ($2/M input, $12/M output) substantially.
The practical catch: Plus carries its multimodal stack on every request. Independent testing showed roughly 14% latency overhead versus Max on text-only tasks, and 7% higher token cost on agentic workflows that never invoke vision. For workloads where visual inputs appear less than 20% of the time, Max remains the leaner option.
What This Does to the Computer Use Market
Three weeks ago, the frontier-tier GUI grounding options were Claude Computer Use and OpenAI Operator — both expensive, both from US labs. Qwen3.7 Plus at ScreenSpot Pro 79.0 and $0.40/M changes the calculus for teams building browser automation, screenshot-to-code pipelines, and GUI testing infrastructure. The moat that Western labs held on computer use capability has a credible challenger.
The model doesn’t generate images or video. It reads them. The generation side stays in separate Alibaba model families. For pure computer use and visual agent work, Plus is the most cost-effective frontier option in the market as of this week.