GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Grok 4.5 Tops Perplexity Computer's WANDR at 0.328 — Beats Opus 4.8 at Half the Cost

Perplexity published head-to-head orchestrator data for its Computer agentic product this week, testing six model configurations on WANDR — a benchmark designed around large professional research tasks that require sustained search, computation, and multi-step reasoning. Grok 4.5 leads the field.

The Numbers

ModelWANDR ScoreCost per Trial
Grok 4.50.328$4.76
GPT-5.6 Sol (medium)0.289$2.64
Claude Opus 4.8 (high, thinking)0.254$9.46

Grok 4.5 posted the highest score across all five competing configurations and at roughly half the cost of Opus 4.8. GPT-5.6 Sol in medium-effort mode came second — cheaper than Grok 4.5 but 39 millipoints below it on WANDR. Opus 4.8 with thinking enabled landed last on score while costing nearly twice Grok 4.5.

What Perplexity Computer and WANDR Actually Measure

Perplexity Computer is an agentic product that lets an orchestrator model direct subagents across research, coding, web browsing, and document-processing tasks. The orchestrator breaks down problems, delegates to specialists, and synthesises results.

WANDR — Wide-Area Network Data Research — evaluates exactly that coordination workload: tasks requiring many sources searched, computations run, results organised, duplicates removed, evidence checked, and a complete structured answer produced. It is not a single-pass generation benchmark. It penalises models that delegate poorly, lose track of intermediate results, or produce incomplete synthesis.

Why This Matters

The Opus 4.8 underperformance at $9.46/trial is the most significant finding. Raw intelligence benchmarks consistently place Opus 4.8 at or near the frontier. On WANDR under Perplexity Computer’s orchestration harness, it trails Grok 4.5 by 0.074 points — a 22.6% gap — at twice the price. The implication is that orchestration efficiency matters as a distinct capability from raw task intelligence, and that Grok 4.5 has meaningfully better orchestration discipline on sustained, wide-scope research tasks.

GPT-5.6 Sol at medium effort is the value alternative: cheaper than Grok 4.5 by $2.12/trial but scoring 0.039 lower. For cost-sensitive deployments with large research task volumes, GPT-5.6 Sol medium is the alternative to reach for; for maximum WANDR throughput, Grok 4.5 wins.

Grok 4.5 is now generally available as an orchestrator in Perplexity Computer for Pro and Max subscribers.