Grok 4.5 Tops Perplexity Computer's WANDR at 0.328 — Beats Opus 4.8 at Half the Cost
Perplexity published head-to-head orchestrator data for its Computer agentic product this week, testing six model configurations on WANDR — a benchmark designed around large professional research tasks that require sustained search, computation, and multi-step reasoning. Grok 4.5 leads the field.
The Numbers
| Model | WANDR Score | Cost per Trial |
|---|---|---|
| Grok 4.5 | 0.328 | $4.76 |
| GPT-5.6 Sol (medium) | 0.289 | $2.64 |
| Claude Opus 4.8 (high, thinking) | 0.254 | $9.46 |
Grok 4.5 posted the highest score across all five competing configurations and at roughly half the cost of Opus 4.8. GPT-5.6 Sol in medium-effort mode came second — cheaper than Grok 4.5 but 39 millipoints below it on WANDR. Opus 4.8 with thinking enabled landed last on score while costing nearly twice Grok 4.5.
What Perplexity Computer and WANDR Actually Measure
Perplexity Computer is an agentic product that lets an orchestrator model direct subagents across research, coding, web browsing, and document-processing tasks. The orchestrator breaks down problems, delegates to specialists, and synthesises results.
WANDR — Wide-Area Network Data Research — evaluates exactly that coordination workload: tasks requiring many sources searched, computations run, results organised, duplicates removed, evidence checked, and a complete structured answer produced. It is not a single-pass generation benchmark. It penalises models that delegate poorly, lose track of intermediate results, or produce incomplete synthesis.
Why This Matters
The Opus 4.8 underperformance at $9.46/trial is the most significant finding. Raw intelligence benchmarks consistently place Opus 4.8 at or near the frontier. On WANDR under Perplexity Computer’s orchestration harness, it trails Grok 4.5 by 0.074 points — a 22.6% gap — at twice the price. The implication is that orchestration efficiency matters as a distinct capability from raw task intelligence, and that Grok 4.5 has meaningfully better orchestration discipline on sustained, wide-scope research tasks.
GPT-5.6 Sol at medium effort is the value alternative: cheaper than Grok 4.5 by $2.12/trial but scoring 0.039 lower. For cost-sensitive deployments with large research task volumes, GPT-5.6 Sol medium is the alternative to reach for; for maximum WANDR throughput, Grok 4.5 wins.
Grok 4.5 is now generally available as an orchestrator in Perplexity Computer for Pro and Max subscribers.