Artificial Analysis Benchmarks 7 Search APIs: Parallel Search Leads, Tavily Costs 9x More for Worse Results
Artificial Analysis released the AA Search Index on August 18, benchmarking 11 configurations across 7 search API providers on three tasks that measure how well search augments a base language model on real research problems. The verdict: search roughly doubles model accuracy, but the quality gap between providers is wide, and cost does not track quality.
The benchmark
The AA Search Index is the equal-weighted mean of three sub-benchmarks:
- DeepSearchQA: Deep, multi-hop factual research queries
- BrowseComp: Complex research that requires reading and synthesising web pages
- AA-Omniscience: Proprietary AA knowledge-intensive benchmark
All providers were tested using the same underlying model (GPT-5.6 Luna, medium) in the same agent harness. The baseline — the same model without any search — scores 33 overall.
Results
| Provider | Overall | DeepSearchQA | BrowseComp | AA-Omni | Search cost/task |
|---|---|---|---|---|---|
| Parallel Search (advanced) | 75 | 81 | 77 | 67 | $47.93 |
| Exa Search (auto) | 74 | 78 | 74 | 70 | $65.57 |
| Firecrawl Search | 73 | 74 | 74 | 73 | $30.48 |
| Parallel Search (basic) | 73 | 79 | 73 | 68 | $45.14 |
| Exa Search (fast) | 68 | 76 | 61 | 69 | $78.11 |
| You.com Search | 68 | 63 | 74 | 66 | $68.93 |
| Parallel Search (turbo) | 67 | 70 | 75 | 56 | $13.64 |
| Keenable Search (pro) | 67 | 70 | 66 | 65 | $23.69 |
| Keenable Search (realtime) | 67 | 70 | 64 | 66 | $26.87 |
| Tavily Search (basic) | 66 | 74 | 59 | 64 | $126 |
| Brave Search | 65 | 62 | 67 | 65 | $71.61 |
| Model only (baseline) | 33 | 45 | 17 | 38 | — |
What stands out
BrowseComp is where search matters most. The baseline model scores 17 on BrowseComp without any search access. The top providers push that to 75-77. This benchmark requires actually reading and synthesising retrieved pages — tasks where retrieval quality directly limits the model.
Tavily’s cost-quality ratio is severe. At $126 per task, Tavily is 9x the cost of Parallel Search (turbo) and still ranks 10th overall (66). It scores 74 on DeepSearchQA but falls apart on BrowseComp (59) and AA-Omniscience (64).
Firecrawl tops AA-Omniscience at 73, beating providers that rank above it overall. For knowledge-intensive research tasks specifically, Firecrawl’s web scraping approach outperforms the pure search-index approaches.
Budget option: Parallel Search (turbo) at $13.64. Scores 67 overall — 2 points below Exa (auto) at nearly 5x the cost. It leads all providers on BrowseComp (75). For cost-constrained agentic pipelines, turbo mode is a serious option.
Context
AA has been building out its agent-focused benchmark suite, alongside AA-Briefcase (knowledge work), Optima (custom benchmarks), and AgentPerf (hardware). The Search Index fills the gap for teams choosing search API providers for AI agent infrastructure.
Coverage launches with 7 providers; Artificial Analysis says it will expand.