GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Nari Labs Tops Coval Voice AI Benchmarks with 44ms STT at $0.12/hr — 3.75x Cheaper Than AssemblyAI

Nari Labs published benchmark results on September 14, 2026, placing its Qwen3-ASR and Qwen3-TTS models at the top of Coval’s public voice AI leaderboard on the latency-cost Pareto frontier.

Speech-to-Text

Qwen3-ASR Fast ranked first on time-to-final-segment (TTFS) at 44ms p50 — the latency between a user’s final utterance and the text output. Word error rate came in at 3.6%, placing it second behind AssemblyAI’s Universal 3.5 Pro at 3.5%.

On pricing, Nari Labs charges $0.12 per hour, which ties for second-lowest among models with publicly listed rates on Coval’s pricing directory. AssemblyAI Universal 3.5 Pro runs $0.45/hr — 3.75x the price. Deepgram Nova 3 costs $0.29/hr — 2.4x.

A Standard endpoint at $0.06/hr would be the cheapest available if accuracy requirements can accommodate the trade-off.

Text-to-Speech

Qwen3-TTS Fast ranked second on time-to-first-audio (TTFA) at 63ms p50, and first on word error rate among TTS models in the benchmark.

Why Coval

Coval measures production-relevant latency — not processing time, but the wall-clock gap that determines whether a voice AI agent feels responsive. TTFA and TTFS are the metrics that matter in live conversation, which is why the benchmark is widely cited in voice agent deployment decisions.

Nari Labs notes Coval’s rankings refresh every 30 minutes, so the exact positions shift. The company published these results as of mid-September 2026, representing a snapshot rather than a permanent ranking.

Context

The voice AI benchmark market has become competitive. PolyAI’s Raven 3.5 claimed top marks on customer service benchmarks in May 2026. Soniox TTS v2 targeted production agents in August 2026 with frontier quality claims. Artificial Analysis launched a speech-to-speech index in June 2026 showing GPT-Realtime-2 leading at 77.2%.

Nari Labs differentiates on the cost-latency combination specifically — its models sit at the intersection of sub-50ms response times and sub-$0.15/hr pricing, an edge that matters for high-volume deployments where inference cost compounds.

Key Numbers

  • STT latency: 44ms p50 TTFS (#1 on Coval)
  • STT word error rate: 3.6% (#2, behind AssemblyAI Universal 3.5 Pro at 3.5%)
  • STT price: $0.12/hr Fast endpoint ($0.06/hr Standard)
  • TTS latency: 63ms p50 TTFA (#2 on Coval)
  • TTS word error rate: #1 on Coval
  • Benchmark: Coval Voice AI Leaderboard, September 2026