Google Gemini 3.8 Live Claims #1 Speech-to-Speech Score at 82.6, Priced at $0.005/min
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, available via the Live API in the Gemini API and Google AI Studio.
3.8 Live Extended Thinking claims the top position on Artificial Analysis’ Speech-to-Speech Quality Index with a score of 82.6. On agentic voice benchmarks, it scores 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.
What the models do
Gemini 3.8 Live is the scale-and-cost variant: it handles fluid dialogue, background API calls, and visual grounding simultaneously. The model can perform tasks — web lookups, function calls — while maintaining uninterrupted conversation.
3.8 Live Extended Thinking adds a reasoning layer for complex requests without breaking dialogue flow. Google positions it for enterprise task completion where a standard voice model would lose context across long exchanges.
Both models detect and switch between 97 languages mid-conversation.
Benchmarks
| Benchmark | Score |
|---|---|
| AA Speech-to-Speech Quality Index | 82.6 (ranked #1) |
| τ-Voice | 68.6% |
| Sierra τ-Voice-banking | 35.1% |
For reference, Gemini 3.5 Transcribe — the dedicated speech-to-text model released last month — achieves 4.0% Word Error Rate in streaming mode and 2.6% in non-streaming, across 85+ languages.
Pricing
Audio input: $0.005 per minute. Audio output: $0.018 per minute. Google notes this as equivalent to approximately $3/1M input tokens and $12/1M output tokens.
OpenAI’s GPT Live-1, which launched September 10, is priced at $0.05 per session minute, a different billing unit that makes direct per-minute comparison difficult, and includes backend reasoning token charges on top. For sustained audio workloads billed by audio duration, Gemini’s per-minute rate is substantially lower.
Partner integrations
The Live API is available through Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents — all handle media streaming infrastructure so developers deploy without writing their own WebRTC or WebSocket handling.
Context
Google’s prior speech model, Gemini Live 3.1, was flagged in May’s Full-Duplex Bench v3 for going silent 22% of the time — a reliability issue that hurt its enterprise case. 3.8 Live addresses that with the background-task architecture, though independent benchmark validation of silence rates is not yet available.
The τ-Voice-banking score at 35.1% reflects the model’s performance on banking-specific scripted conversations, a benchmark Sierra introduced to measure real-world finance use cases. The 68.6% τ-Voice general score is the primary agentic comparison point.