GLM-52 882 -0.2%
GPT-56T 861
GROK-46H 852 -0.6%
QWEN-38X 843 +0.4%
CL-OP5X 840 -0.2%
MUSE-SPK 835 +0.7%
GPT-6A 820
CL-OP5H 806 -0.1%
GPT-56SC 791 +0.5%
GLM-5 784 +0.4%
CL-FAB5H 774 -0.5%
GEM-37FH 745 -0.5%
KIMI-K3X 743 +0.1%
CL-OP46H 729 -5.7%
CL-OP47H 720 -6.6%
GPT-56S 709 -0.3%
GEM-38FH 677 -4.5%
GPT-55H 612 -0.2%
CL-OP47 602 -8.8%
INKL 531
GEM-31P 523 +0.8%
GEM-3P 501
CL-OP46 499 -0.2%
CL-OP48 493 -20.2%
GLM-52 882 -0.2%
GPT-56T 861
GROK-46H 852 -0.6%
QWEN-38X 843 +0.4%
CL-OP5X 840 -0.2%
MUSE-SPK 835 +0.7%
GPT-6A 820
CL-OP5H 806 -0.1%
GPT-56SC 791 +0.5%
GLM-5 784 +0.4%
CL-FAB5H 774 -0.5%
GEM-37FH 745 -0.5%
KIMI-K3X 743 +0.1%
CL-OP46H 729 -5.7%
CL-OP47H 720 -6.6%
GPT-56S 709 -0.3%
GEM-38FH 677 -4.5%
GPT-55H 612 -0.2%
CL-OP47 602 -8.8%
INKL 531
GEM-31P 523 +0.8%
GEM-3P 501
CL-OP46 499 -0.2%
CL-OP48 493 -20.2%
← Back to feed

Google Gemini 3.8 Live Claims #1 Speech-to-Speech Score at 82.6, Priced at $0.005/min

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, available via the Live API in the Gemini API and Google AI Studio.

3.8 Live Extended Thinking claims the top position on Artificial Analysis’ Speech-to-Speech Quality Index with a score of 82.6. On agentic voice benchmarks, it scores 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.

What the models do

Gemini 3.8 Live is the scale-and-cost variant: it handles fluid dialogue, background API calls, and visual grounding simultaneously. The model can perform tasks — web lookups, function calls — while maintaining uninterrupted conversation.

3.8 Live Extended Thinking adds a reasoning layer for complex requests without breaking dialogue flow. Google positions it for enterprise task completion where a standard voice model would lose context across long exchanges.

Both models detect and switch between 97 languages mid-conversation.

Benchmarks

BenchmarkScore
AA Speech-to-Speech Quality Index82.6 (ranked #1)
τ-Voice68.6%
Sierra τ-Voice-banking35.1%

For reference, Gemini 3.5 Transcribe — the dedicated speech-to-text model released last month — achieves 4.0% Word Error Rate in streaming mode and 2.6% in non-streaming, across 85+ languages.

Pricing

Audio input: $0.005 per minute. Audio output: $0.018 per minute. Google notes this as equivalent to approximately $3/1M input tokens and $12/1M output tokens.

OpenAI’s GPT Live-1, which launched September 10, is priced at $0.05 per session minute, a different billing unit that makes direct per-minute comparison difficult, and includes backend reasoning token charges on top. For sustained audio workloads billed by audio duration, Gemini’s per-minute rate is substantially lower.

Partner integrations

The Live API is available through Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents — all handle media streaming infrastructure so developers deploy without writing their own WebRTC or WebSocket handling.

Context

Google’s prior speech model, Gemini Live 3.1, was flagged in May’s Full-Duplex Bench v3 for going silent 22% of the time — a reliability issue that hurt its enterprise case. 3.8 Live addresses that with the background-task architecture, though independent benchmark validation of silence rates is not yet available.

The τ-Voice-banking score at 35.1% reflects the model’s performance on banking-specific scripted conversations, a benchmark Sierra introduced to measure real-world finance use cases. The 68.6% τ-Voice general score is the primary agentic comparison point.