Google Ships Gemini 3.5 Live Translate: Continuous Speech-to-Speech in 70+ Languages, SynthID-Watermarked
Google released Gemini 3.5 Live Translate on June 9, 2026 — a new audio model that translates live speech continuously rather than waiting for a speaker to finish a sentence before generating output.
The model supports 70+ languages and more than 2,000 language combinations. It auto-detects the source language without manual configuration. All AI-generated audio carries an imperceptible SynthID watermark, a first for Google’s translation products.
What’s Different From Existing Translation
Most real-time translation tools work in turns: the speaker finishes, the system processes, then the translation plays. Gemini 3.5 Live Translate operates on a streaming architecture — it processes audio as it arrives, balancing context quality against latency and staying a few seconds behind the original speaker without awkward pauses.
Google says the model preserves the speaker’s intonation, pacing, and pitch in the translated output. Independent demos posted since launch show it handling mid-sentence language switches and background noise without configuration changes.
The model’s input context window is 128K tokens (audio). Output supports audio and text up to 64K tokens.
Where It’s Live Today
| Surface | Availability |
|---|---|
| Google Translate (iOS/Android) | Generally available today |
| Gemini Live API + AI Studio | Public developer preview |
| Google Meet | Enterprise private preview (broader rollout later in 2026) |
Google Meet is the high-stakes deployment. Before today, Meet’s speech translation supported five languages and required English as one endpoint. With 3.5 Live Translate, that expands to 70+ languages and 2,000+ language pairs in a single meeting. The rollout starts with select business Google Workspace customers.
For Android users, a new “listening mode” streams translations through the phone’s earpiece — no headphones required. The use case is one-sided reception: following a presentation, guided tour, or support call in a language you don’t speak, without broadcasting the translation to the room.
Developer API
Developers can access the model via the Gemini Live API and Google AI Studio as of today. Documented use cases include live dubbing of video content, multilingual customer support routing, and real-time interpretation for broadcasts and virtual meetings.
Ride-hailing platform Grab is the named early enterprise tester, reportedly evaluating the model for driver-passenger communication across Southeast Asia’s 11 major national languages.
Why This Matters
Continuous speech translation has been a stated goal for Google for years. Prior attempts — including earlier Pixel Buds features — required offline phrase models with limited vocabulary. 3.5 Live Translate is a cloud-streamed frontier model applied to a consumer and enterprise interface for the first time at this language breadth.
The SynthID watermarking is notable as a policy lever: all translated audio is marked as AI-generated, which matters for legal proceedings, broadcast compliance, and any environment where provenance of speech matters. Google has not published the false-positive or detection rate for the watermark under adversarial conditions.
Pricing for the Gemini Live API access to 3.5 Live Translate has not been published separately — it will likely be bundled into Gemini API audio tiers.