Google introduces Gemini 3.5 Live Translate, an audio model designed to translate spoken conversations almost in real time. The tool detects more than 70 languages and generates translated speech that aims to preserve the speaker’s intonation, rhythm, and tone.
The difference is not just in how many languages it recognizes. It is also in how it responds: instead of waiting for someone to finish each sentence, the system processes audio as it arrives and produces translation continuously. What’s the result? Fewer awkward pauses and a delay of just a few seconds during the conversation.
Voice translation without waiting for each turn
Many translation systems work in turns. They listen to a complete sentence, process it, and then deliver the result. That approach can be accurate, but it disrupts the rhythm of a natural conversation.
Gemini 3.5 Live Translate attempts to solve that problem through streaming audio processing. The model receives speech progressively and must balance two goals that often conflict: waiting for enough context to improve quality or responding quickly enough to stay synchronized with the speaker.
It can also work with multilingual input without requiring users to manually configure each language. Google also highlights its resistance to noise, an important feature for calls, meetings, classes, broadcasts, and conversations in unpredictable environments.
Translation stops feeling like a series of separate responses and comes closer to a continuous conversation.
Where Gemini 3.5 Live Translate will be available
The launch reaches different products and audiences:
- Developers: public preview access through the
Gemini Live APIand Google AI Studio. - Businesses: private preview in Google Meet for selected Google Workspace customers.
- General users: global rollout in Google Translate for Android and iOS.
For developers, the API makes it possible to create live interpreting, dubbing, and simultaneous translation features in multiple languages. Platforms such as Agora, Fishjam, LiveKit, Pipecat, and Vision Agents already integrate multimedia streaming infrastructure, allowing teams to focus on designing the final experience.
This opens up possibilities for multilingual calls, international meetings, courses, events, and broadcasts. Grab, for example, is testing the model to make communication between drivers and passengers easier during pickup. The company processes more than 10 million voice calls per month among its users.
Google Meet expands its translation languages
Google is also preparing an update to voice translation in Meet. The feature will move from a previous limit of five languages to more than 70, and it will support more than 2,000 language combinations in a single meeting.
Another change will be the ability to translate between languages without relying exclusively on English as an intermediate language. The interface will also be updated to provide more direct access to voice translation.
The feature begins in private preview for some business customers during this month and will receive a broader rollout later in 2026.
Translation from your phone and headphones
In the Google Translate app for Android and iOS, users will be able to activate Live Translate and connect any pair of headphones. The translation will be played back while attempting to reflect the tone of the person speaking in more than 70 languages.
On Android, a listening mode is also beginning to roll out. You can hold the phone next to your ear, as if you were answering a call, to hear the translation directly through the device’s earpiece.
This mode can be useful when you do not have headphones available or when you want to hear a translation without playing it aloud for people nearby. Google gives the example of a guided tour in Spanish being translated into English and delivered directly to the user’s ear.
A voice model with a SynthID watermark
Audio generated by Gemini 3.5 Live Translate includes an imperceptible SynthID watermark. This signal is embedded directly into the audio output and makes it possible to detect that it was created by artificial intelligence.
The measure aims to make synthetic content easier to identify and reduce risks related to misinformation. Google says that the technical and safety details are available in the model card.
Real-time translation will still face challenges: accents, noise, double meanings, and cultural expressions do not disappear just because you use a faster model. However, Gemini 3.5 Live Translate shows where this technology is headed: less time waiting for a response and more tools for having conversations without language being an immediate barrier.
Original source
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate
