Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two audio models designed to make conversations with artificial intelligence more natural, faster, and more useful. The proposal targets both users of its applications and companies and developers building voice agents.
The main difference lies in each model’s purpose: Gemini 3.8 Live prioritizes scale, cost, and fluency, while Gemini 3.8 Live Extended Thinking is designed for complex tasks that require multi-step reasoning. What’s the result? Systems that can talk with you while working in the background.
Two models for two types of conversation
Gemini 3.8 Live combines conversational intelligence, fluid dialogue, and near-real-time visual understanding. This means it can analyze images or elements appearing in front of the camera and use them as context during a conversation.
It can also automatically detect and switch between 97 supported languages in the middle of a dialogue. For someone who switches between languages or serves international customers, this capability can make the interaction feel less rigid and more like talking to another person.
The model can also run tools and make API calls while maintaining the conversation. For example, it can receive a request, confirm that it has started processing it, and continue talking with you while it looks up information or completes an action in the background.
Extended reasoning without breaking the conversation
Gemini 3.8 Live Extended Thinking is designed for workflows that require more analysis. Its distinctive feature is that it can reason and speak at the same time, rather than going completely silent while it works through a task.
The model uses early verbal cues such as “Let me check that...” and can describe the progress of a multi-step operation. This provides context and reduces one of the common frustrations with voice agents: not knowing whether the system is still working or has stopped responding.
Of course, speaking while reasoning does not mean every response will be instant. The advantage is that users receive confirmations and updates during the process, which is especially useful when the agent needs to consult external systems, run tools, or complete business tasks.
Results in voice and agent tasks
According to Google, Gemini 3.8 Live Extended Thinking took first place overall in Artificial Analysis’s Speech to Speech Quality Index, with a score of 82.6. This index evaluates the quality of a voice-to-voice interaction, including aspects such as naturalness, responsiveness, and conversational experience.
In agentic tasks, meaning those in which the model must take actions and complete goals, it achieved the following results:
- 68.6% on the τ-Voice benchmark.
- 35.1% on Sierra’s τ-Voice banking benchmark.
- 97.7% on Big Bench Audio, a test focused on audio reasoning capabilities.
Google also says the model maintains a competitive price compared with other advanced models. These figures were obtained through the Live API on the Gemini Enterprise Agent Platform, so they should be interpreted within that evaluation environment and not as an automatic guarantee for every application.
Gemini 3.8 Live, for its part, took second place in the Speech Agent Arena based on user preference. Its strength lies in offering a combination of capability and efficiency for large-scale deployments, where the cost per interaction and latency matter just as much as response quality.
What changes for developers
The new models are arriving on the Gemini Live API, which makes it possible to build real-time voice experiences. However, developing this type of agent is not just a matter of connecting a model and turning on a microphone.
You also have to manage audio streaming, interruptions, conversational turns, latency, tool calls, and stable connections. Platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents provide infrastructure to handle much of that complexity.
This allows teams to focus more on the user experience. For example, a company could create a technical support agent that listens to an explanation, checks the customer’s history, searches a knowledge base, and suggests a solution without making the person wait in silence.
Google also highlighted collaborations with Salesforce, Genspark, and Lumeris. These companies particularly emphasize latency, natural conversations, and the ability to run tools through function calls.
Why latency matters so much
In a text chatbot, waiting a few seconds may be tolerable. In a voice conversation, that same delay feels strange. Long pauses make the system seem confused or disconnected.
That is why voice agents need to respond quickly, detect when someone has finished speaking, and allow natural interruptions. The conversation should feel like an exchange, not a series of recorded messages.
Gemini 3.8 Live aims to address this need with near-real-time audio processing, while the Extended Thinking version adds deeper reasoning without completely leaving the conversational channel.
Availability in Google products
Gemini 3.8 Live began rolling out through the following channels:
- Developers: Gemini API and Google AI Studio.
- Businesses: private preview in Gemini Enterprise and soon in Gemini Enterprise for Customer Experience.
- General users: Search Live.
Gemini 3.8 Live Extended Thinking also began arriving on the Gemini API and Google AI Studio. For businesses, it is available in private preview in Gemini Enterprise and will soon arrive in Gemini Enterprise for Customer Experience and for Google Workspace business customers.
For end users, the Extended Thinking version is being added to Gemini Live. It will also be available to Google AI Pro and Ultra subscribers in Workspace within Docs, as well as arriving in Gmail and Keep for all Google AI subscribers.
Availability may vary depending on the country, account type, and stage of the rollout. In business products, a private preview usually means access is limited to specific organizations or programs.
SynthID adds a provenance signal
All audio generated by Google’s AI products includes a SynthID watermark. This signal is imperceptible to people and is integrated directly into the audio produced by the model.
Its goal is to help detect AI-generated content and reduce the risk of misinformation. It does not make audio true or false on its own, but it adds a layer of traceability that can be useful for platforms, researchers, and moderation teams.
The measure matters because voice agents can produce increasingly convincing conversations. As synthetic speech becomes more natural, knowing whether a clip was generated by AI stops being a technical curiosity and becomes a matter of trust.
Conversation becomes an interface
Gemini 3.8 Live shows where applied AI is heading: fewer windows filled with text and more systems capable of listening, observing, reasoning, and acting. The technology could be useful in customer service, education, internal support, productivity, and accessibility.
But the final experience will depend on more than benchmark scores. Privacy, tool accuracy, the ability to correct mistakes, and how clearly the system explains what it is doing will also be decisive.
The promise is easy to understand: talk to an AI without having to learn how to speak to it like a machine. The challenge will be ensuring that this naturalness comes with control, transparency, and reliable results.
