Google introduced Gemini 3.1 Flash Live, its new real-time audio model designed for more natural, faster, and more reliable conversations. The update is aimed at developers building voice agents, as well as companies and users who interact with Gemini and Search.
What changes when an AI understands a spoken conversation better? It’s not just about responding with less delay. It also needs to follow the thread, interpret pauses, recognize frustration, and carry out multi-step tasks without getting lost along the way.
A voice model designed for real conversations
Gemini 3.1 Flash Live improves Gemini’s ability to maintain fluid dialogue. According to Google, the model responds faster than its previous version and can preserve the context of a conversation for twice as long in Gemini Live.
This is especially useful during long brainstorming sessions, technical explanations, or problem-solving. Instead of forcing you to repeat every detail, the model can maintain a broader conversational thread and make use of what you’ve already said.
The improvement also applies to the way it speaks. Gemini 3.1 Flash Live analyzes acoustic signals such as the tone and rhythm of your voice. This allows it to detect, for example, that someone is confused or frustrated and adjust its response to be clearer and more helpful.
The naturalness of a voice AI doesn’t depend only on pronouncing words correctly. It also depends on knowing when to stop, how to react, and which information to keep in context.
More capacity for carrying out complex tasks
For developers, one of the key new features is improved tool use through function calling. This capability allows the model not only to hold a conversation, but also to activate external functions, query systems, or complete processes by following several instructions.
In ComplexFuncBench Audio, a benchmark that evaluates function calls involving multiple steps and constraints, Gemini 3.1 Flash Live achieved a score of 90.8%. Google says this result surpasses that of its previous model.
In Audio MultiChallenge, a Scale AI test focused on following complex instructions and reasoning over extended periods, it reached 36.1% with reasoning mode enabled. The benchmark includes interruptions, pauses, and hesitations similar to those that occur in real conversations.
These metrics don’t mean the model is perfect. They do show where voice AI is heading: agents capable of listening to a request, breaking it into steps, and completing a task without relying on rigid commands.
Where developers can use it
Gemini 3.1 Flash Live is available to developers in preview through the Gemini Live API, within Google AI Studio. With this API, teams can experiment with real-time voice applications and agents capable of interacting more flexibly.
Potential uses include:
- Technical support assistants that ask questions and guide users step by step.
- Customer service agents that detect signs of confusion or frustration.
- Educational tools that adapt their explanations to each person’s pace.
- Applications that combine voice, text, images, and external actions.
Google also mentions positive feedback from companies such as Verizon, LiveKit, and The Home Depot, which have tested the model in their workflows.
Availability for companies and users
In the enterprise environment, Gemini 3.1 Flash Live is coming to Gemini Enterprise for Customer Experience. The goal is to help companies build more conversational support experiences that can recognize vocal nuances and respond dynamically.
For the general public, the model is available in Gemini Live and Search Live. In the latter case, Google announced a global expansion that allows users to have real-time multimodal conversations with Search in their preferred language, across more than 200 countries and territories.
This can be useful for resolving everyday questions, getting help with a repair, or exploring a topic without typing every query. Voice thus becomes a more direct gateway to search and digital assistance.
Audio generated with SynthID watermarking
All audio generated by Gemini 3.1 Flash Live includes an imperceptible watermark created with SynthID. This signal is integrated directly into the audio output and makes it possible to reliably detect whether it was produced by artificial intelligence.
The measure aims to help combat misinformation and make synthetic content easier to identify. It doesn’t eliminate every risk associated with AI-generated voice, but it adds an important layer of traceability for platforms, companies, and users.
Gemini 3.1 Flash Live shows that the evolution of conversational AI isn’t only about making models faster. The challenge is getting them to listen better, reason for longer, and respond in a way that makes sense in a human conversation.
For developers, the advance lies in combining dialogue, tools, and context. For the rest of us, the difference shows up in something simpler: being able to talk to an AI without feeling like we need to learn a special language for it to understand us.
Original source
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-live
