Google introduces Gemini 3.1 Flash TTS, a text-to-speech model designed to produce more natural, expressive, and controllable conversations. The proposal is not limited to reading text aloud: it allows you to direct a performance with instructions about tone, pace, accent, and speaking style.
The version is already beginning to roll out in preview to developers through the Gemini API and Google AI Studio, to businesses through Vertex AI, and to Workspace users through Google Vids.
A more natural voice that is easier to direct
Gemini 3.1 Flash TTS improves the overall quality of speech synthesis and adds more precise controls for creating audio experiences. According to Google, the model reached an Elo score of 1,211 in the Artificial Analysis ranking, an evaluation based on thousands of comparisons made by people who did not know which system produced each audio clip.
Why does this matter? The evaluation does not only measure whether a voice pronounces words correctly. It also considers whether it sounds natural, convincing, and pleasant compared with other alternatives.
Artificial Analysis placed the model in its quadrant of most attractive options because it combines generation quality with relatively low costs. In addition, Gemini 3.1 Flash TTS includes native dialogue between multiple speakers and support for more than 70 languages.
Audio tags for controlling the performance
One of the main new features is audio tags, written instructions in natural language that let you modify how a sentence should sound. For example, a developer could request a slower delivery, an excited tone, or a surprised reaction in the middle of a sentence.
This approach brings voice creation closer to directing actors. Instead of configuring dozens of technical parameters, you can describe the scene and ask the model to adjust the performance.
The developer takes on the role of director
Google AI Studio includes configurable controls for organizing a complete voice performance:
- Scene direction: lets you define the setting, context, and dialogue instructions so characters maintain their personality across multiple turns.
- Speaker configuration: each character can receive their own audio profile, along with guidance on pace, tone, and accent.
- Changes within a sentence: tags inserted into the text let you vary the expression at specific moments without changing the character’s entire configuration.
- Export to the API: when the result is ready, the parameters can be converted into code to reproduce consistent voices across different projects and platforms.
Imagine an audiobook with several characters, an educational assistant that adapts its tone to the student’s age, or a video game with dialogue that reacts to the context. These features aim to make generated voices sound less like a uniform reading and more like an intentional performance.
Designed for global applications
Support for more than 70 languages points to international use cases such as customer service, training, content localization, and accessibility tools. The combination of control over accent, pace, and style can help create less generic experiences for each market.
For a company, this means that localizing an application would not have to consist solely of translating the text. It would also be possible to adjust how the message is delivered: more formal in a corporate setting, friendlier in a learning app, or more dynamic in an audiovisual campaign.
SynthID to identify generated audio
All audio produced by Gemini 3.1 Flash TTS includes an imperceptible SynthID watermark. This signal is integrated directly into the audio and is designed to make it easier to detect content generated by artificial intelligence.
The measure aims to contribute to the identification of synthetic voices and reduce risks related to misinformation. However, a watermark does not replace human verification or responsible-use policies, especially when the audio imitates real people or is used in sensitive contexts.
What changes for developers
Gemini 3.1 Flash TTS combines three relevant elements: voice quality, creative control, and availability in development tools. The preview will allow developers to test the model before incorporating it into larger-scale products, while Vertex AI offers a path designed for enterprise environments.
The most interesting point is that speech synthesis is moving closer to a creative direction interface. It is no longer just about choosing a voice and changing its speed. Now you can define a scene, assign personalities, and adjust the performance almost as if you were working with a digital cast.
That said, the final results will depend on the language, context, quality of the instructions, and preview availability conditions. The technology promises more expressive voices, but you will still need to evaluate latency, costs, consistency, and safety controls before using it in production.
Original source
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts
