Google introduces Gemini Omni, a new family of multimodal models that combines reasoning, content generation, and editing. Its first model, Gemini Omni Flash, can receive text, images, video, and audio as references to create high-quality clips, starting with video generation.
The proposal aims to make editing a video as simple as having a conversation. Do you want to change the setting, modify an action, or turn a sculpture into bubbles? Instead of learning a complex editor, you can describe the change in natural language.
Gemini Omni turns editing into a conversation
One of Omni’s core features is editing through successive instructions. Each prompt takes into account what happened before, so you can adjust the result over several turns without losing the continuity of the scene.
According to Google, the model is designed to keep characters consistent, respect the physics of objects, and remember important elements from the original video. In practice, you could start with a shot of a person walking, change the environment, add an object, and then modify the visual style without rebuilding everything from scratch.
It is also possible to transform the action in an existing video. An everyday recording can become a fantasy scene, incorporate new characters, or show an event that was never originally filmed.
The central idea is to move from manually editing every detail to directing the result with clear instructions.
More than realistic images: scenes with context
Gemini Omni is not limited to producing visually appealing images. Google says the model uses Gemini’s general knowledge to connect language, images, and meaning, while also representing physical phenomena such as gravity, kinetic energy, and fluid dynamics more effectively.
This can be useful for creating a chain reaction with objects that move coherently or for representing a scientific explanation through images. For example, a prompt could request a stop-motion animation in clay about protein folding, with a stop-motion aesthetic and an accurate visual explanation.
The combination of reasoning and generation also makes it possible to produce educational content. A user could request a video about history, science, or culture and receive a visual sequence that translates a complex idea into examples that are easier to follow.
Of course, this does not mean that every result is automatically correct. In educational or scientific topics, human review is still necessary. The ability to create a convincing explanation does not guarantee that all its facts are accurate.
Videos from any combination of references
Omni can combine different types of input to produce a cohesive clip. You can provide an image of a character, a drawing of a scene, a reference video, or a written instruction to define the result.
You can also describe the style, movement, and effects using natural language. The tool attempts to bring these references together in one visual piece, rather than treating them as isolated elements.
Google says that, initially, audio can be used as a voice reference. The company plans to add other types of audio inputs later. As for outputs, the first stage focuses on video, although image and audio modalities are expected in the future.
A practical example
Imagine you have a character sketch, a photograph of a setting, and an audio recording with narration. With Omni, you could request a video that uses the character from the drawing, preserves the atmosphere of the photograph, and synchronizes the scene with the reference voice.
The result will depend on the quality of the references and the precision of the instruction, but the workflow eliminates several steps that previously required different programs.
Digital avatars and safety controls
Gemini Omni will also allow users to create videos with a digital avatar based on their own voice. The feature aims to generate a digital version of a person to produce videos that look and sound like them.
Google says it is still responsibly testing features related to audio and speech modification. This is important because tools capable of imitating identities can have legitimate uses, but they can also enable impersonation, fraud, and disinformation.
All videos created with Omni include an imperceptible digital watermark called SynthID. Google says this signal makes it possible to verify whether a video was generated with Gemini Omni from the Gemini app, Gemini in Chrome, and Google Search.
The watermark does not replace journalistic verification or solve the problem of synthetic content on its own. However, it provides an additional layer for identifying the origin of a piece when generation technology is difficult to detect at a glance.
Gemini Omni Flash availability
The first model in the family, Gemini Omni Flash, has begun its global rollout to Google AI Plus, Pro, and Ultra subscribers through the Gemini app and Google Flow.
Google also announced that it will be available at no cost to YouTube Shorts users and users of the YouTube Create app. Availability may vary depending on the product, country, and rollout stage.
In the coming weeks, the company plans to make the model available to developers and enterprise customers through APIs. This could open the door to marketing, education, audiovisual production, and content creation tools integrated directly into other applications.
Gemini Omni represents an important evolution in video generation because it does not simply try to create a scene from a prompt. Its goal is to let a person reason about the scene, modify it step by step, and combine different references without having to master a technical editing workflow.
The question is no longer only whether AI can produce a video. It also matters whether it can maintain a creative intention throughout the process, understand what needs to change, and help turn an incomplete idea into a coherent visual story.
Original source
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni
