Google presented nine demonstrations of Gemini Omni and Gemini 3.5 Flash during Google I/O 2026. The first line focuses on creating and editing video with multiple types of input; the second, on agents capable of executing complex tasks, writing code, and building personalized experiences.
The key difference lies in their approach. Gemini Omni aims to turn images, audio, video, and text into new audiovisual pieces. Gemini 3.5 Flash focuses on reasoning, taking action, and quickly completing multi-step workflows.
Gemini Omni: create and edit video through conversation
1. Change elements in a scene using natural language
With Gemini Omni, someone can take an existing video and request specific changes without having to master a traditional editor. For example, the demonstration transforms a sculpture so that it appears to be made of bubbles.
The idea is for each instruction to build on the previous one. The model tries to preserve the characters’ identities, the scene’s continuity, and physical coherence while modifying the content.
Video stops being the final result of a recording and becomes the starting point for creating something new.
2. Reimagine a video’s action
Omni can also alter what happens within a shot. You can change the lighting, add objects or characters, and turn an everyday action into a completely different sequence.
One demonstration suggests darkening a room and creating a glass sphere floating above a hand. Inside the sphere, a recursive representation of the same hand appears, with rooms repeating until they form an infinite visual effect.
This type of editing shows the model’s potential for working with complex creative instructions, not just simple color or cropping adjustments.
3. Refine a video over multiple rounds
Editing can continue across multiple conversational turns. The user can change the setting, camera angle, visual style, or specific details without having to describe the entire scene again.
Technically speaking, the challenge is maintaining visual context across multiple generations. If the model loses a character’s appearance or changes elements that were supposed to remain untouched, the editing stops being useful. Google says Omni is designed to preserve that narrative thread.
Gemini 3.5 Flash: agents that reason and act
4. Automate tasks with many steps
Gemini 3.5 Flash is designed for tasks known as agentic tasks, in which the model does more than answer a question. It can also plan a sequence, use tools, and review results to achieve a goal.
In a demo with Antigravity, the model renames and classifies disorganized assets according to criteria that can change dynamically. This scenario resembles the work done by design, marketing, or development teams when they need to organize large quantities of files.
5. Coordinate subagents with Antigravity
By combining 3.5 Flash with the updated Antigravity environment, Google shows several subagents collaborating under supervision. Each one can handle part of the problem while the system coordinates the complete workflow.
This approach is especially useful for programming tasks and complex operations. Instead of asking a single model to do everything, responsibilities are distributed and human oversight is maintained to validate important steps.
The speed of the Flash family is part of the proposition. An agent that reasons well but takes too long with each action may not be practical in a real workflow.
6. Create interactive interfaces and graphics
In AI Studio, Gemini 3.5 Flash generates different user experience proposals for a checkout flow in approximately 60 seconds, according to Google.
The goal is not just to produce code. The model can explore several design alternatives and turn an idea into a more interactive interface. For an entrepreneur validating an online store, this could speed up the comparison between designs before investing time in a final implementation.
New agent features in Search and Gemini
7. Information agents that work in the background
Google is also integrating 3.5 Flash’s agentic capabilities into Search. The so-called information agents can follow specific topics and send updates when relevant information appears.
One demonstration shows an agent monitoring whether certain athletes announce collaborations with footwear brands or launch exclusive products. The update would include a summary and links to consult the original sources.
These agents introduce an important difference compared with a traditional search: they do not wait for someone to submit a new query. Instead, they monitor a topic based on prior instructions.
Google says this feature will first reach Google AI Pro and Ultra subscribers during summer 2026.
8. Generate visual tools and personalized experiences
Search will also be able to create interactive answers adapted to the question. Instead of showing only links and text, the system can build a visual tool, a simulation, or a generative interface in real time.
Among the examples presented is an interactive explanation of gyroid patterns. For ongoing tasks, such as planning a wedding or establishing an exercise routine, Search could generate dashboards, trackers, or small applications that the user can consult later.
Google says generative interface features will be available to all Search users during the summer at no cost. The creation of personalized experiences with Antigravity will arrive in the following months, initially for Pro and Ultra subscribers in the United States.
9. Gemini Spark as a personal agent
The ninth demonstration presents Gemini Spark, a personal agent based on Gemini 3.5 and the Antigravity environment. Its purpose is to help manage digital tasks continuously, always under the user’s direction.
Spark is integrated with Workspace tools such as Gmail, Docs, and Slides. In the example, it creates a list of nut-free snacks and then adds them to Instacart.
The usefulness is clear, but it also demands precise controls. An agent that can act on someone’s behalf needs limited permissions, confirmations for sensitive actions, and a simple way to review what it has done.
Gemini Omni and Gemini 3.5 Flash availability
Google reports that Gemini Omni Flash is rolling out to Google AI Plus, Pro, and Ultra subscribers worldwide through the Gemini app and Google Flow. It will also be available at no cost in YouTube Shorts and the YouTube Create app.
The company plans to offer Omni to developers and enterprise customers through APIs in the following weeks. For users of these platforms, the availability of multimodal models will be especially relevant because it will make it possible to incorporate video generation and editing into their own products.
Meanwhile, Gemini 3.5 Flash is generally available in Google Antigravity, the Gemini API within Google AI Studio, Android Studio, Gemini Enterprise Agent Platform, and Gemini Enterprise. It is also available to all users in Search’s AI Mode and is being rolled out globally in the Gemini app.
The presentation sends a clear signal: Google is trying to move AI from generating responses toward executing tasks. Omni wants creating video to be as conversational as asking for an idea. 3.5 Flash aims to let agents turn that intention into concrete actions, from organizing files to building a personalized tool.
The real challenge will not just be proving that these models can do this. It will be making sure their results are consistent, verifiable, and safe when they begin working with real information and services.
Original source
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-3-5-videos
