Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber—three models designed to address a very specific need in enterprise AI: running agents with lower latency, fewer tokens, and more controlled costs. The company also confirmed that Gemini 3.5 Pro is being tested with partners and that training for Gemini 4 has already begun.
Gemini 3.6 Flash reduces costs and reasoning steps
Gemini 3.6 Flash arrives as the flagship model in the Flash family. According to Google, it improves performance in programming, knowledge work, and multimodal tasks, while using fewer output tokens than Gemini 3.5 Flash.
In the Artificial Analysis Index, the company reports a 17% reduction in output token usage. In certain tests, such as Datacurve’s DeepSWE, the decrease can reach 65%. Why does that matter? An agent that needs fewer responses, tool calls, and reasoning steps can complete a task with lower cost and latency.
The announced price is $1.50 per million input tokens and $7.50 per million output tokens, below the price of Gemini 3.5 Flash, according to Google.
In the evaluations shared by the company, Gemini 3.6 Flash outperforms its predecessor in several scenarios:
- DeepSWE: 49% versus 37% on programming tasks.
- MLE Bench: 63.9% versus 49.7% on machine learning research.
- OSWorld-Verified: 83% versus 78.4% on computer use.
- GDPval-AA v2: 1421 versus 1349 on knowledge tasks.
Computer use is also being added as a client-side integrated tool in the Gemini API and Gemini Enterprise. This allows agents to interact with interfaces and applications to complete more complex workflows.
Google says companies such as Hebbia and Harvey have achieved strong results on multimodal tasks, including interpreting documents, analyzing charts and data, and preparing reports.
Stronger safety for high-risk uses
Gemini 3.6 Flash includes additional Frontier Safety safeguards to reduce misuse in chemical, biological, radiological, nuclear, and offensive cybersecurity contexts.
The goal is to make the model more resistant to attempts to bypass its restrictions, known as jailbreaks, without unnecessarily increasing refusals when a request has a legitimate purpose. Google provides more details in the model’s technical report.
Gemini 3.5 Flash-Lite focuses on speed and volume
Gemini 3.5 Flash-Lite is designed for tasks where every millisecond and every fraction of the cost matter. Potential uses include agent-powered search, large-scale document processing, information classification, and automating repetitive tasks.
According to the Artificial Analysis Index, it reaches a speed of 350 output tokens per second. Its price is $0.30 per million input tokens and $2.50 per million output tokens.
The model allows you to configure different reasoning levels. With minimal or low levels, a team can prioritize speed and affordability for high-volume tasks. At higher levels, it can handle processes involving multiple steps or subagents.
It also includes computer use as an integrated tool to support agentic tasks. In tests published by Google, Gemini 3.5 Flash-Lite improves on Gemini 3.1 Flash-Lite in areas such as:
- Terminal-Bench 2.1: 54% versus 31% on programming tasks from the terminal.
- GDM-MRCR v2: 72.2% versus 60.1% on long-context tasks.
- GDPval-AA v2: 1140 versus 642 on real-world task execution.
Google says that in some evaluations, it even surpasses Gemini 3 Flash. For example, it scores 54.2% versus 49.6% on SWE-Bench Pro and 74% versus 65.1% on OSWorld-Verified.
The results are promising for products that need to process thousands or millions of requests. However, as with any benchmark, real-world performance will depend on the type of data, the connected tools, and the quality of the agent’s architecture.
Gemini 3.5 Flash Cyber arrives with restricted access
Gemini 3.5 Flash Cyber is a model specialized in detecting and fixing software vulnerabilities. It is based on Gemini 3.5 Flash and was tuned to identify, validate, and repair security issues at a lower cost per token than larger models.
The model will be integrated into CodeMender, Google’s code security agent. Within this system, multiple Gemini 3.5 Flash Cyber agents work together to produce a combined report on the vulnerabilities they find.
In the CyberGym evaluation, Google says the combination achieves competitive performance at the cybersecurity frontier. The company will not make the model openly available for now because of its dual-use nature: the same capability that helps a defender could be used to find flaws for malicious purposes.
For that reason, Gemini 3.5 Flash Cyber will soon be available only to governments and trusted partners through a limited-access pilot program. The intention is to help fix critical vulnerabilities before they are exploited while also reducing the risk of abuse.
Availability of the new models
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are already available to different types of users:
- Developers: through the Gemini API, Google AI Studio, and Android Studio.
- Google Antigravity: Gemini 3.6 Flash is also available on this platform.
- Businesses: through the Gemini Enterprise Agent Platform. Gemini 3.6 Flash is also coming to the Gemini Enterprise app.
- General users: Gemini 3.5 Flash-Lite is beginning to roll out in the Gemini app and Google Search.
For those building agents, the difference is not simply having a more powerful model. It is also being able to choose when to pay for more reasoning and when to prioritize speed, volume, or cost. That flexibility can be decisive when moving from an interesting demo to a system that works every day in production.
Google says Gemini 3.5 Pro is still being tested with partners and will become publicly available when it is ready. Meanwhile, training for Gemini 4 has already begun, confirming that the race for more efficient and capable models continues to accelerate.
