
According to Google’s X post, these models aim to make AI agents faster, smarter, and cheaper at scale. Google’s new Gemini models: Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber 3.5 are built to deliver higher token efficiency, lower latency, and more reliable performance. Designed to strike the right balance between efficiency and quality, the Flash series enables scaling agentic workflows. Building on Gemini 3.5 Flash, these new models push forward the capabilities of production AI agents.
Exploring the New Gemini Models:
Flash 3.6: Google’s workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it cuts output token usage by 17% compared to Flash 3.5, and in some benchmarks like DeepSWE by Datacurve, the reduction goes up to 65%, all at a lower cost per output token.
3.5 Flash-Lite: Gemini’s quickest and most affordable 3.5-class model. According to the Artificial Analysis Index, it delivers 350 output tokens per second, which notably outperforms earlier Flash-Lite generations in agentic workflows.
3.5 Flash Cyber in CodeMender: Successful cybersecurity depends on combining models with strong agent infrastructure. Google is introducing a highly efficient, cyber‑focused model paired with its CodeMender security agent. Together, they deliver competitive performance at the cutting edge.
Flash 3.6: Smarter, Faster, Better :
Gemini 3.6 Flash was built based on feedback from developers and customers using 3.5 Flash. It takes coding and knowledge work a step further while using tokens more efficiently. For example, tests show it uses 17% fewer output tokens than 3.5 Flash, and it needs fewer reasoning steps and tool calls to handle complex tasks.
On top of that, it’s also cheaper. With pricing at $1.50 per million input tokens and $7.50 per million output tokens, 3.6 Flash lowers the overall cost of running AI agents, making them more affordable to build and use.
Comparing Flash 3.6 and Flash 3.5 :
Flash 3.6 delivers more precise results, with fewer unwanted code edits and fewer repeated execution loops. In DeepSWE tests, it reached 49% compared to 37% for Flash 3.5, and in ML research benchmarks like MLE Bench, it improved to 63.9% vs 49.7%. Flash 3.6 has enhanced computer use skills, as shown in OSWorld‑Verified (83.0% compared to 78.4%). Computer use is now available as a built‑in client‑side tool through the Gemini API and Gemini Enterprise. Flash 3.6 does better than Flash 3.5 in knowledge work, as shown in benchmarks like GDPval‑AA v2 (1421 compared to 1349). Customers such as Hebbia and Harvey have found it especially strong at multimodal tasks, including document parsing, chart and data analysis, and drafting reports.
Customers say Flash 3.6 is a big step forward, offering better quality at a lower cost
Stronger Safety Built In:
Flash 3.6 comes with stronger safety features to prevent misuse in areas like chemicals, biology, radiology, nuclear risks, and cyber attacks. These protections make it harder to break the model, while still allowing it to support helpful uses without unnecessary refusals.
Flash‑Lite 3.5: Designed for Scalable Workflows :
Flash‑Lite 3.5 is the quickest model in the 3.5 series, running at 350 output tokens per second. It’s priced at $0.30 per million input tokens and $2.50 per million output tokens. With much better quality than Flash‑Lite 3.1, it offers developers and customers a strong balance of speed, cost, and performance for handling high‑volume production traffic.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. It’s a significant step up in coding and agentic tasks, as seen in Terminal-Bench 2.1 (54% vs 31%), long context, as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution, as seen in GDPval-AA v2 (1140 vs. 642).
Flash Cyber 3.5: Fast, Reliable Vulnerability Fixes:
Flash Cyber 3.5 builds on Flash 3.5 and is specially tuned to spot and fix code security problems. Its strong speed and efficiency make it a solid base for detecting, checking, and patching vulnerabilities at scale. And because it runs at a lower cost per token than bigger models, it’s an affordable way to strengthen cybersecurity.
In CodeMender, multiple Flash Cyber 3.5 agents work together to create one combined report. On the popular CyberGym benchmark, this setup delivers strong, frontier‑level performance.
Flash Cyber 3.5 will soon be offered only to governments and trusted partners through CodeMender as part of a limited pilot program. This gives frontline defenders an early advantage in spotting and fixing critical vulnerabilities before they can be exploited, while reducing the risk of misuse
Beyond Today:
Flash 3.6 and Flash‑Lite 3.5 are available starting today. Developers can access them through the Gemini API in Google AI Studio and Android Studio, with Flash 3.6 also in Google Antigravity
Open to everyone in the Gemini app, and Flash‑Lite 3.5 is expanding into Google Search.
Beyond today’s launches, Gemini 3.5 Pro is being tested with partners and will be widely available once it’s ready. At the same time, Google and Gemini are already working on the next generation of models. They’ve begun their most ambitious training run yet for Gemini 4.
Original source: in