Gemini 3.6 Flash claims 65% agent token savings, but 3.5 Pro stays in test

Google's DeepMind released three proprietary Gemini models today, with Gemini 3.6 Flash posting token reductions as high as 65% on long-horizon engineering tasks compared to its predecessor, while the company's flagship Gemini 3.5 Pro remains in private testing. The release is efficient in the literal sense: cheaper per million tokens, lighter on output, and built for the kind of multi-step workflows where agentic systems burn through context. The release is strategically narrower: it fills the middle and lower tiers of Google's lineup while leaving the frontier position unfilled, and it gates the cybersecurity variant behind a closed partner program.

The headline number is the 65% reduction in tokens on the DeepSWE long-horizon software engineering benchmark, where Gemini 3.6 Flash also improved its absolute score from 37% to 49%. According to the third-party Artificial Analysis Index, the model cuts total output tokens by 17% against its predecessor across mixed workloads, and the same source reports that on multi-step agentic flows the savings climb much higher. The model cards for both 3.6 Flash and 3.5 Flash-Lite list a 1-million-token input context window, a 64,000-token output ceiling, and a March 2026 knowledge cutoff.

On cost, Google is pricing 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite is the cheaper sibling at $0.30/$2.50. The pricing comparison in the source places Google's prior Gemini 3.1 Flash-Lite at $0.25/$1.50, still the most cost-efficient model in Google's lineup, but Google says the older Lite is twice as slow as the new 3.5 Flash-Lite. Artificial Analysis pegs 3.5 Flash-Lite at 350 output tokens per second, roughly twice the rate of 3.1 Flash-Lite.

The speed and price point both matter, but they are not the same variable. A team running high-volume agentic search or document processing sees the latency win directly. A team running complex multi-step engineering work sees the token-efficiency win. 3.5 Flash-Lite is positioned for the first case; 3.6 Flash is positioned for the second. Google's split maps to the two usage patterns the source identifies, and the structure of the new product line reads as a direct response to both.

The benchmark table attached to the release shows movement beyond token efficiency. 3.6 Flash reaches 63.9% on MLE-Bench against the predecessor's 49.7%, and 83.0% on OSWorld-Verified against 78.4%. On the agentic-knowledge work GDPval-AA v2, the score moves from 1349 to 1421. 3.5 Flash-Lite, despite its lite designation, outperforms the standard Gemini 3 Flash on SWE-Bench Pro (54.2% versus 49.6%) and on OSWorld-Verified (74.0% versus 65.1%). The direction is consistent across the suite, and the magnitudes are large enough that the source does not need to argue capability, only price.

The third product, Gemini 3.5 Flash Cyber, is the one Google is not putting on the open market. The source says it is fine-tuned for vulnerability discovery and integrates with Google's CodeMender agent, which the company released last year as an AI code bug-fixing system. Google has not published a price, stating only that the model runs at a lower per-token cost than larger models. Access is restricted to governments and trusted partners, which the source describes as a closed pilot. The pattern resembles Anthropic's Mythos-Project Glasswing gate and OpenAI's staggered rollout for GPT-5.6: a deliberate tiering of who can probe cyber-offensive capabilities.

The flagship gap is harder to ignore when read against the rest of the release. Gemini 3.1 Pro Preview debuted in February 2026, and rivals have shipped several frontier updates since. Google technical staffer Logan Kilpatrick responded to developer questions about the missing Gemini 3.5 Pro by writing that the model is "currently testing with partners and we plan to make it broadly available as soon as it's ready." The source also reports that pre-training for Gemini 4 has already begun, so the missing 3.5 Pro is not explained as a paused effort; the source frames it as a sequencing decision the company has not closed. The cost-efficient Flash line is what Google is selling today, and the source's own framing treats that line as the immediate future of agentic AI.

Both consumer-facing models are proprietary and distributed through the Gemini API and select consumer surfaces, including the Gemini app and Google Search. Developers do not receive model weights or training data, and there is no self-hosting path without a higher-tier enterprise agreement. The source frames this commercial tethering as a meaningful constraint on deployment flexibility, especially for teams that need air-gapped operation or fixed monthly cost rather than metered usage. For the Cyber variant, that same constraint is the gate that keeps it out of open distribution.

The release is a coherent efficiency play, and the benchmark gains are large enough to take seriously. The Flash-tier efficiency and the agentic scores will need to hold their position against the next wave of frontier releases from OpenAI and Anthropic, neither of which has been standing still in the months since 3.1 Pro shipped. Google's answer, for now, is to ship the lighter models first and keep the flagship in test.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe