Google shipped Gemini 3.6 Flash on Tuesday, July twenty first, and for once the headline number is not a benchmark, it is a bill. The new workhorse model uses about seventeen percent fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, and Google cut the output price from nine dollars to seven dollars fifty per million tokens at the same time. It arrived alongside Gemini 3.5 Flash Lite, and it is already rolling out inside GitHub Copilot for coding and agentic work. If you run agents that fan out tool calls across long tasks, cheaper tokens that also finish in fewer steps is the kind of change that shows up in your invoice, so here is what landed and how to think about it.
The short answer
Google released Gemini 3.6 Flash on July twenty first, a cheaper and more token efficient Flash model built for coding and agentic workflows. It spends about seventeen percent fewer output tokens than Gemini 3.5 Flash, drops the output price to seven dollars fifty per million, and posts higher scores on coding and computer use benchmarks. It shipped with Gemini 3.5 Flash Lite and is already rolling out in GitHub Copilot, while the flagship 3.5 Pro stays in testing.
Most model launches ask you to care about a leaderboard. This one asks you to look at your token meter. Gemini 3.6 Flash is not pitched as a new frontier ceiling, it is pitched as the everyday model you leave running under agents and coding tools, and the whole story is about doing the same work for less. For teams that watch spend on long agentic chains, that is a more useful promise than another point on a benchmark chart.
What Google shipped
On Tuesday, July twenty first, Google DeepMind released Gemini 3.6 Flash as the new default in its fast, low cost Flash line, replacing Gemini 3.5 Flash at the top of that tier. It landed together with Gemini 3.5 Flash Lite, the cheapest option in the family, so the update refreshes the whole workhorse tier rather than a single model.
The pricing is the part to write down. Input stays at one dollar fifty per million tokens, and output drops to seven dollars fifty per million, down from nine dollars on 3.5 Flash. Google describes the model as taking fewer reasoning steps and tool calls to finish multi step workflows, which matters because in agentic work the token count, not the sticker price alone, is what sets your cost.
The numbers that matter
Google reports that Gemini 3.6 Flash uses about seventeen percent fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to sixty five percent fewer on the DeepSWE coding benchmark. Pair that with the lower unit price and the effective cost of a real coding or agent run can fall further than the headline price cut suggests.
On raw capability, the model moves up across the tests this audience cares about. DeepSWE rises to forty nine percent from thirty seven, MLE Bench to sixty three point nine from forty nine point seven, and OSWorld Verified, which measures computer use, to eighty three from seventy eight point four. Its GDPval score on the Artificial Analysis version two index is one thousand four hundred twenty one against one thousand three hundred forty nine. None of that makes it a frontier flagship, and Google does not claim it is one, but a workhorse that codes better and costs less is exactly the trade most teams want.
Now rolling out in GitHub Copilot
The reason this lands for working developers today is Copilot. Gemini 3.6 Flash is generally available and rolling out in GitHub Copilot for Pro, Pro Plus, Max, Business, and Enterprise plans, selectable from the model picker. GitHub positions it for web and app development, coding, and longer horizon agentic tasks, and says early testing showed higher task completion rates and better token efficiency than 3.5 Flash.
Two practical notes. The rollout is gradual, so if the model is not in your picker yet, it should appear shortly. And on Copilot Business and Copilot Enterprise, an administrator has to switch on the Gemini 3.6 Flash policy in Copilot settings before your team can pick it, so if you manage a workspace, that toggle is the gate.
A workhorse tier, with Gemini 4 on the horizon
There are two bits of context worth holding alongside the launch. First, the flagship Gemini 3.5 Pro is still in limited partner testing rather than general release, so for now the Flash tier is where Google's fresh, broadly available capability actually lives. If you were waiting on Pro, 3.6 Flash is the model you can run this week. Second, Google said it has kicked off its most ambitious pre training run yet, for Gemini 4, a signal that the Flash refresh is a stepping stone, not the summer's last word.
For a working team the move is simple. If you already route agents or coding tasks through the Flash tier, test 3.6 Flash against your real workloads, watch the token count as closely as the pass rate, and let the numbers on your own traffic decide. The model supports configurable reasoning effort and parallel tool use, so you can turn thinking up for the hard problems and keep it cheap and fast everywhere else.
Sources and further reading
- Google: Introducing Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber
- TechCrunch: Google releases three new Gemini models, but no 3.5 Pro
- GitHub Changelog: Gemini 3.6 Flash is now available in GitHub Copilot
- MarkTechPost: Google releases Gemini 3.6 Flash, a cheaper, more token efficient Flash tier
- 9to5Google: Google launches Gemini 3.6 Flash and 3.5 Flash Lite, teases Gemini 4
Frequently asked questions
How much cheaper is Gemini 3.6 Flash?
Two ways. The list price for output dropped from nine dollars to seven dollars fifty per million tokens, with input at one dollar fifty per million. On top of that, Google says the model spends about seventeen percent fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to sixty five percent fewer on the DeepSWE coding benchmark. The token saving compounds with the lower unit price, so real agentic workloads can land well below a flat price comparison suggests. Benchmark your own traffic before assuming a fixed percentage.
Is Gemini 3.6 Flash available in GitHub Copilot?
Yes. Gemini 3.6 Flash is generally available and rolling out in GitHub Copilot for Pro, Pro Plus, Max, Business, and Enterprise plans. The rollout is gradual, so it may not appear in your model picker immediately. On Copilot Business and Copilot Enterprise, an administrator has to enable the Gemini 3.6 Flash policy in Copilot settings before anyone in the organization can select it.
How does it compare with Gemini 3.5 Flash on benchmarks?
Google reports gains across coding and agentic tests. On DeepSWE it scores forty nine percent against thirty seven for 3.5 Flash, on MLE Bench sixty three point nine against forty nine point seven, and on OSWorld Verified for computer use eighty three against seventy eight point four. Its GDPval score on the Artificial Analysis version two index is one thousand four hundred twenty one against one thousand three hundred forty nine. Google frames it as a workhorse tier rather than a frontier flagship.
What else did Google release with it?
Gemini 3.6 Flash shipped with Gemini 3.5 Flash Lite, the most cost efficient model in the class, as part of the same Flash tier update. Google also said it has begun its most ambitious pre training run yet, for Gemini 4. Separately, the flagship Gemini 3.5 Pro remains in limited partner testing rather than general release, so the Flash tier is where the fresh, broadly available capability sits right now.
Where can I use Gemini 3.6 Flash today?
Through the Gemini API in Google AI Studio and Android Studio, in the Gemini Enterprise Agent Platform, in Google Antigravity, and in the Gemini app, alongside the GitHub Copilot rollout. The model supports configurable reasoning effort and parallel tool use, so you can dial thinking up for hard tasks and keep it low for cheap, fast calls.