Google released Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash, at $0.75 per million input tokens and $3.75 per million output tokens. That introductory price is the part worth putting in a calendar, because it expires on December 31, 2026 and returns to $1.50 and $7.50 on January 1. The benchmark movement is real and lands almost entirely on agent work: DeepSWE v1.1 goes from 49.0 to 65.3 percent, AutomationBench from 17.0 to 30.4 percent. Google frames the model as a workhorse for coding and agents rather than a frontier release, and the three week gap between versions says roughly the same thing.
The short answer
Google released Gemini 3.7 Flash on August 13, 2026, describing it as its most intelligent workhorse model yet for coding and agents. It bills at $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026, then $1.50 and $7.50 from January 1. Context is 1 million tokens, maximum output is 64K, the knowledge cutoff is March 2026. The largest benchmark moves are on agent work: DeepSWE v1.1 from 49.0 to 65.3 percent, AutomationBench from 17.0 to 30.4 percent.
Read the release notes for a mid tier model and the benchmark table is usually the first thing you look at. This time the interesting line is in the pricing footnote.
What shipped
On August 13, 2026, Google released Gemini 3.7 Flash and called it its most intelligent workhorse model yet for coding and agents. It arrived three weeks after Gemini 3.6 Flash, which is a short enough gap to tell you what kind of release this is. Nobody pretrains a base model in three weeks. Coverage of the launch describes the gains as algorithmic optimisation work rather than a new pretraining run, and the specification supports that reading: the same 1 million token context window, a 64K maximum output, a March 2026 knowledge cutoff.
Availability is broad on day one. The model is in the Gemini API and Google AI Studio, in Google Antigravity and Android Studio, and across the Gemini Enterprise Agent Platform and app. Consumers reach it through Gemini Spark on the AI Pro and Ultra plans.
The price is the announcement
Gemini 3.7 Flash bills at $0.75 per million input tokens and $3.75 per million output tokens. Both numbers carry an expiry date: December 31, 2026. From January 1, 2027 the model costs $1.50 and $7.50.
Those second numbers are worth recognising, because they are exactly what 3.6 Flash cost when it launched. So the accurate description of what happened is not that Google cut the price of its workhorse tier. It shipped a better model at the same list price and put it on sale for four and a half months. That is still a good deal, and it is a different thing to plan around.
The competitive gap during the window is genuine. Claude Sonnet 5 lists at $2.00 and $10.00, GPT-5.6 Terra at $2.00 and $12.00. Against $3.75 on output, where agent workloads spend most of their money, that is roughly a third of the cost for a model scoring within a point or two on the same composite index.
Where the gains actually are
The headline number Google leads with is DeepSWE v1.1, up from 49.0 to 65.3 percent. The one we would look at harder is AutomationBench, up from 17.0 to 30.4 percent, because it nearly doubled and because a 17 percent score is the kind of number that means a capability did not really exist yet.
Line the moves up and the pattern is consistent. FrontierCode 1.1 Main goes from 34.4 to 43.6 percent. WebDev Arena goes from 1538 to 1588 Elo. Document comprehension on GDP.pdf goes from 22.0 to 34.0 percent. Long context recall sits at 97.0 percent on GDM-MRCR v2 at 128k, and CharXiv Reasoning at 84.5 percent. What improved is multi step work over long, messy inputs, which is precisely the workload a cheap model used to be disqualified from.
On the composite view, Artificial Analysis puts the Intelligence Index at 56, against 55 for Claude Sonnet 5, 52 for 3.6 Flash and 57 for both GPT-5.6 Terra and Muse Spark 1.2. Everything in this tier is now within a few points of everything else, which is why the vendors are competing on price and latency instead. The same week brought an OpenAI service tier chasing raw token throughput, and the two announcements are the same argument made from different ends.
What we would do with it
Pin the version. A three week release cadence in your dependency tree is faster than most teams' evaluation cycles, and a model string that silently points at whatever is newest is a change you did not test. Treat the id like any other pinned dependency, evaluate the new one against your own task set, then move.
Budget at the post January price. If a workload only clears its cost target at $0.75 and $3.75, it does not clear it, because that rate has a published end date and the number after it is double. Run the arithmetic at $1.50 and $7.50, and if the workload still works, the discount is four months of margin rather than a load bearing assumption.
Then watch the 64K output ceiling rather than the 1 million token input. The context window gets the attention, but in practice the limit teams hit is the response, usually when a model is asked to emit a whole file or a long structured document in one go. That failure is easy to design around if you know it exists and unpleasant to debug if you do not.
Sources and further reading
- Gemini 3.7 Flash model card, Google DeepMind
- Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut, VentureBeat, August 13, 2026
- Google AI Just Released Gemini 3.7 Flash, MarkTechPost, August 13, 2026
- Gemini 3.7 Flash launches three weeks after last model, live in Spark, 9to5Google, August 13, 2026
- Google's Gemini 3.7 Flash arrives before Gemini 3.5 Pro, Axios, August 13, 2026
- Gemini 3.7 Flash: On the Intelligence vs. Time per Task Pareto frontier, Artificial Analysis
- Google Releases Gemini 3.7 Flash, Competes With GPT 5.6 Terra and Muse Spark 1.2 On Benchmarks, OfficeChai, August 13, 2026
Frequently asked questions
What does Gemini 3.7 Flash actually cost?
The published introductory rate is $0.75 per million input tokens and $3.75 per million output tokens. That rate is explicitly time limited: it runs until December 31, 2026, and from January 1, 2027 the model bills at $1.50 and $7.50, which is exactly what 3.6 Flash cost at its own launch. So the honest way to read the announcement is that Google shipped a better model at the same list price and put it on sale for four and a half months. If you are sizing a budget for next year, size it at $1.50 and $7.50 and treat the discount as a windfall.
Where did the benchmark gains actually land?
On agent and workflow tasks far more than on raw knowledge. DeepSWE v1.1 moves from 49.0 to 65.3 percent and AutomationBench, which Google describes as measuring enterprise workflow automation, moves from 17.0 to 30.4 percent, close to a doubling. FrontierCode 1.1 Main goes from 34.4 to 43.6 percent and WebDev Arena from 1538 to 1588 Elo. Document comprehension moves too, with GDP.pdf going from 22.0 to 34.0 percent. The pattern is consistent: the tasks that improved most are the ones with many steps, tool calls and long inputs, which is where a cheap model previously fell over.
How does it compare with the other models in the same tier?
On Artificial Analysis's composite Intelligence Index, Gemini 3.7 Flash scores 56, ahead of Claude Sonnet 5 at 55 and its own predecessor at 52, and just behind GPT-5.6 Terra and Muse Spark 1.2 at 57. So on measured intelligence it is inside the same band as everything else in the mid tier. The separation is price. Claude Sonnet 5 lists at $2.00 and $10.00, GPT-5.6 Terra at $2.00 and $12.00, against $0.75 and $3.75 during the introductory window. On output tokens, which dominate agent bills, that is roughly a third of the cost.
What are the hard limits we should design around?
The context window is 1 million tokens and the maximum output is 64K tokens per response, which is the limit people actually hit first when they ask a model to emit a large file or a long structured document. The knowledge cutoff is March 2026, so anything more recent has to arrive through retrieval or tool calls rather than from the weights. Thinking configuration is exposed, meaning you can trade latency for depth per call rather than picking one setting for the whole application, and that control is worth wiring into config rather than hardcoding.
Why is Google shipping a new Flash model every three weeks?
Because these are not new pretraining runs. A three week gap is far too short for that, and coverage of the launch describes the gains as coming from ongoing algorithmic optimisation across the DeepMind teams rather than a fresh base model. Practically, that means version churn in this tier is now faster than most teams' evaluation cycles, and pinning a specific model string in production stops being optional. Treat the model id as a dependency with a version, evaluate the new one against your own task set before switching, and keep the old string reachable until you have.