Google's DeepMind team released a new model card for Gemini 3.7 Flash on August 13, 2026, with coding performance jumping to 65.3% on the DeepSWE v1.1 benchmark compared to 49.0% on the previous Flash version. The model comes with introductory pricing slashed by half through the end of the year, but those rates double on January 1, 2027, creating a tight window for teams to evaluate whether the performance gains justify scaled production use. DeepMind positions the release as a mid-tier workhorse built for agentic workflows, coding tasks, and enterprise operations.

The model delivers sharp improvements across multiple coding benchmarks, with FrontierCode 1.1 climbing from 34.4% on Gemini 3.6 Flash to 43.6% on the new release. Long-context performance also rose, with the GDM-MRCR v2 8-needle score moving from 91.8% to 97.0%. Terminal-bench 2.1 hit 85.8%, and Code Arena Web development scored 1588 Elo. Introductory pricing runs at $0.75 per million input tokens and $3.75 per million output through December 31, 2026, then doubles to standard rates of $1.50 and $7.50 starting January 1, 2027. The context window remains at 1 million tokens in and 64,000 tokens out, and the model is available through Google AI Studio, the Gemini API, Google Antigravity, Gemini Enterprise, the Enterprise Agent Platform, and the Spark tier of the consumer Gemini App. Gemini 3.7 Flash arrived three weeks after version 3.6, while the Pro release cadence is slowing.

DeepMind describes the release as incorporating "algorithmic improvements to its core reasoning foundation" and "customizable thinking configurations to control the mix of quality, cost and latency." On safety, the card states the model "performs similarly to Gemini 3.6 Flash across both safety and tone, with low unjustified refusals." The report flags several gaps, noting that Terminal-bench 3.0, the more difficult version of the 2.1 suite, scores just 14.9%, and the card doesn't clarify what drives that gap. The knowledge cutoff sits at March 2026, with some domains capped at January 2025, and all benchmark figures come from Google's own testing without independent replication so far.

The report emphasizes that the 50% introductory price expires on January 1, 2027, doubling the per-token cost for any team that builds production usage on it now. For teams already running coding agents on Flash-class tokens, the practical move is to conduct internal evaluation during the fourth quarter while the cheaper rate still holds, then decide whether the January price jump alters unit economics before committing to scale. The report notes the same cost-benefit calculation is happening at Anthropic, where session patterns are shaping Claude Code expenses. Coverage from other outlets frames the launch within diverging economics by model tier, with Flash commoditizing rapidly while Pro slows, and enterprise governance still blocking agent adoption despite falling prices. The three-week release cadence between Flash versions signals a two-speed product strategy by tier, with faster iteration on the lower-cost line. Pricing cliffs and benchmark performance matter most when they shift the math on whether to automate, and that window closes in less than five months.