Google has released Gemini 3.8 Flash — the third Flash model in six weeks, just three weeks after Gemini 3.7 Flash. The introductory price matches the previous version, but on complex tasks the model works harder: it performs more reasoning steps, calls tools more often, and consumes more tokens, while reduced effort levels and full support for 3.7 Flash remain available for cost-conscious scenarios. At the same time, Gemini 3.8 Flash Cyber was introduced for autonomous vulnerability discovery and patching, showing frontier-level performance on CyberGym, although its distribution is limited to the Fairwind program.



What happened
Google introduced Gemini 3.8 Flash just three weeks after Gemini 3.7 Flash — the third Flash model in the past six weeks. The introductory pricing matches 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026, then $1.50 and $7.50 respectively. The model scores 54.9% on HLE-Verified, accepts up to 1 million tokens of context, outputs up to 64K tokens, and has knowledge of the world as of March 2026. In parallel, Gemini 3.8 Flash Cyber was released for autonomous vulnerability discovery and patching: on CyberGym it outperforms 3.5 Flash Cyber and larger models, showing 47.2% pass@1 on CWE-Bench on the Pareto frontier versus 47.8% for the leading model at a significantly lower price, demonstrates over 70% success on an internal benchmark across 20 programming languages, and found a critical vulnerability in Google Cloud in less than two hours instead of months.
Context
The release solidifies Google's shift to nearly monthly releases of cheap workhorse models: a third Flash release in six weeks means there is almost no time between releases for independent reproductions and comparative runs. The economic logic of the release is that quality gains at an unchanged per-token price are achieved by increased compute consumption on complex tasks, so the correct metric for comparing models becomes not the per-token price but the price per solved task. Google explicitly leaves reduced effort levels and full support for 3.7 Flash for cost-conscious users, acknowledging that the new model is not needed in all scenarios. Cybersecurity is being shaped as a separate product track with closed distribution — following the same limited-access logic as Anthropic's and OpenAI's cyber models: leading results in a dual-use area are demonstrated on vendor and internal benchmarks, while access is distributed by application.
Why this matters for the industry
For the industry, the release is not a new price point but a new quality knob: with an unchanged per-token rate, effort levels, token-consumption telemetry, and routing of cheap requests to 3.7 Flash — which remains fully supported — become mandatory elements of LLM products. Startups gain a capability boost in the Flash class — up to 1 million tokens of context and up to 64K output allow working with entire codebases and long documents — but pay for it in billing predictability: those who calculate cost on their own tasks win. If the cadence of "a Flash model every few weeks" continues, one-off migrations will be replaced by a permanent eval pipeline with regression suites on real tasks, automatic version comparison, and routing by cost per solved task. By the end of the introductory period on December 31, 2026, when the reduced rate expires, teams without such calculations risk facing increased expenses, and product unit economics will be counted in solved tasks rather than tokens; a wave of tools around automatic model and effort-level selection, cost telemetry, and quality regression testing is also expected.
Why this matters for users
Gemini 3.8 Flash can be tried today in the Gemini API and Google AI Studio, as well as in Antigravity, Android Studio, Gemini Enterprise, and the Gemini app with an AI Pro or Ultra subscription. The per-token price has not changed, but on complex tasks the bill may increase due to higher token consumption, so when working with the API it is worth setting default effort levels and tracking actual per-session consumption, and before migrating pipelines from 3.7 to 3.8 — measure the cost per solved task on your own data rather than relying on the rate. Migration is not forced: 3.7 Flash is fully supported, and cost-conscious scenarios can stay on it. Up to 1 million tokens of context and up to 64K output open up scenarios with long documents and entire codebases right now. Gemini 3.8 Flash Cyber is not available to regular developers — access is distributed by application through the Fairwind program.
What is still unknown / limitations
The key figure of 54.9% on HLE-Verified is given without an evaluation protocol: the sources do not disclose task subsets, effort settings, or the number of runs, so the result should be considered a vendor report rather than a reproduced one; there are no independent third-party runs yet, and the frequent release cadence leaves little time for them to appear. There is no latency data in the sources at all — real-time services will require their own measurements. Gemini 3.8 Flash Cyber results were obtained on vendor and internal benchmarks, including CyberGym and an internal set across 20 languages, and closed distribution through the Fairwind program makes independent verification of these metrics difficult.
Sources
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber — The Keyword, Google
- Gemini 3.8 Flash — Model Card, Google DeepMind
- Fairwind Program — Google DeepMind
Author
Look at AI, editorial team
