Gemini 3.8 Flash and 3.8 Flash Cyber: Google's Third Flash Release in Six Weeks

Google has released Gemini 3.8 Flash and a security-focused sibling, Gemini 3.8 Flash Cyber, on September 2. By Google's own count it is the third Flash release in six weeks, arriving about three weeks after Gemini 3.7 Flash on August 13. The company describes 3.8 as its best reasoning and coding model yet, delivered at the same speed and price as 3.7. For developers and businesses building AI agents, the release is a strong signal that the cheap, fast tier of models is now where much of the competition is happening. Here is what was announced, what the numbers do and do not tell you, and how to decide whether to switch.
Two models, one shared core
Google positions the two models as tailored for different deployment settings but powered by the same underlying intelligence. It says the gains come from a set of innovations, including training in the demanding domain of cybersecurity and long-running agentic loops designed to recursively evaluate and refine the models. That description is high-level, and the company does not spell out the details, so treat it as marketing context, not as a technical explanation.
- Gemini 3.8 Flash is the general model, pitched at long-horizon coding and autonomous agents.
- Gemini 3.8 Flash Cyber is a specialist model for vulnerability detection and automated patching, available only to trusted defenders.
Gemini 3.8 Flash: more diligence, more tokens
Google says 3.8 Flash makes significant improvements over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning in specialized domains, and often approaches the performance of higher-cost frontier models. On DeepSWE v1.1, a benchmark for long-horizon software engineering, it is said to outperform most larger frontier models at solving complex engineering problems end to end, at a fraction of the cost. On professional benchmarks such as Vals Finance Agent V2 and Harvey's Legal Agent Benchmark it is said to beat 3.7 Flash and other frontier models, and it scores 54.9 percent on HLE-Verified, a broad multi-step reasoning test.
Two cautions apply. First, these are vendor-reported benchmarks and the comparison sets are chosen by the vendor, so you should confirm them on your own tasks. Our explainer on what AI benchmarks do and do not measure shows why a leaderboard position rarely settles a purchasing decision. Second, Google is candid about a trade-off: the model works harder. On complex tasks it takes extra reasoning steps and calls tools iteratively, and it can use more tokens to maximize performance, especially at higher effort levels. If cost efficiency is your main concern, Google suggests lower effort settings or staying with 3.7 Flash, which remains fully supported.
That last point is worth dwelling on. A model that is cheaper per token but uses more tokens per task does not automatically cost less overall, so measure cost per completed task, not the headline rate.
Pricing and where you can use it
The introductory price for 3.8 Flash matches 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. Google notes that this introductory pricing expires on December 31, 2026, and from January 1, 2027 the standard rates of $1.50 per million input tokens and $7.50 per million output tokens apply. If you are budgeting for next year, use the higher figures.
Availability is broad:
- Developers can use it in Google Antigravity, the Gemini API through Google AI Studio and Android Studio, and Stitch for generating interfaces.
- Enterprises can access it in Gemini Enterprise.
- Consumers on Google AI Pro and Ultra get it in the Gemini app, AI Mode in Google Search and Gemini in Google Sheets.
The consumer note is important: 3.8 Flash is not automatically available to every free user, so check your subscription before expecting it in the app. For a practical comparison of how Gemini fits into everyday work next to other assistants, see our Gemini and ChatGPT workflow comparison.
Gemini 3.8 Flash Cyber and the Fairwind Program
The Cyber model is the more unusual half of the release. Google calls it its most capable cybersecurity model, with frontier-level performance in finding vulnerabilities and automatically patching them, and it is offered through a new Fairwind Program to trusted government authorities, critical infrastructure operators and software maintainers who apply for access. Google says it deliberately prioritized fixing over offensive capabilities such as exploitation.
The reported numbers are notable:
- On CyberGym, a standard benchmark for finding vulnerabilities, it shows frontier-level performance and beats both 3.5 Flash Cyber and much larger frontier models.
- On an internal benchmark covering 20 programming languages, its success rate exceeds 70 percent.
- On CWE-Bench, a patching benchmark run by Collinear, it reaches a pass@1 of 47.2 percent, compared with 47.8 percent for a leading frontier model, at a significantly lower cost.
Google also gives examples of internal use. It says Chrome's security team found the model produced 2.6 times as many correct patches as the best commercial models, which are much larger, and that Wiz measured 7.5 to 9.7 percent higher recall on its penetration-testing benchmark at 2.3 to 5.2 times lower cost. In another example, Google's Cloud Vulnerability Research team used it to find a critical foundational vulnerability in under two hours, a task that Google says normally takes months. These are the company's own accounts and have not been independently verified here.
Restricted access is the safety story. Google says 3.8 Flash ships with safeguards against misuse in chemical, biological, radiological and nuclear areas and in cyber offense, while the Cyber model has a more permissive set of mitigations and is therefore limited to defenders. Both models are also reported to be significantly more robust against prompt injection, as measured by Gray Swan. Prompt injection is a real risk for any AI that reads untrusted content and takes actions, which is central to the rise of agents; our guide to how AI agents work covers why it matters.
What the demos show
Google's examples include a 3D wizard-in-a-castle game built from a simple prompt with a looping instruction in Antigravity, a playable DOS-style version of a maps app, a topographic explorer using U.S. Geological Survey data, and an interactive hardware teardown visualizer built in AI Studio. Demos are chosen to impress, so they show what is possible, not what is typical. Real projects are messier, and the useful question is how often the model finishes a task without you stepping in. For coding work specifically, our comparison of AI coding assistants outlines how to evaluate them on your own codebase.
How to decide whether to switch
Run a small evaluation before changing anything in production. Take twenty to fifty real tasks, run them on 3.7 Flash and 3.8 Flash at a couple of effort levels, and compare the results on success rate, tokens used and time. Keep a note of the cost per successful task. If the improvement is real for your work, the switch is easy because the introductory price is identical. If you handle sensitive code or run a security team, look into the Fairwind Program and read the eligibility terms. If you are a casual user, the practical step is simply to check whether your subscription includes the new model and try it on a task you already know well. The main takeaway is that faster, cheaper models keep closing the gap on more expensive ones, but the only benchmark that counts is your own workload.


