Zendoric
← Back to the day · September 4, 2026

Google launches Gemini 3.8 Flash and 3.8 Flash Cyber, its most powerful models for agentic coding and cyber defense

🕒 Published on Zendoric: September 4, 2026 · 09:12

✨ AI-generated · how it's made

Google has unveiled Gemini 3.8, its new generation of «Flash» models, in two variants: Gemini 3.8 Flash, designed as a workhorse for software engineering tasks, agentic workflows and multi-step reasoning in specialized domains; and Gemini 3.8 Flash Cyber, a model built specifically for…

Google has unveiled Gemini 3.8, its new generation of "Flash" models, in two variants: Gemini 3.8 Flash, designed as a workhorse for software engineering tasks, agentic workflows and multi-step reasoning in specialized domains; and Gemini 3.8 Flash Cyber, a cybersecurity-specific model geared toward vulnerability detection and automatic patching, available only to a group of "trusted defenders" through a new program called Fairwind. The launch comes barely three weeks after Gemini 3.7 Flash and is the third version of the Flash family in just six weeks, confirming the accelerated release pace Google has adopted in its competition with other AI labs to dominate the field of autonomous coding agents.

According to the company itself, both variants share a common intelligence base, reinforced through long-running agentic loops that recursively evaluate and refine the models, as well as intensive training in the demanding domain of cybersecurity, which Google says has also contributed to the broader gains in reasoning and coding.

Gemini 3.8 Flash keeps the same introductory price as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens (that rate expires on December 31, 2026; from January 1, 2027 it will rise to $1.50 and $7.50 respectively). Google claims that the model, despite maintaining its predecessor's cost and speed, achieves a substantial performance jump, in several cases approaching far more expensive frontier models. On the DeepSWE v1.1 benchmark, focused on long-horizon autonomous software engineering, it says the model outperforms most larger frontier models by solving complex problems end to end at a fraction of the cost. Google also highlights results superior to 3.7 Flash and to other frontier models on professional-domain benchmarks such as Vals Finance Agent V2 (finance) and Harvey's Legal Agent Benchmark (law), as well as 54.9% on HLE-Verified, which measures multi-step reasoning across science, the humanities and professional fields.

Google explains this improvement as a deliberate design decision: the model "works harder", running additional reasoning steps and calling tools iteratively, which can translate into higher token consumption, especially at higher effort levels. For those who prioritize computational efficiency, the company offers lower effort levels or recommends continuing to use Gemini 3.7 Flash, which remains available. As practical demonstrations, the article mentions examples built with a single prompt in Google Antigravity: a 3D video game set in a castle with textures generated using Nano Banana, a working version of Google Maps for DOS that can be navigated with directions and Street View, and an interactive topographic map with real data from the United States Geological Survey (USGS). It also cites "Hardware Anatomy", a 3D visualizer built in Google AI Studio that generates Three.js renderings of exploded views of electronic devices.

On the cybersecurity front, Gemini 3.8 Flash Cyber is presented as Google's most capable model to date in this area, with an explicit focus on giving defensive teams an edge over attackers: the company says it prioritized from the outset the ability to fix vulnerabilities over offensive capabilities such as exploitation. On the industry-standard CyberGym benchmark, the model reportedly showed frontier-level performance in autonomous vulnerability discovery, outperforming both 3.5 Flash Cyber and larger frontier models. In its own internal evaluation, covering 20 different programming languages —beyond the C/C++ that predominates in CyberGym—, the model reportedly reaches a success rate above 70%. As for automatic patching, on the external CWE-Bench benchmark (managed by Collinear) the model reportedly sits on the "Pareto frontier", with a pass@1 of 47.2% versus 47.8% for a leading frontier model, but at a significantly lower cost.

Google also offers examples of internal use of the model to secure its own code: the Chrome Security team reportedly obtained 2.6 times more correct patches for vulnerabilities than the best larger commercial models; the security firm Wiz reports between 7.5% and 9.7% more recall on its internal penetration-testing benchmark, at a cost 2.3 to 5.2 times lower than other frontier models; and Google's Cloud Vulnerability Research team says it found a critical vulnerability in under two hours using the model, a process that would normally require months of research.

On safety and misuse, Google indicates that Gemini 3.8 Flash incorporates safeguards against improper use in the chemical, biological, radiological and nuclear (CBRN) and cyber-offensive domains, in line with its Frontier Safety Framework, while 3.8 Flash Cyber has a more permissive set of mitigations but is restricted to trusted defenders who need broader cyber capabilities. The company also points to progress in robustness against prompt injection attacks, according to measurements by the firm Gray Swan.

As for availability, Gemini 3.8 Flash can already be used in Google Antigravity, the Gemini API via Google AI Studio, Android Studio and Stitch for developers; in Gemini Enterprise for companies; and by consumers subscribed to Google AI Pro and Ultra, through the Gemini app, Google Search's AI Mode and Gemini in Google Sheets. Access to Gemini 3.8 Flash Cyber, by contrast, is channeled exclusively through the Fairwind program, aimed at trusted government authorities, critical infrastructure operators and software maintainers, who must request access.

The release pace —three Flash models in six weeks— and the shared emphasis on agentic coding workflows and automated cybersecurity suggest that Google is trying to consolidate its position both in the race for autonomous software agents and in the application of AI to cyber defense, an area where the company says it has deliberately chosen to restrict access to the most sensitive capabilities.

🔗 Related on Zendoric

Sources & references