
On September 2, 2026, Google introduced two new models on the Gemini blog: Gemini 3.8 Flash, a general-purpose workhorse model, and Gemini 3.8 Flash Cyber, a cybersecurity-focused variant with restricted access. The announcement came from Tulsee Doshi, Senior Director of Product Management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind, on blog.google.
Three releases, six weeks
Google frames this as its third Flash-tier release in six weeks. The cadence started with Gemini 3.6 Flash on July 21, 2026, continued with Gemini 3.7 Flash on August 13, 2026, and now lands on 3.8 Flash and 3.8 Flash Cyber. Google describes 3.8 as building on the momentum of 3.7 Flash, released three weeks earlier, per the announcement.

Gemini 3.6 Flash launched alongside Gemini 3.5 Flash Cyber, the model that 3.8 Flash Cyber now replaces. The cyber-focused line has been iterating on the same six-week schedule as the general-purpose Flash models.
What changed in Gemini 3.8 Flash
Google positions 3.8 Flash as its most intelligent workhorse model yet, with gains in software engineering, agentic tasks, and multi-step reasoning over 3.7 Flash, according to the official announcement. On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash scores 73.7 percent, up from 65.3 percent for 3.7 Flash. That places it ahead of GPT-5.6 Sol, GPT-5.6 Terra, and Claude Sonnet 5 in Google's own comparison table, though Claude Opus 5 still leads the group at 74.0 percent.
The model also gained ground on general reasoning and agentic coding. Gemini 3.8 Flash scores 54.9 percent on HLE-Verified, a benchmark spanning STEM, humanities, and professional fields, edging out every model in the same table, including Claude Opus 5's 54.4 percent and GPT-5.6 Sol's 54.5 percent. On Terminal-bench 2.1, agentic terminal coding, 3.8 Flash posts 89.4 percent against 85.8 percent for 3.7 Flash.
Google also credits 3.8 Flash with beating every other model in its comparison table on two specialized agent benchmarks. On Vals Finance Agent V2, 3.8 Flash scores 61.4 percent against 59.0 percent for 3.7 Flash and 58.6 percent for the next-best model, Claude Opus 5. On Harvey's Legal Agent Benchmark, 3.8 Flash reaches a 10.0 percent full pass rate, roughly double GPT-5.6 Sol's 2.5 percent.
Google attributes these gains to the model working harder: taking extra reasoning steps and calling tools iteratively on complex tasks, at the cost of more tokens at higher effort settings. Developers who need lower latency can drop the effort level, or stay on 3.7 Flash, which remains fully supported.
Pricing stays where 3.7 Flash left it, for now. 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, per the pricing footnote on the announcement. Starting January 1, 2027, that rate doubles to $1.50 and $7.50 per million tokens, the same rate 3.6 Flash charged back in July.
How Gemini 3.8 Flash stacks up against rivals
Google's own announcement includes a head-to-head table against Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra across fifteen benchmarks. These are Google's self-reported numbers, so treat them the way you would any vendor's own comparison chart, pending independent replication.
Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
Input price, $ per 1M tokens | $0.75* | $0.75* | $5.00 | $2.00 | $4.00 | $2.00 |
Output price, $ per 1M tokens | $3.75* | $3.75* | $25.00 | $10.00 | $20.00 | $12.00 |
DeepSWE v1.1, long-horizon SWE | 73.7% | 65.3% | 74.0% | 53.8% | 72.7% | 69.6% |
GDPVal-AA v2, knowledge work (Elo) | 1545 | 1482 | 1824 | 1584 | 1710 | 1528 |
Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.9% | 53.8% | 54.4% |
Harvey's Legal Agent Benchmark, full pass | 10.0% | 8.8% | 6.7% | 5.0% | 2.5% | 0.8% |
Terminal-bench 2.1, agentic terminal coding | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
Terminal-bench 4.0, general agent | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
GDP.PDF, expert PDF comprehension, full pass | 35.0% | 34.0% | 37.0% | 28.0% | 40.0% | 29.0% |
CharXiv Reasoning, no tools | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
LVBench, agentic / static, long video | 87.8% / 87.1% | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
HLE-Verified, multidisciplinary reasoning | 54.9% | 53.6% | 54.4% | 31.0% | 54.5% | 51.1% |
OSWorld-2.0, agentic computer use, partial | 59.0% | 50.6% | 75.4% | 42.6% | 62.6% | 50.2% |
BioMysteryBench, human-solvable | 88.8% | 87.1% | 90.1% | 87.5% | 79.5% | 83.8% |
BioMysteryBench, human-difficult | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
LABBench2, biology research tasks | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |
* Introductory price through Dec 31, 2026. Regular price is $1.50 input and $7.50 output per 1M tokens.
Source: Google's Gemini 3.8 Flash announcement; methodology listed at deepmind.google/models/evals-methodology/gemini-3-8-flash.
Claude Opus 5 leads on five of the fourteen benchmark rows: DeepSWE v1.1, GDPVal-AA v2, Terminal-bench 4.0, OSWorld-2.0, and the human-solvable split of BioMysteryBench. GPT-5.6 Sol leads on one, GDP.PDF. Gemini 3.8 Flash leads the remaining eight rows, and undercuts every rival on both price columns by a wide margin.
What changed in Gemini 3.8 Flash Cyber, and who can use it
Gemini 3.8 Flash Cyber is not a public model. Access runs through the newly launched Fairwind Program, which Google says prioritizes trusted government authorities, critical infrastructure operators, and software maintainers.
On CyberGym, the standard industry benchmark for vulnerability discovery, Google reports that 3.8 Flash Cyber delivers frontier-level performance, surpassing both 3.5 Flash Cyber and larger frontier models. Because CyberGym leans heavily on C and C++ codebases, Google also ran an internal benchmark spanning 20 programming languages, where 3.8 Flash Cyber crossed a 70 percent vulnerability discovery success rate.
Patching is where the numbers get most specific. On CWE-Bench, an external benchmark run by Collinear, 3.8 Flash Cyber posted a pass@1 of 47.2 percent, just behind a leading frontier model's 47.8 percent, but at meaningfully lower cost per rollout, per Google's writeup of the result. Google calls this landing on the Pareto frontier for the benchmark.
According to the primary announcement, real-world results reinforce the benchmark story. Google's own Chrome Security team found 3.8 Flash Cyber produced 2.6 times more correct vulnerability patches than larger commercial models. Wiz reported 7.5 to 9.7 percent higher recall on its internal penetration-testing benchmark, at 2.3 to 5.2 times lower cost than other leading frontier models. Google's Cloud Vulnerability Research team used the model to find a critical foundational vulnerability in under two hours, a process that normally takes months.
One core model, two access envelopes
Both variants share the same foundational intelligence, refined through what Google calls long-running agentic loops that recursively evaluate and improve the underlying model. What separates them is the safety envelope wrapped around each release channel, not the underlying architecture.
flowchart TD
A[Shared foundational model, refined via agentic training loops] --> B[Gemini 3.8 Flash]
A --> C[Gemini 3.8 Flash Cyber]
B --> D[Developers: Antigravity, AI Studio, Android Studio, Stitch]
B --> E[Enterprises: Gemini Enterprise]
B --> F[Consumers: Gemini app, AI Mode, Sheets, Pro/Ultra only]
C --> G[Fairwind Program: application required]
G --> H[Governments, critical infrastructure operators, software maintainers]
3.8 Flash ships broadly across developer, enterprise, and consumer surfaces. 3.8 Flash Cyber ships to a narrow, vetted list of defenders, reflecting the dual-use risk of a model built to find and patch vulnerabilities at scale.
Side by side
Metric | Gemini 3.7 Flash | Gemini 3.8 Flash | Gemini 3.8 Flash Cyber |
|---|---|---|---|
Release date | Aug 13, 2026 | Sep 2, 2026 | Sep 2, 2026 |
Access | General availability | General availability | Fairwind Program only |
Price, input / output per 1M tokens | $0.75 / $3.75 | $0.75 / $3.75 through Dec 31, 2026 | Not publicly listed |
DeepSWE v1.1 | 65.3% | 73.7% | N/A |
Terminal-bench 2.1 | 85.8% | 89.4% | N/A |
HLE-Verified | 53.6% | 54.9% | N/A |
CyberGym, vulnerability discovery | N/A | N/A | Frontier-level, beats 3.5 Flash Cyber |
CWE-Bench patching, pass@1 | N/A | N/A | 47.2%, vs 47.8% for top frontier model |
Source: blog.google.
Safety notes worth reading
3.8 Flash ships with safeguards against misuse in chemical, biological, radiological, and nuclear domains, plus cyber offense, under Google's Frontier Safety Framework, per the announcement. 3.8 Flash Cyber carries a more permissive set of mitigations, which is why it stays behind the Fairwind Program gate rather than shipping broadly.
Google also reports a significant jump in prompt injection robustness for the 3.8 generation, measured against Gray Swan's indirect prompt injection benchmark. The company frames this as protection for everyday Gemini users against prompt-injection attacks, not only a cyber-specialist concern.
Where to find it
Developers reach 3.8 Flash through Google Antigravity, Google AI Studio, and Android Studio, or generate interfaces with it in Stitch. Enterprises get access through Gemini Enterprise. Consumers on a Google AI Pro or Ultra plan can reach it inside the Gemini app, AI Mode in Search, or Gemini in Sheets.
3.8 Flash Cyber only reaches people who apply through the Fairwind Program and get accepted as a trusted defender. Google lists early Fairwind partners including Armadin, Palo Alto Networks, Snowflake, and Wiz in the announcement, though it has not published the program's exact acceptance criteria.
The pattern to watch
Three Flash-tier releases in six weeks is a fast cadence for any model family, and it says as much about Google's release strategy as about any single model. If the pattern holds, 3.8 Flash will not be the last word here for long. Developers building on Gemini should treat these version numbers as a moving target and budget review time whenever a new one drops.
Comments (0)
Login to post a comment.