Gemini 3.8 Flash and Cyber arrive for agents and security

What is Gemini 3.8 Flash
Google launched the Gemini 3.8 Flash and the Gemini 3.8 Flash Cyber variant on 02/09/2026, according to the official announcement on The Keyword. The introductory price is the same as 3.7 Flash, US$ 0.75 and US$ 3.75 per million tokens, and it expires on 31/12/2026, doubling on 01/01/2027.
Published on 10/09/2026. The event is from 02/09/2026; prices, deadlines, and benchmarks follow the primary source's reading as of this date.
Eight days separate the announcement from this analysis, and the distance is deliberate. A launch story is spent in a day; what remains is the decision the launch forces. The decision this announcement leaves has three fronts, each with a published date or number: the deadline on the introductory price, the effort dial that changes cost per task, and the restricted access to the Cyber variant. Each one changes the spreadsheet of anyone consuming models through an API.
What happened
The announcement of 02/09/2026, signed by Tulsee Doshi and Raluca Ada Popa of Google DeepMind, presents two models. Gemini 3.8 Flash is the workhorse model for software engineering, agentic tasks, and multi-step reasoning. Gemini 3.8 Flash Cyber is the cybersecurity variant, with access restricted to the Fairwind program.
It went out on The Keyword at 15:00 UTC on 02/09/2026. One post, two models, two access models: one open through the usual API channels, the other closed inside a trust program. The asymmetry between the two is the most important routing fact in the announcement.
Gemini 3.8 Flash is available on the channels Google lists: Gemini API with AI Studio, Google Antigravity, Android Studio, Stitch, and Gemini Enterprise, plus the Gemini apps for Pro and Ultra subscribers. The list is a fact of the 02/09 event. What each company does with it, where the model enters the mix, and on which route, is its own decision.
The Cyber variant does not enter that open circuit. According to the model page on Google DeepMind and the official post on the Fairwind program, access is exclusive to Fairwind, intended for trusted government authorities, critical infrastructure operators, and software maintainers. The announcement records that Cyber operates with more permissive mitigations, made possible by the restricted access. The primary source does not publish a public price for Cyber.
The cadence is also a fact: this is the third release of the Flash line in six weeks. 3.7 Flash arrived three weeks earlier and remains supported for efficiency-first workloads, according to the same announcement. This is not the first Gemini event covered here; the previous one, on agentic video and token consumption, is in the analysis of agentic video in Gemini. Anyone receiving a new model every three weeks needs an evaluation process, and the process starts with cost.
Why it matters
For the CFO, the introductory price has an expiration date: it expires on 31/12/2026 and doubles on 01/01/2027. For the CTO, cost per task gained an effort dial, and token consumption now varies with configuration. For the CEO, the frontier security route became a question of access, not of model choice.
Gemini 3.8 Flash enters at US$ 0.75 per million input tokens and US$ 3.75 per million output tokens, the same introductory price as 3.7 Flash. The official footnote of the announcement closes the deadline: on 31/12/2026 the introductory tier expires, and from 01/01/2027 US$ 1.50 and US$ 7.50 apply. The January number is double the September number.
The direct arithmetic of the published prices: a workload of 100 million input tokens and 20 million output tokens per month pays US$ 150 per month under the introductory tier and US$ 300 from 01/01/2027. It is rare for a vendor to hand over the date of the increase together with the low price. Here the date came in the footnote, and the 2027 cost line of the budget already has a number. Price as a routing argument has its own chapter on this blog: the cost of AI models as a routing argument.
The promised capability comes with vendor numbers. On DeepSWE v1.1, a long-horizon software engineering benchmark, Google states that 3.8 Flash beats most larger frontier models at a fraction of the cost. The model scores 54.9% on HLE-Verified and appears on the Vals Finance Agent V2 and Harvey's Legal Agent Benchmark agentic leaderboards. All of it from the 02/09/2026 announcement, not from independent evaluation; these figures enter the decision as a thesis to validate, not as settled fact.
The effort dial is the operational change. On complex tasks, the model runs extra reasoning steps and calls tools iteratively, and at the higher effort levels Google warns that token consumption rises; the lower levels minimize overhead. Cost per task stops being a passive derivative of the per-token price and becomes a configuration. Whoever measures cost per task already measures what matters; whoever measures only the token price will be measuring the wrong variable in 2027. For the reader arriving at the topic now, the LLM gateway and AI router explainer covers the role of the layer that decides which model answers each call.
On the security side, the general 3.8 Flash brings a gain in prompt injection resilience measured by Gray Swan, plus CBRN and cyber offense safeguards under the Frontier Safety Framework. The same package in the announcement reports production impact: Chrome's security team recorded 2.6x more correct patches than the best much larger commercial models, Wiz measured from +7.5% to 9.7% recall on an internal pentest benchmark at 2.3x to 5.2x lower cost, and the Cloud Vulnerability Research team found a critical vulnerability in under 2 hours, against months in the usual process. Vendor numbers, in cases the vendor chose. The detail that changes the routing decision is in the access.
The model behind that security front has no public price and no open channel. The Cyber variant is exclusive to the Fairwind program, for trusted government authorities, critical infrastructure operators, and software maintainers, and the primary source gives no expansion criteria and no deadline. Frontier security stopped being a model choice and became also an access question. Whoever fits the profile has one route to evaluate; whoever does not routes with the general 3.8 Flash and with the controls of the company's own gateway layer, the subject of the enterprise gateway for routing and security.
What changes in practice
The decision regime changes on four lines: the price stays at US$ 0.75 and US$ 3.75, but now with a published deadline; cost per task gains the effort dial; access to the Cyber variant goes through Fairwind; and agent resilience now has a disclosed metric, with Gray Swan measuring prompt injection.
The table compares the 3.7 regime with the 3.8 regime on the four fronts that change the decision spreadsheet.
| Decision | 3.7 regime (before) | 3.8 regime (after) |
|---|---|---|
| Price and validity | US$ 0.75 and US$ 3.75 per million tokens, introductory price in force | Same price, now with a deadline: expires on 31/12/2026 and doubles to US$ 1.50 and US$ 7.50 on 01/01/2027 |
| Cost per task | Read as price per token, the cost variable of the line's announcement | Effort dial: high levels consume more tokens, with extra reasoning and iterated tool calls; lower levels minimize overhead |
| Cyber availability | Previous generation (3.5 Flash Cyber) cited in a benchmark in the announcement, with no named access program in this piece | Exclusive to Fairwind: trusted government authorities, critical infrastructure operators, and software maintainers |
| Agent resilience | No disclosed prompt injection metric in the previous generation | Significant gain measured by Gray Swan, with CBRN and cyber offense safeguards under the Frontier Safety Framework |
The price is the line that bites first. The introductory regime ends on 31/12/2026, and the announcement's footnote fixes the 2027 tier: US$ 1.50 per million input tokens and US$ 7.50 per million output tokens.
It doubles on both sides.
Whoever builds architecture and contracts on top of the September price will reopen the spreadsheet in January. The difference is that the deadline was written from day one, and a price with a published validity is a contract condition, not a billing surprise.
The table has a limit worth naming: the benchmarks that support the capability line are the vendor's. DeepSWE v1.1, CyberGym, and CWE-Bench measure what Google says the models do, per the reading of the 02/09/2026 announcement. Independent validation comes later. The router that waits for it does not miss the price deadline; the deadline belongs to the calendar, not the leaderboard.
What to do now
Five decisions concentrate the practical value of the launch: anchor the 2027 budget on the January price, configure the effort level per workload, decide the organization's position on Fairwind, review the workloads that stay on 3.7 Flash, and validate the vendor benchmarks with real load. The list, in order of deadline.
- Anchor the 2027 budget on the 01/01/2027 price. US$ 1.50 and US$ 7.50 per million tokens are the numbers that hold in planning; the introductory tier of US$ 0.75 and US$ 3.75 is a window until 31/12/2026. A workload of US$ 150 per month under the current regime costs US$ 300 in January, by the arithmetic of the published prices.
- Set the effort level per workload, not per account. High levels pay off on complex agentic tasks, where the model iterates tools and runs extra reasoning; lower levels minimize overhead where the bill is dominated by volume. This configuration is the new variable in cost per task.
- Classify the organization against Fairwind. Trusted government authorities, critical infrastructure operators, and software maintainers have a security route to evaluate with the Cyber variant. Outside that profile, the route is the general 3.8 Flash with the company's own gateway controls.
- Review the efficiency-first workloads on 3.7 Flash. The model remains supported, according to the announcement, and the migration decision goes through the effort dial and the price deadline, not through the version number. The primary source gives no end-of-support date.
- Treat the benchmarks as a thesis to validate. DeepSWE v1.1, CyberGym, and CWE-Bench are vendor numbers from the 02/09/2026 announcement. The pilot with the company's real workload is what converts a leaderboard into a decision.
Frequently asked questions about Gemini 3.8 Flash
How much does Gemini 3.8 Flash cost?
US$ 0.75 per million input tokens and US$ 3.75 per million output tokens, the same introductory price as 3.7 Flash, according to the announcement of 02/09/2026. The official footnote fixes the deadline: the price expires on 31/12/2026, and from 01/01/2027 US$ 1.50 and US$ 7.50 apply.
What is Gemini 3.8 Flash Cyber?
It is the cybersecurity variant of 3.8 Flash, with frontier performance in autonomous vulnerability discovery and automated patching. The announcement cites 47.2% pass@1 on CWE-Bench, against 47.8% for a leading frontier model at a significantly lower cost, and a success rate above 70% on the internal benchmark across 20 languages.
Is Gemini 3.8 Flash Cyber available to any company?
No. At the time of publication (10/09/2026), access is exclusive to the Fairwind program, intended for trusted government authorities, critical infrastructure operators, and software maintainers. The primary source gives no expansion criteria and no date for broader access.
Is 3.7 Flash still available?
Yes. According to the announcement of 02/09/2026, 3.7 Flash remains supported for efficiency-first workloads. Google did not give an end-of-support date in the material cited in this piece.
What changes in the price in 2027?
It doubles. From 01/01/2027, Gemini 3.8 Flash costs US$ 1.50 per million input tokens and US$ 7.50 per million output tokens, double the introductory tier, according to the footnote of the announcement of 02/09/2026.
References and further reading
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, Tulsee Doshi and Raluca Ada Popa, The Keyword (Google), 02/09/2026
- Gemini 3.8 Flash Cyber, model page, Google DeepMind
- Fairwind program, official program page
- Fairwind program, official post on the Google blog
- LLM gateway and AI router explainer
- Enterprise gateway: routing and security
- The cost of AI models as a routing argument
- The analysis of agentic video in Gemini
What to watch
The calendar is in charge. The introductory price expires on 31/12/2026, the doubled tariff applies from 01/01/2027, and between the two dates sits the decision of which route Gemini 3.8 Flash occupies in the mix. The absence that remains unanswered is Fairwind's: no deadline and no expansion criteria at the time of publication.
Two signals are worth following. The first is the cadence of the Flash line: three releases in six weeks indicate that the evaluation window for a Flash model is shorter than the quarterly review cycle of many companies, and the next generation will reopen the spreadsheet again. The second is independent validation of the announcement's benchmarks; until it arrives, Google's numbers hold as a thesis.
For anyone routing models, the launch changes three variables of the decision, and none of them is which model is better. The price has a horizon with a date. The cost per task has a dial. The security route has a door, and the one holding the key is Fairwind. A tariff policy with a deadline, an effort configuration, and an access rule are exactly the kind of parameter that a routing layer turns into an automatic decision, and that layer is where the Nexforce Router operates. The restricted access of Cyber and the price that doubles in January are not catalog changes; they are policy changes, and policy is what stays under the control of whoever routes.

Accelerate your company'sbusiness and operational efficiency
We design the technology of tomorrow to boost your business operational scale
Talk to a SpecialistRelated articles

GPT-6 Astra: pricing, benchmarks and OpenAI safety rating
OpenAI launched GPT-6 Astra at US$ 10/US$ 50 per million tokens, with a Critical cybersecurity rating and worse monitorability than its predecessor. The piece shows what that combination changes in routing decisions for anyone consuming the API in production.
Read more
ChatGPT Ads reaches $1B run rate and launches global self-serve
OpenAI reaches $1 billion annualized advertising revenue run rate on ChatGPT within 200 days and unlocks global self-serve ad buying across 40+ countries.
Read more
Quasar 438B: the European 438B model enters the routing leaderboard
Multiverse Computing launched Quasar 438B, a 438B-class reasoning model in English and Spanish that scores 43 on Intelligence Index v4.1.1 and returns 500 tokens with thinking in 15.3 seconds. For anyone routing models, Europe now has a flash-latency reasoning route with long context near the frontier.
Read more