Google shipped a cyber model with deliberately relaxed safeguards. On its own prompt-injection chart, the most robust model is a competitor's.
The second lab in twenty-four hours to loosen the same control, and the first to make the compensating control an access list rather than a capability limit.
By Parminder Kumar Sharma · · 8 min read

Google announced Gemini 3.8 Flash Cyber on 2 September. The sentence that matters is in the safety section, and it is stated plainly:
"3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities."
Read that against the line immediately before it, describing the standard model: 3.8 Flash "ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense".
So this is a model shipped with one of those safeguard classes deliberately loosened. That is the second frontier lab in twenty-four hours to do it, and the two chose opposite compensating controls.
Two labs, one week, two different answers
Anthropic's Fable 5.1 announcement on 1 September said its model "can now be used to discover software vulnerabilities, though not to develop exploits for them". The control there is a capability boundary: find but not exploit, and the model ships generally available.
Google's control is an access gate. There is no stated capability line at all. Instead there is the Fairwind Program, where participating organisations "agree to strict operational standards, including limiting access to employees within their internal cybersecurity, incident response, or penetration testing teams and deploying protections like multi-factor authentication".
Both are defensible. They are not the same bet, and the difference is worth naming because it determines who carries the risk. A capability boundary is enforced by the vendor and fails closed for everybody. An access gate is enforced by the customer's own joiners and leavers process, and it fails wherever that process fails.
If you are in Fairwind, the eligibility criteria are your control. They are stated in prose, in a blog post, and Google does not publish how they are audited.
The prompt-injection claim, and Google's own chart
Google writes that "Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan".
They published the chart. On it, sixteen models are plotted on an axis running to 80%, because DeepSeek V4 Pro reaches 60.1%. Seven models cluster between 4.8% and 9.2%, which at that scale renders as seven near-identical stubs.
The same Gray Swan numbers, at a readable scale
At a scale where the cluster is legible, the ordering is visible, and the most robust model in Google's own set is Claude Opus 5 at 4.8%. Gemini 3.8 Flash is second at 5.5%. Gemini 3.8 Flash Cyber, the model being announced, is third at 6.0%.
Two things follow, and they pull in opposite directions.
Google published a chart on which a competitor wins. That is more than most vendors do and it should be said before anything else.
And the prose above it says Gemini made "a significant leap", which is true generation on generation, 9.2% down to 5.5%, and is not a claim of leadership. A reader who takes "significant leap in prompt injection robustness" as "most robust" has been misled by the axis rather than by the sentence.
The two numbers that do not survive their own charts
CWE-Bench. Google reports Gemini 3.8 Flash Cyber "is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost".
That is an honest sentence describing a loss. 47.2 is behind 47.8. The Pareto framing is legitimate, because cost is a real axis and being close at much lower price is a genuine result. But the headline patching figure is not a win, and coverage that reports it as one has dropped the comparison Google supplied.
CyberGym. The text says the model "demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models." No number appears in the prose.
The number is on the chart. Gemini 3.8 Flash Cyber reaches 86.2%. The next model, GPT-5.5-Cyber, reaches 85.6%. That is a margin of six tenths of a percentage point, against models Google describes as significantly larger.
Six tenths may well be real. It is also the kind of margin that needs a confidence interval before it can carry the word "surpasses", and none is given.
What each figure is, and where it comes from
| The claim | Where it is stated | What the chart shows |
|---|---|---|
| “A significant leap in prompt injection robustness” | Prose | True generation on generation, 9.2% to 5.5%. The most robust model in the set is Claude Opus 5 at 4.8% |
| “On the Pareto frontier” on CWE-Bench | Prose, with both figures given | 47.2% against 47.8% for a leading frontier model. A loss on accuracy, at lower cost |
| “Surpasses significantly larger frontier models” on CyberGym | Prose, with no figure | 86.2% against 85.6%. A margin of 0.6 points, with no interval stated |
| “Success rate exceeding 70%” on real-world discovery | Prose | 71.0%, on Google’s own internal benchmark across 20 languages. No external comparator appears |
| 2.6x more correct Chrome patches | Prose, credited to Google’s Chrome Security team | Not charted. An internal result, against unnamed “best commercial models” |
| +7.5 to 9.7% higher recall at 2.3 to 5.2x lower cost | Prose, credited to Wiz | Not charted. The only externally attributed result in the announcement |
The Wiz line is worth pulling out of that table. It is the one result in the whole announcement attributed to a named external party running their own benchmark, and it is expressed as a range rather than a point. That is the most credible number Google published, and it is the least quoted.
What a defender should take from this
The capability claims are plausible and mostly self-measured, which is normal and not a criticism. What changes your posture is narrower.
There is now a model, in limited release, explicitly shipped with cyber-offence mitigations loosened, whose access control is a partner programme rather than a capability boundary. That is a new category of thing to have an opinion about, and if your organisation is in Fairwind then your joiners and leavers process is now part of a frontier lab's safety architecture.
The second thing is the pattern. Two labs relaxed the same class of control within a day of each other, and both said so in public. Neither hid it. The direction of travel is now established rather than speculative, and an AI usage policy written on the assumption that vendors tighten cyber controls over time is written against the wrong trend.
Take this with you
If this is in scope for you
- Establish whether your organisation is in the Fairwind Program, and if so, who inside it holds access. The eligibility rule limits use to internal cyber, incident response and penetration testing staff, which makes your access review the enforcement mechanism.
- Do not quote the CWE-Bench figure as a win. Google gives both numbers: 47.2% against 47.8% for a leading frontier model. The claim is Pareto optimality at lower cost, which is a different and narrower claim.
- Treat the CyberGym result as a tie until an interval is published. 86.2% against 85.6% is six tenths of a point, and the word in the prose is “surpasses”.
- Note which results are external. The Wiz recall range is the only figure attributed to a named third party running its own benchmark; the 71% and the Chrome 2.6x are internal.
- Update any AI policy that assumes cyber safeguards tighten over time. Two frontier labs loosened the same class of control within twenty-four hours and both documented it.
- Ask any vendor claiming prompt-injection robustness for the axis, not the adjective. On Google’s own chart the leader is a competitor, and that is invisible at the scale they plotted.
The position
The most useful thing about this announcement is how much of it is checkable, and that is not faint praise. Google published the competitor comparison on CWE-Bench with both numbers in the sentence. It published a Gray Swan chart that a rival wins. It named Wiz and gave a range rather than a point.
The gap is between the prose and the images. Three of the most quotable sentences are qualified only by figures that appear nowhere except inside a chart, and two of those charts are drawn at a scale where the qualification is invisible.
That is not deception. It is the ordinary result of a launch post being written to be read quickly. The correction is equally ordinary: when a vendor gives you a chart, read the chart.
Sources
- PrimaryGemini 3.8 Flash and 3.8 Flash Cyber, 2 September 2026. The permissive-mitigations sentence, the CWE-Bench comparison and the published chartsGoogleaccessed 2026-09-03
- PrimaryThe Fairwind Program, its eligibility conditions and the figure of more than 650 participating partnersGoogleaccessed 2026-09-03
- PrimaryClaude Fable 5.1, 1 September 2026, the capability-boundary comparison: discover vulnerabilities but not develop exploitsAnthropicaccessed 2026-09-03


