P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Google shipped a cyber model with deliberately relaxed safeguards. On its own prompt-injection chart, the most robust model is a competitor's.

The second lab in twenty-four hours to loosen the same control, and the first to make the compensating control an access list rather than a capability limit.

By Parminder Kumar Sharma · · 8 min read

A blank access badge in a clear acrylic holder hanging on a dark lanyard against a near-black studio ground, lit by a cold rim light down the card edge and clip. The card face is entirely unprinted.

Google announced Gemini 3.8 Flash Cyber on 2 September. The sentence that matters is in the safety section, and it is stated plainly:

"3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities."

Read that against the line immediately before it, describing the standard model: 3.8 Flash "ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense".

So this is a model shipped with one of those safeguard classes deliberately loosened. That is the second frontier lab in twenty-four hours to do it, and the two chose opposite compensating controls.

Two labs, one week, two different answers

Anthropic's Fable 5.1 announcement on 1 September said its model "can now be used to discover software vulnerabilities, though not to develop exploits for them". The control there is a capability boundary: find but not exploit, and the model ships generally available.

Google's control is an access gate. There is no stated capability line at all. Instead there is the Fairwind Program, where participating organisations "agree to strict operational standards, including limiting access to employees within their internal cybersecurity, incident response, or penetration testing teams and deploying protections like multi-factor authentication".

Both are defensible. They are not the same bet, and the difference is worth naming because it determines who carries the risk. A capability boundary is enforced by the vendor and fails closed for everybody. An access gate is enforced by the customer's own joiners and leavers process, and it fails wherever that process fails.

If you are in Fairwind, the eligibility criteria are your control. They are stated in prose, in a blog post, and Google does not publish how they are audited.

The prompt-injection claim, and Google's own chart

Google writes that "Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan".

They published the chart. On it, sixteen models are plotted on an axis running to 80%, because DeepSeek V4 Pro reaches 60.1%. Seven models cluster between 4.8% and 9.2%, which at that scale renders as seven near-identical stubs.

The same Gray Swan numbers, at a readable scale

The same Gray Swan numbers, at a readable scaleGoogle plotted these on a 0 to 80% axis. Here is 0 to 10%.Attack success rate within 15 attempts, lower is better0%2%4%6%8%10%Claude Opus 54.8%Gemini 3.8 Flash5.5%Gemini 3.8 Flash Cyber6.0%Claude Fable 56.5%Claude Sonnet 56.7%Claude Opus 4.88.0%Gemini 3.7 Flash9.2%best in the set, and it is not a GeminiFigures read from Google’s published Gray Swan chart, 2 September 2026. Blue marks Google’s own models, following their colour convention.
Google published a chart on which a competitor wins, which is more than most vendors do and deserves saying. The observation here is only about the axis: with one model at 60.1% setting the scale, the seven that actually compete are compressed into a centimetre, and the ordering among them is the part a reader would want.
Figures read from Google's published chart of 2 September 2026, redrawn on a 0 to 10% axis. Blue marks Google's own models, following their colour convention. Attack success rate within 15 attempts, lower is better.

At a scale where the cluster is legible, the ordering is visible, and the most robust model in Google's own set is Claude Opus 5 at 4.8%. Gemini 3.8 Flash is second at 5.5%. Gemini 3.8 Flash Cyber, the model being announced, is third at 6.0%.

Two things follow, and they pull in opposite directions.

Google published a chart on which a competitor wins. That is more than most vendors do and it should be said before anything else.

And the prose above it says Gemini made "a significant leap", which is true generation on generation, 9.2% down to 5.5%, and is not a claim of leadership. A reader who takes "significant leap in prompt injection robustness" as "most robust" has been misled by the axis rather than by the sentence.

The two numbers that do not survive their own charts

CWE-Bench. Google reports Gemini 3.8 Flash Cyber "is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost".

That is an honest sentence describing a loss. 47.2 is behind 47.8. The Pareto framing is legitimate, because cost is a real axis and being close at much lower price is a genuine result. But the headline patching figure is not a win, and coverage that reports it as one has dropped the comparison Google supplied.

CyberGym. The text says the model "demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models." No number appears in the prose.

The number is on the chart. Gemini 3.8 Flash Cyber reaches 86.2%. The next model, GPT-5.5-Cyber, reaches 85.6%. That is a margin of six tenths of a percentage point, against models Google describes as significantly larger.

Six tenths may well be real. It is also the kind of margin that needs a confidence interval before it can carry the word "surpasses", and none is given.

What each figure is, and where it comes from

The claimWhere it is statedWhat the chart shows
“A significant leap in prompt injection robustness”ProseTrue generation on generation, 9.2% to 5.5%. The most robust model in the set is Claude Opus 5 at 4.8%
“On the Pareto frontier” on CWE-BenchProse, with both figures given47.2% against 47.8% for a leading frontier model. A loss on accuracy, at lower cost
“Surpasses significantly larger frontier models” on CyberGymProse, with no figure86.2% against 85.6%. A margin of 0.6 points, with no interval stated
“Success rate exceeding 70%” on real-world discoveryProse71.0%, on Google’s own internal benchmark across 20 languages. No external comparator appears
2.6x more correct Chrome patchesProse, credited to Google’s Chrome Security teamNot charted. An internal result, against unnamed “best commercial models”
+7.5 to 9.7% higher recall at 2.3 to 5.2x lower costProse, credited to WizNot charted. The only externally attributed result in the announcement
Prose quotations from Google's announcement of 2 September 2026, read in full. Chart figures read from the charts published alongside it. The distinction between the two matters here, because three of the most quotable claims are in the prose while the numbers that qualify them are only in the images.

The Wiz line is worth pulling out of that table. It is the one result in the whole announcement attributed to a named external party running their own benchmark, and it is expressed as a range rather than a point. That is the most credible number Google published, and it is the least quoted.

What a defender should take from this

The capability claims are plausible and mostly self-measured, which is normal and not a criticism. What changes your posture is narrower.

There is now a model, in limited release, explicitly shipped with cyber-offence mitigations loosened, whose access control is a partner programme rather than a capability boundary. That is a new category of thing to have an opinion about, and if your organisation is in Fairwind then your joiners and leavers process is now part of a frontier lab's safety architecture.

The second thing is the pattern. Two labs relaxed the same class of control within a day of each other, and both said so in public. Neither hid it. The direction of travel is now established rather than speculative, and an AI usage policy written on the assumption that vendors tighten cyber controls over time is written against the wrong trend.

Take this with you

If this is in scope for you

  • Establish whether your organisation is in the Fairwind Program, and if so, who inside it holds access. The eligibility rule limits use to internal cyber, incident response and penetration testing staff, which makes your access review the enforcement mechanism.
  • Do not quote the CWE-Bench figure as a win. Google gives both numbers: 47.2% against 47.8% for a leading frontier model. The claim is Pareto optimality at lower cost, which is a different and narrower claim.
  • Treat the CyberGym result as a tie until an interval is published. 86.2% against 85.6% is six tenths of a point, and the word in the prose is “surpasses”.
  • Note which results are external. The Wiz recall range is the only figure attributed to a named third party running its own benchmark; the 71% and the Chrome 2.6x are internal.
  • Update any AI policy that assumes cyber safeguards tighten over time. Two frontier labs loosened the same class of control within twenty-four hours and both documented it.
  • Ask any vendor claiming prompt-injection robustness for the axis, not the adjective. On Google’s own chart the leader is a competitor, and that is invisible at the scale they plotted.

The position

The most useful thing about this announcement is how much of it is checkable, and that is not faint praise. Google published the competitor comparison on CWE-Bench with both numbers in the sentence. It published a Gray Swan chart that a rival wins. It named Wiz and gave a range rather than a point.

The gap is between the prose and the images. Three of the most quotable sentences are qualified only by figures that appear nowhere except inside a chart, and two of those charts are drawn at a scale where the qualification is invisible.

That is not deception. It is the ordinary result of a launch post being written to be read quickly. The correction is equally ordinary: when a vendor gives you a chart, read the chart.

Sources

  1. PrimaryGemini 3.8 Flash and 3.8 Flash Cyber, 2 September 2026. The permissive-mitigations sentence, the CWE-Bench comparison and the published chartsGoogleaccessed 2026-09-03
  2. PrimaryThe Fairwind Program, its eligibility conditions and the figure of more than 650 participating partnersGoogleaccessed 2026-09-03
  3. PrimaryClaude Fable 5.1, 1 September 2026, the capability-boundary comparison: discover vulnerabilities but not develop exploitsAnthropicaccessed 2026-09-03

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.