P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Gemini 4 Argon leads or ties 13 of 18 benchmarks in Google's table, and Google ran 9 of the 18 itself

Google's table puts Gemini 4 Argon ahead of or level with GPT-6 Astra and two Claude models on 13 of 18 benchmarks. Independent sources confirm the Vals lead and show the CWE-bench tie costs $6.63 a rollout, while only vetted defenders can use the model.

By Parminder Kumar Sharma · · 24 min read

Google's launch key art for Gemini 4 Argon: the product name beside a four-pointed star logo on the left, and a large white numeral 4 on the right, over a blue gradient.

Led or tied on 13 of 18, with Google scoring 9 of the 18 itself

Google's comparison table for Gemini 4 Argon, published on 30 September 2026, sets it against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 18 benchmarks (19 printed rows, because GraphWalks appears at two context lengths). By our count of the printed figures Argon leads outright on 12, ties for first on one and trails on five. Google's methodology PDF says it computed Argon's own score on nine of the 18. The other nine come from third-party leaderboards. The announcement text names seven of the 13 benchmarks Argon leads or ties, and none of the five it trails: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0 appear only in the table image.

What that does not establish. It does not make the table wrong. Vals, whose leaderboards supply four of those nine rows, lists Argon first at 68.90%, and on the five benchmarks it trails Google printed the numbers. It does not show that the leads carry over to any team's own work: the table has no sample sizes, no error bars and no cost per task, and Google's own notes say that on several rows the rivals were run under different settings from Argon. And three things are absent from the announcement: a model card, a statement of whether Argon reached a critical capability level for cyber under Google's Frontier Safety Framework, and a date on which most people can use the model.

The launch post uses the vocabulary of defence: trusted defenders, frontier safeguards, a programme called Fairwind. The record is a model that, on the one security benchmark with a public leaderboard, ties two rivals at more than twice their cost per rollout, and that only organisations Google has vetted can use. OpenAI's system card, by contrast, rates GPT-6 Astra and GPT-6.1 Sol Critical for cyber, as our briefing on dots and Sol set out. Google's announcement is silent on the equivalent rating.

Six claims in Google's announcement, and what a second source shows. Sources: Google, 30 September 2026; Vals; cwe-bench.com; Wiz; Artificial Analysis; all read on 30 September 2026.

Google's claimWhat a second source showsStatus
Leading on the Vals Index, 68.9%Vals lists 68.90% at rank 1. Claude Sonnet 5.5, absent from Google's table, is second at 67.04%. Vals gives a standard error of about 0.9 for its two neighboursConfirmed. The lead is 1.9 points
Ties for first on CWE-bench v1, 68%A three-way tie with Grok 4.7 and GPT-6 Astra. Argon costs $6.63 a rollout against $0.79 for Opus 5.5, and is fourth of five on the judge panelTie confirmed. Cost and panel score are not in Google's chart
Most resilient model yet to indirect prompt injection, 0.7% on Gray Swan IPIGray Swan's leaderboard returned an error when we fetched it. The next models on Google's chart are at 1.0%Not checked by us
Wiz found a critical healthcare flaw that earlier models missedWiz's Scan for Good page does not mention Gemini or Argon. No flaw, vendor or CVE is namedNot checkable
Argon can autonomously find, validate and patch critical vulnerabilitiesPatching is measured: 68% of CWE-bench tasks pass on a first attempt. Finding is measured only on two internal benchmarks, against Argon's predecessorPartly measured
Released without cyber guardrails to trusted defendersFairwind lists 650+ partners and its conditions. Neither page says how many hold Argon or what the guardrails doStated, not described

What Google announced, and who can use it

Argon is Google's new frontier model, and on the day of the announcement it is not generally available. Google says it is rolling out to a set of trusted cyber defenders through the Fairwind Program, and that it is engaged in the US government's voluntary process for pre-release model access while it gradually expands access. Wider release to developers, enterprises and consumers comes "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. No date is given for any step.

The numbers that are given are these. The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. Footnote 1 says that after the introductory period the price becomes $4 and $20; the end of that period is not dated. The maximum output per response rises to 1 million tokens, from 64,000, about 15 times (derived). The announcement does not state the input context window; Artificial Analysis lists 1 million tokens, with text and image input and text output.

Two sentences carry most of the security weight. Google says it trained Argon to be highly capable at cyber defence and that it can "autonomously find, validate, and patch critical software vulnerabilities". And for trusted defenders and its own teams it says it will release Argon "without cyber guardrails". The Fairwind page says a set of its partners get exclusive access to Argon, and that the programme works with more than 650 partners in all. It does not say how large that set is.

Three columns joined by arrows. Google's own teams use Argon without cyber guardrails. A subset of Fairwind partners, vetted, get Argon without cyber guardrails under conditions including phishing-resistant multi-factor authentication and use only by internal security teams. Paid API customers and Google AI Ultra subscribers come later with no date, at two dollars in and ten out, rising to four and twenty. A red panel lists what is not stated.
Drawn from Google's announcement, the Fairwind Program page, DeepMind's Gemini page, Artificial Analysis and cwe-bench.com, all read on 30 September 2026. The dashed column and the red panel mark what is intention or silence, not fact.

What is stated and what is not about access, price and documentation. Sources: Google's announcement and the Fairwind Program page, 30 September 2026.

QuestionStatedNot stated
Who has Argon nowGoogle's own teams; a set of Fairwind partners (650+ partners in the programme)How many partners hold Argon, which, or where they are
What they getArgon "without cyber guardrails"What the guardrails cover, and what is removed
ConditionsVetting, phishing-resistant MFA, internal security teams only, no resale, dual-use tasks only; zero data retention on Gemini EnterpriseWho audits compliance, and what is logged
Everyone else"As soon as possible"; paid API customers and AI Ultra firstA date, a region list, or whether UK customers are included
Price$2 in, $10 out, introductory; $4 and $20 afterwards; cached input 95% offWhen the introductory period ends
Safety documentationThe Frontier Safety Framework is cited; four safeguard areas are describedA model card, a critical capability level, a safety case

The benchmark table, row by row

The table is the centre of the launch. It is reproduced below exactly as Google published it, and DeepMind's own Gemini page prints the same figures as text, which confirms every number we read from the image. Blue cells mark Argon's lead or tie; grey cells mark a rival ahead of it. A dash means the rival has no score on that row.

Google's benchmark table with four columns: Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, across 19 rows in nine groups. Argon is highlighted in blue where it leads or ties, for example Vals Index 68.9%, DeepSWE v1.1 77.9% and CWE-bench v1 68.0%. Rivals are highlighted in grey where they lead, for example FrontierSWE v2 and Terminal-bench 4.0.
Google, Gemini 4 Argon announcement, 30 September 2026. The methodology link at the foot of the image is deepmind.google/models/evals-methodology/gemini-4-argon.

Sorting the rows by who produced the numbers changes the picture, but not in the direction a sceptic might expect. Where Google ran every model itself (four benchmarks) Argon leads three. Where Google ran only Argon and took the rivals from elsewhere (five benchmarks) Argon trails three. Where the numbers come from third-party leaderboards (nine benchmarks) it leads or ties eight. So provenance does not show self-flattery. It shows which rows compare like with like, and that is a separate question.

A bar chart of Argon's margin over the best of three rivals on 19 rows in three groups. Third-party leaderboards: from plus 12.9 on Harvey legal to a tie on CWE-bench and minus 10.5 on FrontierSWE. Google ran Argon only: from plus 3.7 on DeepSWE to minus 10.5 on Terminal-Bench Science. Google ran every model: from plus 12.4 on long GraphWalks to minus 4.0 on PostTrainBench.
Computed from Google's table and its methodology PDF. Margin is Argon minus the highest of the three rivals' scores on that row. GraphWalks is one benchmark at two context lengths.

Six of the rows where Argon leads or ties are by 2 points or less: Vals Index, Vibe Code Bench, GraphWalks to 128k, Agent's Last Exam, Chartography and the CWE-bench tie. Google prints no sample sizes or error bars. Vals does give one: a standard error of about 0.9 points for Sonnet 5.5 and Opus 5.5, so Argon's 1.9-point Vals lead is about two of them (derived; Vals gives none for Argon). The large margins are Harvey's Legal Agent Benchmark (+12.9, 19.6% against 6.7%), the 256k to 1M context range of GraphWalks (+12.4, which rests on 200 problems) and AutomationBench (+8.8). The large deficits are FrontierSWE v2 and Terminal-Bench Science (both -10.5) and Terminal-bench 4.0 (-9.0).

The methodology PDF is candid about how the numbers were produced, and its notes are where the comparisons stop being like for like. The PDF dates its results "as of October, 2026", though the post that links it is dated 30 September; we read that as a slip. The table below lists the rows where the method differs between models. We have covered how much this can matter before: Grok 4.7 scored 38% on Terminal-Bench in its maker's harness and 26% in an independent one.

Method notes from Google's methodology PDF for the rows where models were not run the same way. Source: Google DeepMind, Gemini 4 Argon model evaluation, read on 30 September 2026.

RowWhat the PDF saysWhat it does to the comparison
DeepSWE v1.1Argon run by Google in a mini-swe-agent harness. Astra from the public leaderboard; Fable 5.1 and Opus 5.5 from their system cardsThe 3.7-point lead compares Google's own run with figures taken from elsewhere. The PDF does not state the rivals' harnesses
LVBenchArgon sees 1 frame per second; Astra 800 frames, Fable 5.1 300, Opus 5.5 600, "due to API limitations"Not like for like. Argon's frame count depends on video length and is not stated
OSWorld-2.0Argon's score is the best of three runs, on the offline subset, partial score. Astra from OpenAI's blog post; Anthropic's omitted because it reports online and offline togetherArgon's best of three still trails Astra by 3.4 points
Terminal-Bench Science 0.1Argon run with a 6 times verifier timeout "to address timeout issues". Others from the public leaderboardArgon's setting was changed and it still trails by 10.5 points
Agent's Last ExamArgon run in a 5-hour window with safety filters on; flagged responses returned as empty strings. Fable 5.1 has no leaderboard entryA 1.3-point lead with one of three rivals missing
PostTrainBench, LABBench 2, GraphWalksGoogle ran every model. PostTrainBench on one H100 for 10 hoursConsistent within Google's set-up; not independently reproduced

The coding headline is DeepSWE v1.1: Google says Argon "sets a new state of the art" at 77.9%. The chart below is honest about its axis, which starts at zero. The lead over Opus 5.5 is 3.7 points, over Astra 3.8, and over Fable 5.1 10.5. Two of those rivals' figures come from their own system cards, not from a common run.

A bar chart of DeepSWE v1.1, long-horizon software engineering, higher is better, y-axis from 0 to 100 per cent: Gemini 4 Argon 77.9%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%, Claude Opus 5.5 74.2%.
Google, Gemini 4 Argon announcement, 30 September 2026. The axis starts at zero; the rivals' scores come from a leaderboard and two system cards.

The cybersecurity claims: a 68 per cent tie, and what it cost

Google says that on CWE-bench v1, which tests whether a model can remediate security vulnerabilities, Argon "ties for first place with a top score of 68%". The benchmark's own leaderboard confirms it, and adds what Google's chart leaves out. CWE-bench v1 is 120 held-out audit-and-patch tasks, private to its publisher Collinear AI, each run four times per model. A task passes only if the exploit no longer works and the project's existing tests still pass.

A bar chart titled CWE-bench v1 leaderboard, Pass@1 with ties broken by Pass@4, higher is better. Eleven models with their harness: Grok 4.7 68% opencode, Gemini 4 Argon 68% Antigravity, GPT-6 Astra 68% Codex, Claude Opus 5.5 67% Claude Code, Claude Fable 5.1 58% Claude Code, Grok 4.6 57%, DeepSeek-V4.1-Flash 55%, Muse Spark 1.3 55%, Hy4 Preview 53%, GPT-6 Sol 52% Codex, Inkling 37%.
Google, Gemini 4 Argon announcement, 30 September 2026, drawing on the CWE-bench leaderboard. Each bar measures a model together with its own harness: Argon ran in Google's Antigravity, the OpenAI models in Codex, the Claude models in Claude Code.

Three things follow from the leaderboard itself. First, the tie is three ways, not one: Grok 4.7 and GPT-6 Astra also score 68%, and Opus 5.5 is a point behind at 67%. One point is 1.2 tasks of 120 (derived), so the gap between first and fourth place is about one task. Second, the tie-break puts Argon second of the three: the leaderboard orders equal pass@1 scores by pass@4, and Grok 4.7 scores 81%, Argon 75% and Astra 74%. On pass@4 Opus 5.5, at 79%, is ahead of Argon as well. Third, Google's chart drops the leaderboard's other columns: the judge panel and the cost per rollout.

Six rows of bars from the CWE-bench v1 leaderboard. Pass@1, programmatic: Grok 4.7 68, Gemini 4 Argon 68, GPT-6 Astra 68, Claude Opus 5.5 67, Claude Fable 5.1 58, DeepSeek V4.1 Flash 55. Pass@4: 81, 75, 74, 79, 73, 74. Judge panel pass@1: 64, 62, 66, 67, 60, 53. Cost per rollout in dollars: 2.75, 6.63, 2.85, 0.79, 3.47, 0.09.
Drawn from the CWE-bench v1 leaderboard, read on 30 September 2026. Cost is the mean billed spend for one rollout on one task; Collinear prices Argon from its token counts at $2 per million input, $0.20 per million cached input and $10 per million output.

Argon's mean cost is $6.63 per rollout, against $2.85 for Astra, $2.75 for Grok 4.7 and $0.79 for Opus 5.5: 2.3, 2.4 and 8.4 times (derived). Collinear's cost chart makes the point in its own words: "74× the cost buys 1.2× the pass rate", Argon's 68% at $6.63 against DeepSeek V4.1 Flash's 55% at $0.09. Two cautions on the cost figure. It is priced at the introductory token prices, so at $4 and $20 the same tokens would cost about twice as much (our inference, about $13). And Collinear lists the cached input rate as $0.20 beside a $2 input price, which is 90% off, not the 95% Google states; the effect depends on the cached share, which the page does not give.

The second independent read is the judge panel. Collinear added three rubric-based judges from different model families, which each pass a fix only if the exploit is blocked, normal behaviour is kept, variants of the attack are closed, the grader was not gamed and the codebase still builds. The programmatic score is the ranked one, but the panel's pass@1 is reported beside it. Argon's panel score is 62%, six points below its programmatic 68%. Opus 5.5 scores 67 on both, Astra 66 against 68. Argon is fourth of the five leading models on the panel. The leaderboard does not say why, and a score gap is not a verdict. But for a defender the distinction is the point: a patch that passes the test is not the same as a patch that fixes the flaw.

Two more items sit in the small print. Google says its result "builds on 3.8 Flash Cyber's frontier performance on CWE-bench v0". On v0 that model scored 47.2%, second to Claude Fable 5 at 47.8%. But the v1 leaderboard carries no Flash Cyber row, and Collinear says that for the four models run on both, scores rose "because v1's verifiers are fairer, not because the models got better", so the two versions cannot be read against each other. And the leaderboard does not say whether Argon's entry used the version without cyber guardrails or the standard one. Collinear's launch post, dated 28 September, reports that provider safety filters flagged defensive tasks: Astra on 13 of 120 tasks, Sol on 5, Fable 5.1 on 39 and Opus 5.5 on 9, while Grok and DeepSeek were never flagged. It does not mention Argon. If the tested model was the guardrail-free one, then 68% describes a version only vetted partners can use. If it was the standard one, the score of the version trusted defenders receive is unpublished.

The other two security charts are Google's own. Both come from internal benchmarks, both compare Argon only with its predecessor Gemini 3.8 Flash Cyber, and in both the vertical axis starts at 20, so the bars are drawn from 20, not zero. On the left panel the true ratio is 1.21 and the drawn ratio 1.29; on the right 1.22 and 1.33 (derived).

Two bar charts titled Discovering security vulnerabilities, Pass@1, higher is better. Left, Real-world Vulnerability Discovery, scanning source code across 20 languages: Gemini 4 Argon 85.8%, Gemini 3.8 Flash Cyber 71.0%. Right, Wiz Penetration Test Benchmark, exploiting web vulnerabilities without source code: Gemini 4 Argon 70.9%, Gemini 3.8 Flash Cyber 58.2%. Both vertical axes start at 20.
Google, Gemini 4 Argon announcement, 30 September 2026. Both axes start at 20, so the bars are drawn from 20, not zero.

What the two internal security benchmarks are, from Google's methodology PDF, and what is missing. Source: Google DeepMind, read on 30 September 2026.

BenchmarkWhat the PDF saysWhat is missing
Real-world Vulnerability Discovery: 85.8% against 71.0%An internal dataset testing recall over recent, confirmed historical vulnerabilities in popular open source projects, with source code, in an internal Antigravity harnessNo rival model. Whether the vulnerabilities postdate the training data is not stated
Wiz Penetration Test Benchmark: 70.9% against 58.2%An internal benchmark measuring the ability to write exploits against real web application vulnerabilities, without the codebaseNo rival model, no task count, no outside run

The testimonial has the same shape. Google says Wiz is using Argon through its Scan for Good initiative and that the model found a critical flaw exposing personal information in healthcare software used by hospitals worldwide, "a severe risk that previous frontier models had missed". Wiz's Scan for Good page does not mention Gemini or Argon. It reports programme totals, 17,461 root domains in scope, 326,891 monitored endpoints and 475 high or critical findings, none attributed to Argon. Its examples are anonymised. No vulnerability, vendor or CVE is named, and the claim about earlier models cannot be checked. Wiz is owned by Google, and our earlier briefing found that the terms an organisation accepts for a Scan for Good scan disclaim liability for disruption.

The safeguards: one chart and three absences

Google lists four safeguard areas before broad release: defending against misuse, defending against prompt injection, monitoring for misalignment and hardening systems. Only one comes with a number. The Gray Swan indirect prompt injection chart shows Argon's attack success rate after 15 attempts at 0.7%, against 1.0% for Claude Opus 5.5 and Claude Fable 5.1, 4.6% for Claude Opus 5, 5.5% for Gemini 3.8 Flash, 8.5% for GPT-6 Astra and 10.1% for GPT-6 Sol. Argon's figure is one twelfth of Astra's (derived).

A bar chart titled Gray Swan IPI, attack success rate at k attempts, lower is better. The number above each bar is the rate for 15 attempts: Gemini 4 Argon 0.7%, Claude Opus 5.5 1.0%, Claude Fable 5.1 1.0%, Claude Opus 5 4.6%, Gemini 3.8 Flash 5.5%, Gemini 3.8 Flash Cyber 6.0%, GPT-6 Astra 8.5%, GPT-6 Sol 10.1%, Muse Spark 1.3 15.9%, GPT-5.6 Sol 27.0%, GLM 5.3 31.5%, Grok 4.6 51.8%, Kimi K3 52.7%. Each bar is layered for 1, 10 and 15 attempts.
Google, Gemini 4 Argon announcement, 30 September 2026. Google's methodology PDF says the Gray Swan IPI results were sourced from Gray Swan. We could not load Gray Swan's leaderboard, so the figures rest on this chart.

The 0.3-point gap between Argon and the two Claude models is the claim behind "most resilient yet". The chart gives the attempt budget, up to 15, but neither the chart nor the methodology gives the number of attacks in the set or an interval, so a gap that small cannot be separated from noise on the published evidence. The large gaps, between the top group and models at 27% to 53%, are much less likely to be noise, though without the sample size that is a judgement, not a measurement. We could not check any of it against Gray Swan directly.

The four safeguard areas in Google's announcement and what was published about each. Sources: Google's announcement; Google's agent control roadmap, deepmind.google, 18 June 2026. Read on 30 September 2026.

SafeguardWhat Google saysWhat was measured and published
Defending against misuseDesigned to refuse harmful cyber and CBRN requests; monitoring of internal activations; tested by internal and external red teamsNo refusal rate, no red-team result, no account of what the trusted-defender version removes
Prompt injection"Most resilient model yet"; automated red teaming and adversarial trainingOne chart, from Gray Swan: 0.7% at 15 attempts. No sample size or interval
Misalignment monitoringMonitors chain-of-thought and actions and stops execution "when necessary"Nothing. Google's June roadmap says it measures coverage, recall and time-to-response; none is given here
Hardening systemsSandboxes isolated and sealed before high-risk training or evaluationNo detail. For context, Collinear reported a different vendor's model using an API key from its environment to get round a no-internet sandbox, in 4 of nearly 1,000 rollouts

The framework Google cites says it conducts safety case reviews before external launches when relevant critical capability levels are reached. The announcement does not say whether Argon reached a cyber level, does not link a safety case and, as far as we could find, there is no model card: the announcement, the Fairwind page, the methodology PDF and DeepMind's Gemini page link none, and three likely addresses returned a not-found error. That is an absence from what we could reach, not proof that no document exists. It is, though, the comparison to draw with OpenAI, whose system card rates two of its models Critical for cyber in its first section.

What it costs: the list price, and the price of a task

The list price is easy to state and easy to misread. The introductory $2 and $10 is half the post-introductory $4 and $20, a fact that appears only in footnote 1. A cached input token costs 5% of the input price: $0.10 per million at the introductory rate and $0.20 afterwards (derived). A single maximum-length response of 1 million output tokens costs $10 at the introductory rate and $20 afterwards, before any input (derived).

Argon's token prices and what they imply. Source: Google, 30 September 2026, footnote 1; derived figures are our arithmetic.

ItemIntroductoryAfter the introductory period
Input, per million tokens$2$4
Output, per million tokens$10$20
Cached input, per million tokens (derived)$0.10$0.20
One 1-million-token response, output only (derived)$10$20

A price per token is not a price per task, because a model that thinks longer writes more tokens. Three independent measurements of cost per task point in different directions. On Artificial Analysis's index, Argon (High) costs $1.99 a task against $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5, but it writes 110 million output tokens on the index against 60 million for Astra. On Vals, at the post-introductory $4 and $20, Argon costs $15.68 a test against $18.46 for Astra, $21.34 for Claude Sonnet 5.5 and $32.14 for Opus 5.5. On CWE-bench it costs more than every model that ties or nearly ties it. The workload decides.

Cost per task from three independent sources, read on 30 September 2026. Prices differ by source: Artificial Analysis and Collinear use the introductory $2 and $10; Vals uses $4 and $20.

Source and unitGemini 4 ArgonComparators
Artificial Analysis Intelligence Index, per task$1.99 (index 53, rank 8 of 223)GPT-6 Astra $3.26 (index 53); Opus 5.5 $5.98 (index 58, rank 1)
Vals Index, per test$15.68 (68.90%)Astra $18.46; Sonnet 5.5 $21.34; Opus 5.5 $32.14; GPT-6.1 Sol $3.24 (61.15%)
CWE-bench v1, per rollout$6.63 (68%)Astra $2.85; Grok 4.7 $2.75; Opus 5.5 $0.79 (67%)

Two readings matter. Artificial Analysis scores Argon at 53, equal to GPT-6 Astra and five points behind Claude Opus 5.5 at 58, ranked 8th against 1st, although it runs Argon at its High setting and the others at Max. That is a different set of tests from Google's. It is not a contradiction, but it is the reason to read a vendor's table as a selection. And GPT-6.1 Sol, which our briefing covered, lists at the same $2 input price as Argon's introductory rate, scores 61.15% on the Vals Index at $3.24 a test, against Argon's 68.90% at $15.68. That is 7.75 points for 4.8 times the cost (derived). Whether the extra points are worth it is a decision about your work, not about the model.

What Google says it did with Argon itself

The announcement lists four internal results. They are the most concrete evidence of capability in the post, and they are all Google's own, unreviewed and, in places, unfinished.

Four internal uses claimed in Google's announcement. Source: Google, 30 September 2026.

ClaimStatedNot stated
Quantum subroutine optimisationBeat "the published baseline" by 40% in minutesWhich baseline, which subroutine, or whether it was reviewed
Memory efficiency in data centresOver 300 TiB freed "once rolled out"; an estimated 500 TiB to 1 PiB in totalWhether the 300 TiB is yet realised; the rollout date
C and C++ to Rust migrationUp to 800,000+ lines for the Fuchsia Zircon kernel; under "rigorous automated and manual auditing" before productionAny result for the kernel; defects found; the audit's outcome
libgav1 video decoder32,000 lines of SIMD code replaced; 2.7 times faster than the existing Rust port, identical video outputThe baseline against the C++ original; how the output was compared

The Rust migration matters for this site's readers more than the speed figure. Memory-safety work at kernel scale, done by agents, is a claim about defensive capability in the plain sense: fewer bug classes in shipped code. Google itself says it is not yet in production and is being audited. Until that audit reports, the 800,000-line figure is a scale claim, not a quality one.

What this means for UK readers

Nothing in the announcement is about the UK. It names the US government's pre-release process and no UK body. The Fairwind page prioritises governments and national cyber authorities, critical infrastructure operators and core technology platforms, and does not name the NCSC or any other UK organisation. It does not say whether Fairwind is open to UK organisations, and Google has given no UK date for paid API or AI Ultra access. The Fairwind page does say that, accessed as a managed model on Gemini Enterprise, Argon supports zero data retention, which is the feature a UK organisation handling personal data would ask about first.

For most UK security teams the practical position is that they cannot use Argon yet, and will have to decide what to do when they can. The decision to prepare for is not whether the model is good. It is what your policy says about a model released without cyber guardrails to a vetted group and, on our reading of Google's wording, with them to everyone else, before your suppliers start using one on your code.

Four comforting words, and what each one does not establish

The launch is built from reassuring labels. Each is a fair description of something. None is a control by itself.

Labels in Google's launch and what they describe. Sources: Google's announcement; the Fairwind Program page; Google's methodology PDF.

LabelWhat it describesWhat it does not establish
"Defender" and "trusted"Organisations that passed Google's background checks and accepted Fairwind's termsThat the model is safer in their hands by design. The same methodology PDF describes a benchmark of the model's ability to write exploits against real web applications
"Without cyber guardrails"A version released to Google's teams and vetted partnersWhat was removed, who audits it, or that the published scores describe it
"Fairwind" and "Scan for Good"Programmes with conditions and, in Wiz's case, free scanning of public infrastructureThat the programme's findings came from Argon; Wiz's page does not mention it
"Most resilient yet"The lowest figure on one chart from one benchmark publisherA difference from the next models that the published sample can resolve

Separate the method from the accusation. Google is the model's maker, the author of two of the security benchmarks, the operator of the programme that controls who gets it, and the owner of the company that supplied the testimonial and the penetration-test benchmark. The outside parties are companies too, and each has its own interest in being cited. None of that makes the numbers wrong. It is why the rows that an independent party published, Vals, CWE-bench and Artificial Analysis, carry more weight than the rows that only Google can see, and why the most useful part of this launch is the methodology PDF, which says plainly which is which.

What to do, in order

Take this with you

Before Argon reaches your organisation

  • Read Google's table as a vendor's selection. For each row that resembles your work, note who ran each model and in which harness, using the methodology PDF, before you quote a number.
  • Ask Google, or your account team, for a model card and a statement of whether Argon reached a cyber critical capability level. Do not approve it for your code or data until you have an answer or a written reason for the gap.
  • If you are a Fairwind candidate, read the partner conditions first: phishing-resistant MFA, internal security teams only, no sharing, dual-use tasks only. Check that your contract covers zero data retention.
  • Test on your own code. Take your last 20 fixed vulnerabilities, run each candidate model through the same harness, and record first-attempt pass rate and cost per rollout. Have a person judge whether a fix that passes the test is a fix, because CWE-bench's panel shows the two can differ.
  • Put cost per task, not list price, in the business case, and record tokens written. The same model was cheaper than its rivals on two independent measures and dearer on a third.
  • Decide your own policy for guardrail-free models now: who may use one, on which repositories, with what logging and approval.
  • Watch for four dates: general availability, the end of the introductory price, a model card, and Gray Swan's or Vals's updates to their leaderboards.

The question this leaves

Google has published a methodology note that says plainly which numbers it ran itself, and the announcement text then draws on the rows where Argon is strongest. Five rows where Argon trails are in the table and not in the text. One security score ties at a cost the chart omits, and the safeguards that matter most for a model of this kind have either one chart or none. So the question for anyone deciding about Argon is a narrow one. Which of the eighteen rows in Google's table is the work you would hand this model, who ran that row, and what did it cost?

Sources

  1. PrimaryGemini 4 Argon: our next era of frontier intelligence, 30 September 2026. The announcement. Read in full, including footnote 1 on the price after the introductory period.Googleaccessed 2026-09-30
  2. PrimaryGemini 4 Argon model evaluation: approach, methodology and results (PDF, five pages). Used for who computed each number, the harnesses, the frame counts and the settings.Google DeepMindaccessed 2026-09-30
  3. PrimaryGemini models page. Prints the benchmark table as text, which confirms every figure read from the table image; links no model card.Google DeepMindaccessed 2026-09-30
  4. PrimaryFairwind Program page. Used for the partner count, who is prioritised, the conditions and the statement that a set of partners gets Argon.Google DeepMindaccessed 2026-09-30
  5. PrimaryCWE-bench leaderboard, v1. Used for pass@1, pass@4, judge-panel pass@1, cost per rollout, the task count, the tie-break rule and the pricing note for Argon.Collinear AIaccessed 2026-09-30
  6. PrimaryCWE-bench v1: Defense Is the Harder Test, 28 September 2026. Used for the judge panel, the sandbox findings and the provider safety-filter counts.Collinear AIaccessed 2026-09-30
  7. PrimaryVals Index leaderboard, updated 29 September 2026. Used for Argon's 68.90%, the neighbours, the standard error and cost per test.Vals AIaccessed 2026-09-30
  8. PrimaryGemini 4 Argon (High) model page. Used for the Intelligence Index score, rank, cost per task, output tokens and context window.Artificial Analysisaccessed 2026-09-30
  9. PrimaryGPT-6 Astra (Max) model page. Used for the comparison index score, cost per task and output tokens.Artificial Analysisaccessed 2026-09-30
  10. PrimaryClaude Opus 5.5 (Adaptive Reasoning, Max Effort) model page. Used for the comparison index score, cost per task and output tokens.Artificial Analysisaccessed 2026-09-30
  11. PrimaryScan for Good. Read in full: it does not mention Gemini or Argon. Used for the programme's own counts.Wizaccessed 2026-09-30
  12. PrimaryStrengthening our Frontier Safety Framework, updated 17 April 2026. Used for the statement that safety case reviews take place before external launches when relevant critical capability levels are reached.Google DeepMindaccessed 2026-09-30
  13. PrimarySecuring the future of AI agents, 18 June 2026. Used for the three metrics Google says it uses to measure agent monitoring: coverage, recall and time-to-response.Google DeepMindaccessed 2026-09-30

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

How often

Every new briefing in one email, at 7am, or at 7am, 12:30pm and 6pm. Nothing is sent when nothing is new. Unsubscribe any time.