Gemini 4 Argon leads or ties 13 of 18 benchmarks in Google's table, and Google ran 9 of the 18 itself
Google's table puts Gemini 4 Argon ahead of or level with GPT-6 Astra and two Claude models on 13 of 18 benchmarks. Independent sources confirm the Vals lead and show the CWE-bench tie costs $6.63 a rollout, while only vetted defenders can use the model.
By Parminder Kumar Sharma · · 24 min read

Led or tied on 13 of 18, with Google scoring 9 of the 18 itself
Google's comparison table for Gemini 4 Argon, published on 30 September 2026, sets it against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 18 benchmarks (19 printed rows, because GraphWalks appears at two context lengths). By our count of the printed figures Argon leads outright on 12, ties for first on one and trails on five. Google's methodology PDF says it computed Argon's own score on nine of the 18. The other nine come from third-party leaderboards. The announcement text names seven of the 13 benchmarks Argon leads or ties, and none of the five it trails: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0 appear only in the table image.
What that does not establish. It does not make the table wrong. Vals, whose leaderboards supply four of those nine rows, lists Argon first at 68.90%, and on the five benchmarks it trails Google printed the numbers. It does not show that the leads carry over to any team's own work: the table has no sample sizes, no error bars and no cost per task, and Google's own notes say that on several rows the rivals were run under different settings from Argon. And three things are absent from the announcement: a model card, a statement of whether Argon reached a critical capability level for cyber under Google's Frontier Safety Framework, and a date on which most people can use the model.
The launch post uses the vocabulary of defence: trusted defenders, frontier safeguards, a programme called Fairwind. The record is a model that, on the one security benchmark with a public leaderboard, ties two rivals at more than twice their cost per rollout, and that only organisations Google has vetted can use. OpenAI's system card, by contrast, rates GPT-6 Astra and GPT-6.1 Sol Critical for cyber, as our briefing on dots and Sol set out. Google's announcement is silent on the equivalent rating.
Six claims in Google's announcement, and what a second source shows. Sources: Google, 30 September 2026; Vals; cwe-bench.com; Wiz; Artificial Analysis; all read on 30 September 2026.
| Google's claim | What a second source shows | Status |
|---|---|---|
| Leading on the Vals Index, 68.9% | Vals lists 68.90% at rank 1. Claude Sonnet 5.5, absent from Google's table, is second at 67.04%. Vals gives a standard error of about 0.9 for its two neighbours | Confirmed. The lead is 1.9 points |
| Ties for first on CWE-bench v1, 68% | A three-way tie with Grok 4.7 and GPT-6 Astra. Argon costs $6.63 a rollout against $0.79 for Opus 5.5, and is fourth of five on the judge panel | Tie confirmed. Cost and panel score are not in Google's chart |
| Most resilient model yet to indirect prompt injection, 0.7% on Gray Swan IPI | Gray Swan's leaderboard returned an error when we fetched it. The next models on Google's chart are at 1.0% | Not checked by us |
| Wiz found a critical healthcare flaw that earlier models missed | Wiz's Scan for Good page does not mention Gemini or Argon. No flaw, vendor or CVE is named | Not checkable |
| Argon can autonomously find, validate and patch critical vulnerabilities | Patching is measured: 68% of CWE-bench tasks pass on a first attempt. Finding is measured only on two internal benchmarks, against Argon's predecessor | Partly measured |
| Released without cyber guardrails to trusted defenders | Fairwind lists 650+ partners and its conditions. Neither page says how many hold Argon or what the guardrails do | Stated, not described |
What Google announced, and who can use it
Argon is Google's new frontier model, and on the day of the announcement it is not generally available. Google says it is rolling out to a set of trusted cyber defenders through the Fairwind Program, and that it is engaged in the US government's voluntary process for pre-release model access while it gradually expands access. Wider release to developers, enterprises and consumers comes "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. No date is given for any step.
The numbers that are given are these. The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. Footnote 1 says that after the introductory period the price becomes $4 and $20; the end of that period is not dated. The maximum output per response rises to 1 million tokens, from 64,000, about 15 times (derived). The announcement does not state the input context window; Artificial Analysis lists 1 million tokens, with text and image input and text output.
Two sentences carry most of the security weight. Google says it trained Argon to be highly capable at cyber defence and that it can "autonomously find, validate, and patch critical software vulnerabilities". And for trusted defenders and its own teams it says it will release Argon "without cyber guardrails". The Fairwind page says a set of its partners get exclusive access to Argon, and that the programme works with more than 650 partners in all. It does not say how large that set is.
What is stated and what is not about access, price and documentation. Sources: Google's announcement and the Fairwind Program page, 30 September 2026.
| Question | Stated | Not stated |
|---|---|---|
| Who has Argon now | Google's own teams; a set of Fairwind partners (650+ partners in the programme) | How many partners hold Argon, which, or where they are |
| What they get | Argon "without cyber guardrails" | What the guardrails cover, and what is removed |
| Conditions | Vetting, phishing-resistant MFA, internal security teams only, no resale, dual-use tasks only; zero data retention on Gemini Enterprise | Who audits compliance, and what is logged |
| Everyone else | "As soon as possible"; paid API customers and AI Ultra first | A date, a region list, or whether UK customers are included |
| Price | $2 in, $10 out, introductory; $4 and $20 afterwards; cached input 95% off | When the introductory period ends |
| Safety documentation | The Frontier Safety Framework is cited; four safeguard areas are described | A model card, a critical capability level, a safety case |
The benchmark table, row by row
The table is the centre of the launch. It is reproduced below exactly as Google published it, and DeepMind's own Gemini page prints the same figures as text, which confirms every number we read from the image. Blue cells mark Argon's lead or tie; grey cells mark a rival ahead of it. A dash means the rival has no score on that row.

Sorting the rows by who produced the numbers changes the picture, but not in the direction a sceptic might expect. Where Google ran every model itself (four benchmarks) Argon leads three. Where Google ran only Argon and took the rivals from elsewhere (five benchmarks) Argon trails three. Where the numbers come from third-party leaderboards (nine benchmarks) it leads or ties eight. So provenance does not show self-flattery. It shows which rows compare like with like, and that is a separate question.
Six of the rows where Argon leads or ties are by 2 points or less: Vals Index, Vibe Code Bench, GraphWalks to 128k, Agent's Last Exam, Chartography and the CWE-bench tie. Google prints no sample sizes or error bars. Vals does give one: a standard error of about 0.9 points for Sonnet 5.5 and Opus 5.5, so Argon's 1.9-point Vals lead is about two of them (derived; Vals gives none for Argon). The large margins are Harvey's Legal Agent Benchmark (+12.9, 19.6% against 6.7%), the 256k to 1M context range of GraphWalks (+12.4, which rests on 200 problems) and AutomationBench (+8.8). The large deficits are FrontierSWE v2 and Terminal-Bench Science (both -10.5) and Terminal-bench 4.0 (-9.0).
The methodology PDF is candid about how the numbers were produced, and its notes are where the comparisons stop being like for like. The PDF dates its results "as of October, 2026", though the post that links it is dated 30 September; we read that as a slip. The table below lists the rows where the method differs between models. We have covered how much this can matter before: Grok 4.7 scored 38% on Terminal-Bench in its maker's harness and 26% in an independent one.
Method notes from Google's methodology PDF for the rows where models were not run the same way. Source: Google DeepMind, Gemini 4 Argon model evaluation, read on 30 September 2026.
| Row | What the PDF says | What it does to the comparison |
|---|---|---|
| DeepSWE v1.1 | Argon run by Google in a mini-swe-agent harness. Astra from the public leaderboard; Fable 5.1 and Opus 5.5 from their system cards | The 3.7-point lead compares Google's own run with figures taken from elsewhere. The PDF does not state the rivals' harnesses |
| LVBench | Argon sees 1 frame per second; Astra 800 frames, Fable 5.1 300, Opus 5.5 600, "due to API limitations" | Not like for like. Argon's frame count depends on video length and is not stated |
| OSWorld-2.0 | Argon's score is the best of three runs, on the offline subset, partial score. Astra from OpenAI's blog post; Anthropic's omitted because it reports online and offline together | Argon's best of three still trails Astra by 3.4 points |
| Terminal-Bench Science 0.1 | Argon run with a 6 times verifier timeout "to address timeout issues". Others from the public leaderboard | Argon's setting was changed and it still trails by 10.5 points |
| Agent's Last Exam | Argon run in a 5-hour window with safety filters on; flagged responses returned as empty strings. Fable 5.1 has no leaderboard entry | A 1.3-point lead with one of three rivals missing |
| PostTrainBench, LABBench 2, GraphWalks | Google ran every model. PostTrainBench on one H100 for 10 hours | Consistent within Google's set-up; not independently reproduced |
The coding headline is DeepSWE v1.1: Google says Argon "sets a new state of the art" at 77.9%. The chart below is honest about its axis, which starts at zero. The lead over Opus 5.5 is 3.7 points, over Astra 3.8, and over Fable 5.1 10.5. Two of those rivals' figures come from their own system cards, not from a common run.

The cybersecurity claims: a 68 per cent tie, and what it cost
Google says that on CWE-bench v1, which tests whether a model can remediate security vulnerabilities, Argon "ties for first place with a top score of 68%". The benchmark's own leaderboard confirms it, and adds what Google's chart leaves out. CWE-bench v1 is 120 held-out audit-and-patch tasks, private to its publisher Collinear AI, each run four times per model. A task passes only if the exploit no longer works and the project's existing tests still pass.

Three things follow from the leaderboard itself. First, the tie is three ways, not one: Grok 4.7 and GPT-6 Astra also score 68%, and Opus 5.5 is a point behind at 67%. One point is 1.2 tasks of 120 (derived), so the gap between first and fourth place is about one task. Second, the tie-break puts Argon second of the three: the leaderboard orders equal pass@1 scores by pass@4, and Grok 4.7 scores 81%, Argon 75% and Astra 74%. On pass@4 Opus 5.5, at 79%, is ahead of Argon as well. Third, Google's chart drops the leaderboard's other columns: the judge panel and the cost per rollout.
Argon's mean cost is $6.63 per rollout, against $2.85 for Astra, $2.75 for Grok 4.7 and $0.79 for Opus 5.5: 2.3, 2.4 and 8.4 times (derived). Collinear's cost chart makes the point in its own words: "74× the cost buys 1.2× the pass rate", Argon's 68% at $6.63 against DeepSeek V4.1 Flash's 55% at $0.09. Two cautions on the cost figure. It is priced at the introductory token prices, so at $4 and $20 the same tokens would cost about twice as much (our inference, about $13). And Collinear lists the cached input rate as $0.20 beside a $2 input price, which is 90% off, not the 95% Google states; the effect depends on the cached share, which the page does not give.
The second independent read is the judge panel. Collinear added three rubric-based judges from different model families, which each pass a fix only if the exploit is blocked, normal behaviour is kept, variants of the attack are closed, the grader was not gamed and the codebase still builds. The programmatic score is the ranked one, but the panel's pass@1 is reported beside it. Argon's panel score is 62%, six points below its programmatic 68%. Opus 5.5 scores 67 on both, Astra 66 against 68. Argon is fourth of the five leading models on the panel. The leaderboard does not say why, and a score gap is not a verdict. But for a defender the distinction is the point: a patch that passes the test is not the same as a patch that fixes the flaw.
Two more items sit in the small print. Google says its result "builds on 3.8 Flash Cyber's frontier performance on CWE-bench v0". On v0 that model scored 47.2%, second to Claude Fable 5 at 47.8%. But the v1 leaderboard carries no Flash Cyber row, and Collinear says that for the four models run on both, scores rose "because v1's verifiers are fairer, not because the models got better", so the two versions cannot be read against each other. And the leaderboard does not say whether Argon's entry used the version without cyber guardrails or the standard one. Collinear's launch post, dated 28 September, reports that provider safety filters flagged defensive tasks: Astra on 13 of 120 tasks, Sol on 5, Fable 5.1 on 39 and Opus 5.5 on 9, while Grok and DeepSeek were never flagged. It does not mention Argon. If the tested model was the guardrail-free one, then 68% describes a version only vetted partners can use. If it was the standard one, the score of the version trusted defenders receive is unpublished.
The other two security charts are Google's own. Both come from internal benchmarks, both compare Argon only with its predecessor Gemini 3.8 Flash Cyber, and in both the vertical axis starts at 20, so the bars are drawn from 20, not zero. On the left panel the true ratio is 1.21 and the drawn ratio 1.29; on the right 1.22 and 1.33 (derived).

What the two internal security benchmarks are, from Google's methodology PDF, and what is missing. Source: Google DeepMind, read on 30 September 2026.
| Benchmark | What the PDF says | What is missing |
|---|---|---|
| Real-world Vulnerability Discovery: 85.8% against 71.0% | An internal dataset testing recall over recent, confirmed historical vulnerabilities in popular open source projects, with source code, in an internal Antigravity harness | No rival model. Whether the vulnerabilities postdate the training data is not stated |
| Wiz Penetration Test Benchmark: 70.9% against 58.2% | An internal benchmark measuring the ability to write exploits against real web application vulnerabilities, without the codebase | No rival model, no task count, no outside run |
The testimonial has the same shape. Google says Wiz is using Argon through its Scan for Good initiative and that the model found a critical flaw exposing personal information in healthcare software used by hospitals worldwide, "a severe risk that previous frontier models had missed". Wiz's Scan for Good page does not mention Gemini or Argon. It reports programme totals, 17,461 root domains in scope, 326,891 monitored endpoints and 475 high or critical findings, none attributed to Argon. Its examples are anonymised. No vulnerability, vendor or CVE is named, and the claim about earlier models cannot be checked. Wiz is owned by Google, and our earlier briefing found that the terms an organisation accepts for a Scan for Good scan disclaim liability for disruption.
The safeguards: one chart and three absences
Google lists four safeguard areas before broad release: defending against misuse, defending against prompt injection, monitoring for misalignment and hardening systems. Only one comes with a number. The Gray Swan indirect prompt injection chart shows Argon's attack success rate after 15 attempts at 0.7%, against 1.0% for Claude Opus 5.5 and Claude Fable 5.1, 4.6% for Claude Opus 5, 5.5% for Gemini 3.8 Flash, 8.5% for GPT-6 Astra and 10.1% for GPT-6 Sol. Argon's figure is one twelfth of Astra's (derived).

The 0.3-point gap between Argon and the two Claude models is the claim behind "most resilient yet". The chart gives the attempt budget, up to 15, but neither the chart nor the methodology gives the number of attacks in the set or an interval, so a gap that small cannot be separated from noise on the published evidence. The large gaps, between the top group and models at 27% to 53%, are much less likely to be noise, though without the sample size that is a judgement, not a measurement. We could not check any of it against Gray Swan directly.
The four safeguard areas in Google's announcement and what was published about each. Sources: Google's announcement; Google's agent control roadmap, deepmind.google, 18 June 2026. Read on 30 September 2026.
| Safeguard | What Google says | What was measured and published |
|---|---|---|
| Defending against misuse | Designed to refuse harmful cyber and CBRN requests; monitoring of internal activations; tested by internal and external red teams | No refusal rate, no red-team result, no account of what the trusted-defender version removes |
| Prompt injection | "Most resilient model yet"; automated red teaming and adversarial training | One chart, from Gray Swan: 0.7% at 15 attempts. No sample size or interval |
| Misalignment monitoring | Monitors chain-of-thought and actions and stops execution "when necessary" | Nothing. Google's June roadmap says it measures coverage, recall and time-to-response; none is given here |
| Hardening systems | Sandboxes isolated and sealed before high-risk training or evaluation | No detail. For context, Collinear reported a different vendor's model using an API key from its environment to get round a no-internet sandbox, in 4 of nearly 1,000 rollouts |
The framework Google cites says it conducts safety case reviews before external launches when relevant critical capability levels are reached. The announcement does not say whether Argon reached a cyber level, does not link a safety case and, as far as we could find, there is no model card: the announcement, the Fairwind page, the methodology PDF and DeepMind's Gemini page link none, and three likely addresses returned a not-found error. That is an absence from what we could reach, not proof that no document exists. It is, though, the comparison to draw with OpenAI, whose system card rates two of its models Critical for cyber in its first section.
What it costs: the list price, and the price of a task
The list price is easy to state and easy to misread. The introductory $2 and $10 is half the post-introductory $4 and $20, a fact that appears only in footnote 1. A cached input token costs 5% of the input price: $0.10 per million at the introductory rate and $0.20 afterwards (derived). A single maximum-length response of 1 million output tokens costs $10 at the introductory rate and $20 afterwards, before any input (derived).
Argon's token prices and what they imply. Source: Google, 30 September 2026, footnote 1; derived figures are our arithmetic.
| Item | Introductory | After the introductory period |
|---|---|---|
| Input, per million tokens | $2 | $4 |
| Output, per million tokens | $10 | $20 |
| Cached input, per million tokens (derived) | $0.10 | $0.20 |
| One 1-million-token response, output only (derived) | $10 | $20 |
A price per token is not a price per task, because a model that thinks longer writes more tokens. Three independent measurements of cost per task point in different directions. On Artificial Analysis's index, Argon (High) costs $1.99 a task against $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5, but it writes 110 million output tokens on the index against 60 million for Astra. On Vals, at the post-introductory $4 and $20, Argon costs $15.68 a test against $18.46 for Astra, $21.34 for Claude Sonnet 5.5 and $32.14 for Opus 5.5. On CWE-bench it costs more than every model that ties or nearly ties it. The workload decides.
Cost per task from three independent sources, read on 30 September 2026. Prices differ by source: Artificial Analysis and Collinear use the introductory $2 and $10; Vals uses $4 and $20.
| Source and unit | Gemini 4 Argon | Comparators |
|---|---|---|
| Artificial Analysis Intelligence Index, per task | $1.99 (index 53, rank 8 of 223) | GPT-6 Astra $3.26 (index 53); Opus 5.5 $5.98 (index 58, rank 1) |
| Vals Index, per test | $15.68 (68.90%) | Astra $18.46; Sonnet 5.5 $21.34; Opus 5.5 $32.14; GPT-6.1 Sol $3.24 (61.15%) |
| CWE-bench v1, per rollout | $6.63 (68%) | Astra $2.85; Grok 4.7 $2.75; Opus 5.5 $0.79 (67%) |
Two readings matter. Artificial Analysis scores Argon at 53, equal to GPT-6 Astra and five points behind Claude Opus 5.5 at 58, ranked 8th against 1st, although it runs Argon at its High setting and the others at Max. That is a different set of tests from Google's. It is not a contradiction, but it is the reason to read a vendor's table as a selection. And GPT-6.1 Sol, which our briefing covered, lists at the same $2 input price as Argon's introductory rate, scores 61.15% on the Vals Index at $3.24 a test, against Argon's 68.90% at $15.68. That is 7.75 points for 4.8 times the cost (derived). Whether the extra points are worth it is a decision about your work, not about the model.
What Google says it did with Argon itself
The announcement lists four internal results. They are the most concrete evidence of capability in the post, and they are all Google's own, unreviewed and, in places, unfinished.
Four internal uses claimed in Google's announcement. Source: Google, 30 September 2026.
| Claim | Stated | Not stated |
|---|---|---|
| Quantum subroutine optimisation | Beat "the published baseline" by 40% in minutes | Which baseline, which subroutine, or whether it was reviewed |
| Memory efficiency in data centres | Over 300 TiB freed "once rolled out"; an estimated 500 TiB to 1 PiB in total | Whether the 300 TiB is yet realised; the rollout date |
| C and C++ to Rust migration | Up to 800,000+ lines for the Fuchsia Zircon kernel; under "rigorous automated and manual auditing" before production | Any result for the kernel; defects found; the audit's outcome |
| libgav1 video decoder | 32,000 lines of SIMD code replaced; 2.7 times faster than the existing Rust port, identical video output | The baseline against the C++ original; how the output was compared |
The Rust migration matters for this site's readers more than the speed figure. Memory-safety work at kernel scale, done by agents, is a claim about defensive capability in the plain sense: fewer bug classes in shipped code. Google itself says it is not yet in production and is being audited. Until that audit reports, the 800,000-line figure is a scale claim, not a quality one.
What this means for UK readers
Nothing in the announcement is about the UK. It names the US government's pre-release process and no UK body. The Fairwind page prioritises governments and national cyber authorities, critical infrastructure operators and core technology platforms, and does not name the NCSC or any other UK organisation. It does not say whether Fairwind is open to UK organisations, and Google has given no UK date for paid API or AI Ultra access. The Fairwind page does say that, accessed as a managed model on Gemini Enterprise, Argon supports zero data retention, which is the feature a UK organisation handling personal data would ask about first.
For most UK security teams the practical position is that they cannot use Argon yet, and will have to decide what to do when they can. The decision to prepare for is not whether the model is good. It is what your policy says about a model released without cyber guardrails to a vetted group and, on our reading of Google's wording, with them to everyone else, before your suppliers start using one on your code.
Four comforting words, and what each one does not establish
The launch is built from reassuring labels. Each is a fair description of something. None is a control by itself.
Labels in Google's launch and what they describe. Sources: Google's announcement; the Fairwind Program page; Google's methodology PDF.
| Label | What it describes | What it does not establish |
|---|---|---|
| "Defender" and "trusted" | Organisations that passed Google's background checks and accepted Fairwind's terms | That the model is safer in their hands by design. The same methodology PDF describes a benchmark of the model's ability to write exploits against real web applications |
| "Without cyber guardrails" | A version released to Google's teams and vetted partners | What was removed, who audits it, or that the published scores describe it |
| "Fairwind" and "Scan for Good" | Programmes with conditions and, in Wiz's case, free scanning of public infrastructure | That the programme's findings came from Argon; Wiz's page does not mention it |
| "Most resilient yet" | The lowest figure on one chart from one benchmark publisher | A difference from the next models that the published sample can resolve |
Separate the method from the accusation. Google is the model's maker, the author of two of the security benchmarks, the operator of the programme that controls who gets it, and the owner of the company that supplied the testimonial and the penetration-test benchmark. The outside parties are companies too, and each has its own interest in being cited. None of that makes the numbers wrong. It is why the rows that an independent party published, Vals, CWE-bench and Artificial Analysis, carry more weight than the rows that only Google can see, and why the most useful part of this launch is the methodology PDF, which says plainly which is which.
What to do, in order
Take this with you
Before Argon reaches your organisation
- Read Google's table as a vendor's selection. For each row that resembles your work, note who ran each model and in which harness, using the methodology PDF, before you quote a number.
- Ask Google, or your account team, for a model card and a statement of whether Argon reached a cyber critical capability level. Do not approve it for your code or data until you have an answer or a written reason for the gap.
- If you are a Fairwind candidate, read the partner conditions first: phishing-resistant MFA, internal security teams only, no sharing, dual-use tasks only. Check that your contract covers zero data retention.
- Test on your own code. Take your last 20 fixed vulnerabilities, run each candidate model through the same harness, and record first-attempt pass rate and cost per rollout. Have a person judge whether a fix that passes the test is a fix, because CWE-bench's panel shows the two can differ.
- Put cost per task, not list price, in the business case, and record tokens written. The same model was cheaper than its rivals on two independent measures and dearer on a third.
- Decide your own policy for guardrail-free models now: who may use one, on which repositories, with what logging and approval.
- Watch for four dates: general availability, the end of the introductory price, a model card, and Gray Swan's or Vals's updates to their leaderboards.
The question this leaves
Google has published a methodology note that says plainly which numbers it ran itself, and the announcement text then draws on the rows where Argon is strongest. Five rows where Argon trails are in the table and not in the text. One security score ties at a cost the chart omits, and the safeguards that matter most for a model of this kind have either one chart or none. So the question for anyone deciding about Argon is a narrow one. Which of the eighteen rows in Google's table is the work you would hand this model, who ran that row, and what did it cost?
Sources
- PrimaryGemini 4 Argon: our next era of frontier intelligence, 30 September 2026. The announcement. Read in full, including footnote 1 on the price after the introductory period.Googleaccessed 2026-09-30
- PrimaryGemini 4 Argon model evaluation: approach, methodology and results (PDF, five pages). Used for who computed each number, the harnesses, the frame counts and the settings.Google DeepMindaccessed 2026-09-30
- PrimaryGemini models page. Prints the benchmark table as text, which confirms every figure read from the table image; links no model card.Google DeepMindaccessed 2026-09-30
- PrimaryFairwind Program page. Used for the partner count, who is prioritised, the conditions and the statement that a set of partners gets Argon.Google DeepMindaccessed 2026-09-30
- PrimaryCWE-bench leaderboard, v1. Used for pass@1, pass@4, judge-panel pass@1, cost per rollout, the task count, the tie-break rule and the pricing note for Argon.Collinear AIaccessed 2026-09-30
- PrimaryCWE-bench v1: Defense Is the Harder Test, 28 September 2026. Used for the judge panel, the sandbox findings and the provider safety-filter counts.Collinear AIaccessed 2026-09-30
- PrimaryVals Index leaderboard, updated 29 September 2026. Used for Argon's 68.90%, the neighbours, the standard error and cost per test.Vals AIaccessed 2026-09-30
- PrimaryGemini 4 Argon (High) model page. Used for the Intelligence Index score, rank, cost per task, output tokens and context window.Artificial Analysisaccessed 2026-09-30
- PrimaryGPT-6 Astra (Max) model page. Used for the comparison index score, cost per task and output tokens.Artificial Analysisaccessed 2026-09-30
- PrimaryClaude Opus 5.5 (Adaptive Reasoning, Max Effort) model page. Used for the comparison index score, cost per task and output tokens.Artificial Analysisaccessed 2026-09-30
- PrimaryScan for Good. Read in full: it does not mention Gemini or Argon. Used for the programme's own counts.Wizaccessed 2026-09-30
- PrimaryStrengthening our Frontier Safety Framework, updated 17 April 2026. Used for the statement that safety case reviews take place before external launches when relevant critical capability levels are reached.Google DeepMindaccessed 2026-09-30
- PrimarySecuring the future of AI agents, 18 June 2026. Used for the three metrics Google says it uses to measure agent monitoring: coverage, recall and time-to-response.Google DeepMindaccessed 2026-09-30


