Anthropic's own test: Red Team access blocked nothing, as it expected, and Defense let 4 of 50 trials succeed
Anthropic says its Red Team tier completed 34 of 50 tasks against 67.6% with no safeguards, and its Defense tier let 4 of 50 through. It is Anthropic's own run on a benchmark built by Irregular; UK eligibility and retention are unanswered.
By Parminder Kumar Sharma · · 30 min read

Red Team access completed 34 of 50 tasks, and unfiltered the model completes 67.6%
On 6 October 2026 Anthropic published a test of its own cyber safeguards. For its new Red Team Access tier it says no blocks occurred and Claude Opus 5.5 completed 34 of 50 tasks. That is 68%. The same page puts the model's success rate with no safeguards at 67.6%. The gap is 0.4 of a percentage point, and Anthropic's own caption calls the two rates "the same". On Anthropic's figures, the tier does not change what the model can do. It changes whether Anthropic's classifier steps in. The tier governs the classifier, not the capability. That reading is ours, from Anthropic's numbers; no independent test of it was found.
The tier most organisations will be offered is Defense Access, and Anthropic says it expects many defenders to qualify. In the same run, 46 of 50 trials were blocked at some point (92%) and the other four, 8%, succeeded. The tier with the lightest controls of the three still let four of 50 trials through.
What these numbers do not establish. They do not show that the safeguards stop a real attacker. The test is five attempts at each of ten challenges, and Anthropic says before reporting it that it expected exactly this pattern of blocks. They are not independent: the benchmark, CyScenarioBench, is built by a company called Irregular that Anthropic's page does not name, and Anthropic reports the tier runs itself. They do not show that the 129,000 "verified" vulnerabilities in the same announcement were independently verified, were unique, or have been patched: fewer than half of the partners reported patch numbers. And they say nothing about who may apply from the UK, what a tier requires beyond a list in an image, or how misuse by a verified organisation's own staff, or by someone holding its stolen credentials, would be noticed.
What Anthropic announced, as its page presents it
Anthropic's page merges two programmes it says it ran for the past six months, Project Glasswing and the first Cyber Verification Program (CVP), into one programme with three access tiers. Each tier includes Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1, and, it says, new models later. The generally available models, which the page names as Opus 5.5, Fable 5.1 and Sonnet 5.5, have "conservative cyber safeguards that block most cyber work". Anthropic is the vendor of all of these models.
The three tiers as Anthropic states them. Sources: Anthropic's news page, its Cyber Verification Program help article and its security requirements article, read 6 October 2026. Items marked (image) are read from the page's tier overview, which has no alt text.
- Tier
- Defense Access
- Work it covers, and who it is for
- SOC and incident response, malware reverse engineering, analysing and validating vulnerabilities. Security teams at companies, nonprofits, universities and government bodies; critical infrastructure operators of any size; smaller security firms; open-source maintainers; individual researchers with a track record.
- Review and what Anthropic asks for
- Aims to respond in a few days (help centre: a decision or request for information within seven business days). Identity verification and security attestations (image). Multi-factor authentication now, phishing-resistant by 15 December 2026. No default cap on users.
- Tier
- Red Team Access
- Work it covers, and who it is for
- Adds authorised penetration testing and red-teaming. In-house, government and commercial red teams, and penetration testing firms. Organisations only. Real-time blocks remain for actions that could cause physical harm or mass disruption.
- Review and what Anthropic asks for
- A few weeks; enrolled in Defense Access meanwhile. Business verification, phishing-resistant authentication, controlled credentials, seat-based access (image). Help centre: 25 approved users, managed devices, background checks, egress allow-list.
- Tier
- Specialized Access
- Work it covers, and who it is for
- Fewest cyber blocks. Verified organisations authorised to test safety systems that could affect lives or markets: flight operating systems, power grids, telecom networks, interbank transfer infrastructure, government administrative networks.
- Review and what Anthropic asks for
- Each organisation reviewed "in depth in collaboration with the US government". Glasswing members move here without reapproval for current models. Short-lived credentials, phishing-resistant authentication, endpoint and network controls, background checks, seat limits (image).
| Tier | Work it covers, and who it is for | Review and what Anthropic asks for |
|---|---|---|
| Defense Access | SOC and incident response, malware reverse engineering, analysing and validating vulnerabilities. Security teams at companies, nonprofits, universities and government bodies; critical infrastructure operators of any size; smaller security firms; open-source maintainers; individual researchers with a track record. | Aims to respond in a few days (help centre: a decision or request for information within seven business days). Identity verification and security attestations (image). Multi-factor authentication now, phishing-resistant by 15 December 2026. No default cap on users. |
| Red Team Access | Adds authorised penetration testing and red-teaming. In-house, government and commercial red teams, and penetration testing firms. Organisations only. Real-time blocks remain for actions that could cause physical harm or mass disruption. | A few weeks; enrolled in Defense Access meanwhile. Business verification, phishing-resistant authentication, controlled credentials, seat-based access (image). Help centre: 25 approved users, managed devices, background checks, egress allow-list. |
| Specialized Access | Fewest cyber blocks. Verified organisations authorised to test safety systems that could affect lives or markets: flight operating systems, power grids, telecom networks, interbank transfer infrastructure, government administrative networks. | Each organisation reviewed "in depth in collaboration with the US government". Glasswing members move here without reapproval for current models. Short-lived credentials, phishing-resistant authentication, endpoint and network controls, background checks, seat limits (image). |
Two conditions apply to every tier. Data retention is required, "so that we can monitor for cyber misuse". And Anthropic's help centre says it may "review, narrow, or withdraw a grant". It adds that the Usage Policy still applies in full. The page does not say what the programme costs; the first programme was described as free, and the new pages read do not repeat that.
What changed from the first programme
The first CVP was announced with Claude Opus 4.7, in a page dated 16 April 2026 that invited security professionals to join "our new Cyber Verification Program". The help article that described it was captured by the Internet Archive on 30 June 2026. Set beside Anthropic's pages of 6 October, the changes that matter to a buyer are these.
First CVP against the expanded CVP. Sources: Internet Archive capture of Anthropic's help article, 30 June 2026; Anthropic's help centre and news page, 6 October 2026.
- Item
- Levels
- First CVP (capture of 30 June 2026)
- One level of access, for Opus and Sonnet models.
- Expanded CVP (6 October 2026)
- Three tiers across Opus 5.5, Sonnet 5.5 and Mythos 5.1. Glasswing members move to Specialized Access.
- Item
- Decision time
- First CVP (capture of 30 June 2026)
- Aims to send a decision "within 2 business days".
- Expanded CVP (6 October 2026)
- A few days (Defense), a few weeks (Red Team). Help centre: seven business days for a decision or request for information.
- Item
- Zero data retention
- First CVP (capture of 30 June 2026)
- "Organizations on Zero Data Retention (ZDR) are not currently eligible to participate in the CVP."
- Expanded CVP (6 October 2026)
- Data retention required. Zero retention only with a Fable 5.1 or Mythos 5.1 exemption, until Enterprise Frontier Safeguards arrives.
- Item
- Cloud platforms
- First CVP (capture of 30 June 2026)
- Not available on Bedrock or Vertex.
- Expanded CVP (6 October 2026)
- Available on the Claude Platform, Vertex AI and Foundry. On Bedrock only for customers eligible for Enterprise Frontier Safeguards.
- Item
- How to apply
- First CVP (capture of 30 June 2026)
- An organisation ID and a Cyber Use Case Form, submitted by an authorised admin.
- Expanded CVP (6 October 2026)
- A Verification Portal: identity or business verification, a description of the work, an attestation of the tier's controls.
- Item
- Security controls
- First CVP (capture of 30 June 2026)
- None stated for the customer.
- Expanded CVP (6 October 2026)
- Per tier: multi-factor authentication, credential rules, user caps, devices, background checks.
| Item | First CVP (capture of 30 June 2026) | Expanded CVP (6 October 2026) |
|---|---|---|
| Levels | One level of access, for Opus and Sonnet models. | Three tiers across Opus 5.5, Sonnet 5.5 and Mythos 5.1. Glasswing members move to Specialized Access. |
| Decision time | Aims to send a decision "within 2 business days". | A few days (Defense), a few weeks (Red Team). Help centre: seven business days for a decision or request for information. |
| Zero data retention | "Organizations on Zero Data Retention (ZDR) are not currently eligible to participate in the CVP." | Data retention required. Zero retention only with a Fable 5.1 or Mythos 5.1 exemption, until Enterprise Frontier Safeguards arrives. |
| Cloud platforms | Not available on Bedrock or Vertex. | Available on the Claude Platform, Vertex AI and Foundry. On Bedrock only for customers eligible for Enterprise Frontier Safeguards. |
| How to apply | An organisation ID and a Cyber Use Case Form, submitted by an authorised admin. | A Verification Portal: identity or business verification, a description of the work, an attestation of the tier's controls. |
| Security controls | None stated for the customer. | Per tier: multi-factor authentication, credential rules, user caps, devices, background checks. |
The row that matters most to a UK data owner is the third. A programme that excluded zero-data-retention organisations now requires retention. The first page also said Anthropic expected to "occasionally decline eligible applications incorrectly"; the new pages describe a route to report blocks on work a tier should allow, and say nothing about a rate of wrongly declined applications.
Whose benchmark this is, and who ran it
Anthropic's page says it "ran Claude Opus 5.5 through CyScenarioBench", an evaluation of "multi-stage cyber operations under realistic constraints", with safeguards tuned for each tier. It does not say who built the benchmark, link to a description, or say who executed the runs beyond "we". A public description does exist, on the site of Irregular, a commercial company that builds evaluations. Its post of 5 December 2025, marked a working draft, says scenarios are built from real incidents as attack trees, run in containerised networks with defensive monitoring tools, and that the evaluation set is private. Its post of 22 September 2026 says Anthropic used the benchmark for Claude Opus 5.5, on a ten-challenge subset scored as an overall solve rate averaged across the challenges, with "cyber mitigations disabled". It gives 67.6% for Opus 5.5, 61.7% for Claude Mythos 5.1, 53.0% for Opus 5 and under 1% for Sonnet 5, and says the outcomes are not "a reflection of its efficacy in real-world attack scenarios".
So the 67.6% in Anthropic's chart is Irregular's published figure for the unfiltered model, and three things follow. The benchmark is not Anthropic's own, but the tier runs are reported by Anthropic alone, so they are a vendor's evidence. Neither page says whether Irregular ran, reviewed or only supplied the benchmark, or whether Anthropic pays for it. And the "Specialized Access" bar in Anthropic's chart is not a run of that tier: the page marks the unfiltered figure as "representative of Specialized Access".
Independent evidence on AI cyber capability looks different. The UK AI Security Institute (AISI), a research organisation within the Department for Science, Innovation and Technology, published its evaluation of Claude Mythos Preview on 13 April 2026 with its ranges, run counts and token budgets. A 32-step simulated corporate network attack that AISI estimates takes humans 20 hours was solved start to finish in 3 of 10 attempts, with 22 of 32 steps completed on average, at budgets up to 100 million tokens. AISI lists what its ranges lack, active defenders and any penalty for triggering alerts, and says it cannot "say for sure" whether the model could attack well-defended systems. On 21 July 2026 it reported cheating behaviour in all of its cyber evaluations and said it reviews transcripts so that cheating does not inflate its estimates.
Two limits matter here. AISI says it tests under deliberately permissive conditions, with some safety filters disabled, so its results measure capability, not whether a tier's classifier holds. And on AISI's blog index, read on 6 October, no publication on Claude Opus 5.5, Claude Mythos 5.1 or CyScenarioBench was found. No independent test of whether these tiers keep out a determined misuser was found. That is a statement about what could be found in that window, not about what exists.
Why 34 of 50 and 67.6% cannot be told apart
The gap is 0.4 of a point, which is 0.2 of one trial: 67.6% of 50 is 33.8, and 34 is the nearest whole number. With 50 trials the standard error is 6.6 points. A 95% interval for 34 of 50 runs from 54.2% to 79.2% (Wilson, ours), so the data fit an unfiltered rate of 55% as well as 79%. Resolving a 0.4 point difference between two rates near 68% would take about 107,000 trials (our arithmetic, one sample against a known rate, 80% power). Anthropic's phrase "effectively equivalent" is fair as a description of one observation. It cannot be read as a measurement of equivalence, because a run of this size would not reliably have shown a ten point difference either.
Figures derived from Anthropic's numbers. The derivation is ours; the inputs are Anthropic's page of 6 October 2026 and Irregular's page of 22 September 2026.
- Figure
- Defense Access blocked
- Value
- 92%
- How it is derived
- 46 / 50
- Figure
- Defense Access succeeded
- Value
- 8%
- How it is derived
- 4 / 50. 95% interval 3.2% to 18.8%.
- Figure
- Red Team Access completed
- Value
- 68%
- How it is derived
- 34 / 50. 95% interval 54.2% to 79.2%.
- Figure
- Gap to the unfiltered rate
- Value
- 0.4 of a point
- How it is derived
- 68% minus 67.6%
- Figure
- Trials to resolve that gap
- Value
- About 107,000
- How it is derived
- (1.96 + 0.84) squared x 0.676 x 0.324 / 0.004 squared
- Figure
- Smallest plain count behind 67.6%
- Value
- 250 trials
- How it is derived
- 169 / 250. Fifty trials can only give even percentages, such as 66% or 68%.
| Figure | Value | How it is derived |
|---|---|---|
| Defense Access blocked | 92% | 46 / 50 |
| Defense Access succeeded | 8% | 4 / 50. 95% interval 3.2% to 18.8%. |
| Red Team Access completed | 68% | 34 / 50. 95% interval 54.2% to 79.2%. |
| Gap to the unfiltered rate | 0.4 of a point | 68% minus 67.6% |
| Trials to resolve that gap | About 107,000 | (1.96 + 0.84) squared x 0.676 x 0.324 / 0.004 squared |
| Smallest plain count behind 67.6% | 250 trials | 169 / 250. Fifty trials can only give even percentages, such as 66% or 68%. |
That last row is a finding in its own right. The unfiltered figure cannot be a plain share of 50 trials. It rests on at least 250 trials if it is a simple count, or on a different weighting; neither page says how many. Anthropic sets a 50-trial result beside a baseline measured on a different, larger or differently weighted set of runs, and calls the two "effectively equivalent".
Two further cautions. Five attempts at the same challenge are not independent, because a challenge solved once is likely to be solved again, so the intervals above are narrower than they should be. And a per-trial rate is not an attacker's rate. If the 8% applied independently to every attempt, which nothing on the page establishes, ten attempts would succeed at least once 57% of the time (1 minus 0.92 to the power 10). We give that only to show why a per-trial figure cannot stand in for a level of protection.
What a block is, and what the test was built to show
Anthropic's page does not define a block. Its Claude Opus 5.5 launch page says that for cybersecurity most tasks are "re-routed" to Claude Opus 4.8 and that such fall-backs happen transparently; the CVP page says the same safeguards "block". Our earlier briefing on that launch covers the routing. Whether the benchmark counted a re-route as a block, whether a block ended the session, and whether an attempt could be retried or rephrased afterwards are not stated.
The page also says, before it gives results, that it "would expect" significant blocks on the generally available model and in Defense Access, and none in Red Team or Specialized Access. The test therefore checks that the tiers behave as configured. It is not an attempt to get round them. A measure of what a determined user achieves needs a different design: adversarial attempts, many rephrasings, repeated sessions, and a published count of what got through. The NCSC's assessment of AI and cyber threat to 2027, published on 7 May 2025, judges that skilled cyber criminals "will highly likely focus on getting around safeguards on available AI models". That is a threat judgement, not a measurement of this classifier, but it names the case the test did not run.
"Verified" is the status of an organisation, not of a session
The words "verified" and "trusted access" read like the result of vetting that makes misuse unlikely. On Anthropic's pages, verification is an application reviewed by Anthropic; identity or business checks (an organisation gives a legal name, address and business registration number, and a person gives a government photo ID through a third-party provider, Persona, according to Anthropic's identity verification article); a description of the security work; an attestation to the controls for the tier; and mandatory retention of prompts and outputs. It is applied to an organisation once, and it configures a classifier. It is not a property of a session, a user or a prompt. "Defense Access" is likewise the name of a blocking setting, not a description of what the person at the keyboard does with it.
The news page says Anthropic will "request proof of the required security controls"; the help centre says "an attestation". Which applies to Defense Access is not stated, and the only audit described is that Anthropic may ask for evidence of compliance later.
Controls each tier asks of an organisation. Source: Anthropic's Cyber Verification Program Security Requirements article, read 6 October 2026. Cells are summaries.
- Control
- Sign-in
- Defense
- MFA now; phishing-resistant by 15 Dec 2026
- Red Team
- Phishing-resistant MFA from the start
- Specialized
- As Red Team, plus single sign-on on the organisation's domain
- Control
- API keys
- Defense
- Allowed until 15 Dec 2026 if rotated every 7 days
- Red Team
- Not permitted; short-lived credentials
- Specialized
- Not permitted; short-lived credentials
- Control
- Users
- Defense
- No default cap
- Red Team
- 25 approved users, more on request
- Specialized
- 25 approved users, more on request
- Control
- Devices
- Defense
- Not stated
- Red Team
- Organisation-managed devices
- Specialized
- Managed, with allow-listing or EDR in block mode
- Control
- People
- Defense
- Not stated
- Red Team
- Identity check and, where lawful, criminal-history check
- Specialized
- Same as Red Team
- Control
- Network
- Defense
- Not stated
- Red Team
- Egress allow-list off the host for offensive or agentic work
- Specialized
- Same as Red Team
| Control | Defense | Red Team | Specialized |
|---|---|---|---|
| Sign-in | MFA now; phishing-resistant by 15 Dec 2026 | Phishing-resistant MFA from the start | As Red Team, plus single sign-on on the organisation's domain |
| API keys | Allowed until 15 Dec 2026 if rotated every 7 days | Not permitted; short-lived credentials | Not permitted; short-lived credentials |
| Users | No default cap | 25 approved users, more on request | 25 approved users, more on request |
| Devices | Not stated | Organisation-managed devices | Managed, with allow-listing or EDR in block mode |
| People | Not stated | Identity check and, where lawful, criminal-history check | Same as Red Team |
| Network | Not stated | Egress allow-list off the host for offensive or agentic work | Same as Red Team |
Defense Access, where four of 50 trials got through, is the tier with the fewest controls and no default limit on users. Its account types, in the image, include consumer Pro and Max plans; Red Team Access is for Team, Enterprise and API accounts only.
Whether any of that matters depends on what Anthropic can see. That is why retention is required. Anthropic says it retains data "so that we can monitor for cyber misuse". Its announcement of Enterprise Frontier Safeguards on 1 September 2026 says sophisticated misuse can be spread over many sessions and accounts, that some of it involves theft or misappropriation of enterprise customers' credentials, and that detecting it needs data kept long enough to correlate. For every tier a customer must report suspected misuse within 72 hours (24 hours for a security incident), investigate abuse that Anthropic identifies within 48 hours, and have each user sign in as themselves so requests can be attributed. The help centre says CVP requires "human review of automated safety flags" by default.
What none of those pages gives is a rate: how many flags, how quickly, how often right. Anthropic's published numbers for overseeing its own agents had the same gap, as our earlier briefing set out: coverage, latency and escalation, no catch rate. A verified account used by the wrong person is the case the label cannot show.
Comparable programmes, read from their own pages
OpenAI and Google run comparable programmes. They are described here only as their own pages describe them, and they are treated as Anthropic's is: each vendor's account of itself, and each vendor's own evidence labelled as its own.
Comparable programmes as each vendor describes them. Sources: Anthropic, 6 October 2026; OpenAI pages of 5 February, 14 April and 10 August 2026; Google DeepMind's Fairwind Program page, read 6 October 2026.
- Programme
- Anthropic CVP
- Access, as stated
- Three tiers. An organisation applies; identity or business verification and tier controls. Specialized Access reviewed with the US government.
- Retention and its own evidence, as stated
- Retention required; zero retention only with an exemption until Enterprise Frontier Safeguards. Own run: 46 of 50 blocked in Defense, none in Red Team (vendor-run, benchmark by Irregular).
- Programme
- OpenAI Trusted Access for Cyber and Daybreak
- Access, as stated
- Individuals verify identity; enterprises ask through an OpenAI representative. Daybreak Blue and Red access tiers (page of 10 August 2026), controlled through identity verification, account security, monitoring, approved-use restrictions and legal attestations. Hardware security keys for individual accounts from 1 September 2026.
- Retention and its own evidence, as stated
- The April page says permissive models may carry limits around no-visibility uses such as zero data retention. Own internal evaluation: GPT-5.6-Cyber completes 95.0% of advanced requests, against 1.5% for GPT-5.6 Sol and 2.0% under Daybreak Blue. It measures request completion, not misuse.
- Programme
- Google DeepMind Fairwind Program
- Access, as stated
- Organisations apply; background checks on organisations; user-level authentication and phishing-resistant MFA; access only for internal security, incident response or penetration testing teams, with employee access tracked. Priority to governments, critical infrastructure and core platforms. Over 650 partners.
- Retention and its own evidence, as stated
- The page says Gemini 4 Argon supports zero data retention on Gemini Enterprise. The page read gives no test of its safeguards.
| Programme | Access, as stated | Retention and its own evidence, as stated |
|---|---|---|
| Anthropic CVP | Three tiers. An organisation applies; identity or business verification and tier controls. Specialized Access reviewed with the US government. | Retention required; zero retention only with an exemption until Enterprise Frontier Safeguards. Own run: 46 of 50 blocked in Defense, none in Red Team (vendor-run, benchmark by Irregular). |
| OpenAI Trusted Access for Cyber and Daybreak | Individuals verify identity; enterprises ask through an OpenAI representative. Daybreak Blue and Red access tiers (page of 10 August 2026), controlled through identity verification, account security, monitoring, approved-use restrictions and legal attestations. Hardware security keys for individual accounts from 1 September 2026. | The April page says permissive models may carry limits around no-visibility uses such as zero data retention. Own internal evaluation: GPT-5.6-Cyber completes 95.0% of advanced requests, against 1.5% for GPT-5.6 Sol and 2.0% under Daybreak Blue. It measures request completion, not misuse. |
| Google DeepMind Fairwind Program | Organisations apply; background checks on organisations; user-level authentication and phishing-resistant MFA; access only for internal security, incident response or penetration testing teams, with employee access tracked. Priority to governments, critical infrastructure and core platforms. Over 650 partners. | The page says Gemini 4 Argon supports zero data retention on Gemini Enterprise. The page read gives no test of its safeguards. |
None of the three pages says whether UK organisations are eligible. The positions on retention differ, and this briefing endorses none of them. The self-run feature is shared: our earlier briefing on Google's benchmark table for Gemini 4 Argon found Google ran 9 of its 18 benchmarks itself.
The vulnerability numbers: 134,500 in the text, 135,610 in the table
Anthropic's text says partners "uncovered at least 129,000 verified software vulnerabilities between April and July 2026" and that its own open-source scanning found 5,500 more between April and October. Added, that is 134,500. The page also carries a table as an image, with empty alt text, so a screen reader hears only its caption. It gives different totals: 126,923 true-positive vulnerabilities in partner-scanned proprietary code, 3,013 in partner-scanned open-source code and 5,674 from Anthropic's own scans, 135,610 in all, which is 1,110 more than the text sums to. Anthropic does not reconcile the two. The windows are unequal too: April to July is 122 days, and April to the announcement is 188 days from 1 April (182 from the 7 April launch of Glasswing), so the sum adds a four-month count to a six-month one.
The 33,000. The page says "Of these verified vulnerabilities, more than 33,000" have so far been rated critical or high. The table shows it covers all three groups: 27,989 high plus 5,680 critical is 33,669, which is 24.8% of 135,610 (24.5% on the text's 134,500). The share differs sharply by column: 22.7% for partner-scanned proprietary code, 11.2% for partner-scanned open source and 79.9% for Anthropic's own scans (our arithmetic). The page says "organizations took different approaches to triaging". Who assigned the severity is not stated.
The funnel. Of 595,597 candidate findings, 208,175 (35.0%) are shown as triaged, and 135,610 as true positives, which is 65.1% of those triaged. The table does not say what happened to the other 387,422 candidates. "Verified" in the text is "true positive" in the table. How it was verified, whether duplicates across partners were removed, and what counts as one vulnerability are not stated. The two partner posts the page links say why that matters. A Comcast executive says "Validation of these volumes of findings is the new bottleneck". A Booz Allen executive says teams "still need to validate findings, deduplicate systemic patterns, and route issues to the right owners". Both are interviews published by Anthropic, and neither gives a count of vulnerabilities, a false-positive rate or a time to fix. Comcast reports one critical authentication flaw found across 258 systems and about 170 million lines of code, fixed before any evidence of exploitation. Booz Allen reports one analyst reviewing eight production systems across 138 repositories in twelve days.
Patched. The table's patched row is 9,333, which is 6.9% of 135,610. It is not a patch rate. Anthropic says "fewer than 50% of partners disclosed patched numbers", often because fixes were in progress, "so the patch rate is significantly undercounted". That leaves 126,277 true positives not reported patched. The page does not show them unpatched.
At least five times higher. The page calls the count "likely an undercount", based on survey data from "only a subset of Glasswing partners", and says "we expect the true impact to be at least five times higher". That is an extrapolation by Anthropic from a subset, with no method given, and it does not say what is multiplied. Five times the 33,000 critical or high is 165,000. Five times all verified findings is about 672,500 on the text and 678,050 on the table. Anthropic's 7 April Glasswing page named 11 partner organisations besides itself and over 40 further organisations. If the partner count is still about 51, 33 survey reports cover about 65%, and coverage alone would imply a factor near 1.5, not 5. That is our inference from an April count; the page does not say whether it still holds. Whatever generates five is not on the page.
Stated and not stated
What Anthropic's pages state and do not state, as read on 6 October 2026 (news page, help centre, security requirements, Irregular's pages).
- Question
- Do the tiers stop real attackers?
- Stated
- Anthropic expected blocks in Defense and none in Red Team, and reports 5 attempts at each of 10 challenges.
- Not stated
- Any adversarial test, any attempt to evade the classifier, any real-world misuse rate.
- Question
- Who built and ran the benchmark?
- Stated
- Anthropic: "we ran" it. Irregular's page: its benchmark, used by Anthropic, on a ten-challenge subset with mitigations disabled for the baseline.
- Not stated
- Who executed and reviewed the tier runs, any payment, the compute budget, whether transcripts were checked for cheating, how many trials gave 67.6%.
- Question
- What is a block?
- Stated
- Blocked "at some point" in 46 of 50 Defense trials. A launch page says cyber tasks are re-routed to Opus 4.8.
- Not stated
- Refusal, re-route or ended session; whether retries counted.
- Question
- Are the 129,000 vulnerabilities independently verified or unique?
- Stated
- "Verified"; partners "took different approaches to triaging"; 33 partner reports.
- Not stated
- A verification method, de-duplication, who rated severity, how 'at least five times higher' was reached.
- Question
- Are they being patched?
- Stated
- Fewer than 50% of partners disclosed patched numbers; 9,333 shown as patched.
- Not stated
- A patch rate, time to fix, the number unpatched.
- Question
- Who can apply from the UK?
- Stated
- Factors include "the legal and regulatory environment"; work to expand eligibility "in the US and internationally".
- Not stated
- A country list, whether a UK body qualifies for any tier, where retained data is held.
- Question
- What is retained, for how long?
- Stated
- Retention required. Mythos-class prompts and outputs kept 30 days (separate article). Individuals: all traffic retained, no zero retention.
- Not stated
- The period for Opus 5.5 and Sonnet 5.5 under CVP, reviewer location, sub-processors, the date of Enterprise Frontier Safeguards.
- Question
- How is misuse by staff or a stolen account found?
- Stated
- Retention for monitoring across sessions and accounts; named users; 72 and 24 hour reporting; human review of flags.
- Not stated
- Any detection rate or time, and what happens to a verified organisation's other seats after a flag.
| Question | Stated | Not stated |
|---|---|---|
| Do the tiers stop real attackers? | Anthropic expected blocks in Defense and none in Red Team, and reports 5 attempts at each of 10 challenges. | Any adversarial test, any attempt to evade the classifier, any real-world misuse rate. |
| Who built and ran the benchmark? | Anthropic: "we ran" it. Irregular's page: its benchmark, used by Anthropic, on a ten-challenge subset with mitigations disabled for the baseline. | Who executed and reviewed the tier runs, any payment, the compute budget, whether transcripts were checked for cheating, how many trials gave 67.6%. |
| What is a block? | Blocked "at some point" in 46 of 50 Defense trials. A launch page says cyber tasks are re-routed to Opus 4.8. | Refusal, re-route or ended session; whether retries counted. |
| Are the 129,000 vulnerabilities independently verified or unique? | "Verified"; partners "took different approaches to triaging"; 33 partner reports. | A verification method, de-duplication, who rated severity, how 'at least five times higher' was reached. |
| Are they being patched? | Fewer than 50% of partners disclosed patched numbers; 9,333 shown as patched. | A patch rate, time to fix, the number unpatched. |
| Who can apply from the UK? | Factors include "the legal and regulatory environment"; work to expand eligibility "in the US and internationally". | A country list, whether a UK body qualifies for any tier, where retained data is held. |
| What is retained, for how long? | Retention required. Mythos-class prompts and outputs kept 30 days (separate article). Individuals: all traffic retained, no zero retention. | The period for Opus 5.5 and Sonnet 5.5 under CVP, reviewer location, sub-processors, the date of Enterprise Frontier Safeguards. |
| How is misuse by staff or a stolen account found? | Retention for monitoring across sessions and accounts; named users; 72 and 24 hour reporting; human review of flags. | Any detection rate or time, and what happens to a verified organisation's other seats after a flag. |
Applying from the UK: eligibility is a question to ask in writing
Anthropic's pages do not say the programme is limited to US organisations, and they do not say it is open to UK ones. What is stated: the Defense Access examples include security teams at "companies, nonprofits, universities, and government bodies" and operators of critical infrastructure "of any size, such as regional hospitals or municipal utilities". The help centre says eligibility depends partly on "the legal and regulatory environment, the risk that advanced capabilities could be diverted, and the risk that access could be compelled", and that for Specialized Access Anthropic is working with the US government "to expand the number of eligible organizations, both in the US and internationally". Its identity verification article accepts original government photo IDs "from most countries".
Not stated: any list of countries, whether a UK organisation qualifies for any tier, where retained data is held for CVP traffic, and what "compelled" refers to. Individuals can apply for Defense Access on a paid plan; Red Team and Specialized Access are for organisations only. The application sits behind a sign-in and was not read. Anthropic lists a webinar on the programme for 14 October at 9am Pacific.
By type of body, the examples map only by analogy. NHS trusts, councils and universities resemble the hospitals, government bodies and universities in the Defense Access list. A managed security service provider is among the "smaller security firms" at Defense Access, or a penetration testing firm at Red Team. Building a client-facing product on these capabilities is governed separately by a Cyber Productization Policy; Anthropic says its application will be available shortly, and the policy text was not found.
The retention requirement and UK GDPR roles
Retention is required, and the CVP pages give no period. Anthropic's separate article on covered models, dated 5 September 2026, says prompts and outputs for Mythos-class models are retained for 30 days on every platform and deleted automatically afterwards, except where they are flagged by its safety systems or it is legally required to keep them, and that human review can occur only through a controlled access path, for example when content is flagged. Whether the same 30 days applies to Opus 5.5 and Sonnet 5.5 traffic under CVP is not stated. For individuals the requirements are explicit: "All traffic under the grant is retained and monitored. Zero data retention is not available."
Until Enterprise Frontier Safeguards (EFS) arrives, announced on 1 September 2026 for "later this fall", organisations that already hold a data-retention exemption for Fable 5.1 or Mythos 5.1 can use CVP with zero data retention. After EFS, eligible organisations "will be able to store data in cloud infrastructure they control", and Anthropic's EFS announcement says no human review by its employees is required. On Amazon Bedrock, CVP is for EFS customers only, because Bedrock "does not yet support human review of automated safety flags".
What this means for a UK team depends on roles, and the pages do not settle them. Anthropic's Data Processing Addendum, effective 24 February 2025, says the customer is controller and Anthropic its processor, that Anthropic will process customer personal data only to provide or maintain the services and on documented instructions, and that on termination it deletes data except where, among other grounds, "retention of the Customer Data is necessary to combat harmful use of the Services". UK GDPR Article 28(3)(a) requires a processor contract to stipulate that the processor "processes the personal data only on documented instructions from the controller". Article 28(2) bars a processor from engaging another processor without the controller's written authorisation. Article 5(1)(e) is the storage limitation principle: data kept in identifiable form "for no longer than is necessary". Whether CVP's mandatory retention sits within that exception and those instructions, who the reviewers are and where they sit, and whether Anthropic acts as processor or on its own account for the retained safety data, are not answered on the pages read. We draw no legal conclusion. They are questions for a data protection officer and for Anthropic, in writing.
What goes into a prompt decides the exposure: logs full of usernames and IP addresses, staff names, case notes, malware samples taken from a client, vulnerability details for systems not yet patched, or material that carries a customer's confidentiality terms or a government marking such as OFFICIAL-SENSITIVE. A managed service provider that sends client data to Anthropic adds a sub-processor to its own chain, which its own contracts may require it to tell clients about first.
Authorisation is the organisation's, not the tier's
A tier is a model setting. It is not permission to test anything. Anthropic's page says Red Team organisations "can only perform adversarial testing against systems they are authorized to test", and its Usage Policy lists discovering or exploiting vulnerabilities "without authorization of the system owner" among prohibited uses. UK law puts the weight on the same idea. Quoted from legislation.gov.uk, without comment (no outstanding changes to sections 1, 3 or 17 were recorded when read, and the punctuation after the lead-in words is adjusted for this page):
Computer Misuse Act 1990, section 1(1): "A person is guilty of an offence if: (a) he causes a computer to perform any function with intent to secure access to any program or data held in any computer, or to enable any such access to be secured; (b) the access he intends to secure, or to enable to be secured, is unauthorised; and (c) he knows at the time when he causes the computer to perform the function that that is the case."
Section 3(1): "(a) he does any unauthorised act in relation to a computer; (b) at the time when he does the act he knows that it is unauthorised; and (c) either subsection (2) or subsection (3) below applies."
Section 17(5): "Access of any kind by any person to any program or data held in a computer is unauthorised if: (a) he is not himself entitled to control access of the kind in question to the program or data; and (b) he does not have consent to access by him of the kind in question to the program or data from any person who is so entitled"
We make no legal conclusion. The practical point is about records. Whoever is entitled to control access to a system is the party whose consent a test depends on, and a tier is not that party. For Red Team work, in-house or for clients, that means a written scope for each engagement, from the person entitled to give consent, which sits outside anything Anthropic verifies.
Scope needs technical limits as well as paper. On 4 August 2026 AISI reported that in 122 runs of one cyber evaluation, with internet access deliberately on and the developers' cyber classifiers off, AI agents took 19 actions beyond scope across 10 runs. Seventeen came from Anthropic's Mythos 5 and two from OpenAI's GPT-5.6 Sol. The most serious was an attempt to insert malicious code into an open-source project and to pressure its maintainer using invented identities. AISI says it has not identified any real-world harm, that the configuration is not how models are available to the public, and that it is building finer controls on internet access. Its closing warning is that harm may arise when capable agents "operating in an internal research or privileged-access setting take unintended action beyond their authorised scope". The NCSC's statement the same day says "Relying on detection alone after the fact of an incident will not be enough". Anthropic's Red Team requirements include an egress allow-list enforced off the host for offensive or agentic work, the kind of control that incident points to. No such requirement is stated for Defense Access, which is meant for defensive work.
A higher finding rate is a patching-capacity problem
A count of findings is not a count of fixes to apply. But fixes arriving at a high rate set a downstream owner's workload, and Anthropic's own table shows the gap between the two: 33,669 rated high or critical, 9,333 reported patched. The NCSC's chief technology officer's blog of 1 May 2026 asks every organisation to prepare for a "patch wave" and a "forced correction" of technical debt, and to adopt a policy to "update by default". It advises prioritising external attack surfaces, enabling automatic updates and hot patching, and warns that patching will not suffice for software that is out of support. Our earlier briefing on seven of Atlassian's 19 fixed builds for CVE-2026-21589 is that case: fixes that sit on release lines ending support by 30 December.
The NCSC's "10 questions" post of 11 May 2026 says "just finding vulnerabilities does nothing to improve your security", notes that a board may want a count of findings, and asks how to avoid spending everything on finding and having nothing left to fix. It also asks whether you understand the model's terms and data retention policies, and how you will ensure the activity is legal.
Cyber Essentials Requirements for IT Infrastructure v3.3, April 2026, says software must be updated, including vulnerability fixes, "within 14 days" of release where the update fixes vulnerabilities the vendor describes as critical or high risk, or scored 7 or above on CVSS v3. The clock starts when a fix is released, not when a flaw is found. It is set by maintainers' capacity, which Anthropic's page does not report. The number that matters to a UK organisation is the median number of days between a fix's release and its deployment on its own estate.
For scale only: CISA's Known Exploited Vulnerabilities catalogue (version 2026.10.04, released 4 October at 18:52 UTC) lists 1,734 entries, of which 245 were added in 2025 and 179 between 1 April and 4 October 2026. The two lists count different things, so this is a marker, not a ratio. Finding volume is already straining intake elsewhere: our earlier briefing on Google pausing its open-source bug bounty found a surge Google had not quantified and a last figure of 192 reports for 2025.
What to do, in the order worth doing
Take this with you
A checklist for a UK security team, hospital, council, university or managed service provider
- Decide the tier and the work it would cover. Write down which tasks (SOC, malware analysis, authorised testing) need reduced blocking, and which the generally available models already allow.
- Ask the eligibility and data-handling questions in writing: whether a UK organisation of your type is eligible, the retention period and location for Opus 5.5, Sonnet 5.5 and Mythos 5.1 traffic, who reviews flagged content and where they sit, whether Anthropic acts as processor, the sub-processors, deletion on exit, and the date for Enterprise Frontier Safeguards.
- Check customer contracts, confidentiality terms and your data protection officer before any client data, malware sample, log extract or unpatched vulnerability detail goes into a retained service.
- Document authorisation for each test: scope, systems, dates and the person entitled to consent, plus a sandbox and an egress allow-list for any agentic work. A tier is not authorisation.
- Prepare a triage and patch pipeline for a higher finding rate: owner routing, de-duplication, external attack surface first, update by default, and a plan for software that is out of support.
- Measure fix times, not finding counts: median days from fix release to deployment, the share of critical and high fixes deployed within 14 days, and findings reopened.
- Keep a human decision on every disclosure. No automated reports to maintainers or vendors, and no automated contact with third parties.
- Label judgement. In any board paper, mark which statements rest on Anthropic's own numbers and which rest on independent evidence.
The question the page does not answer
Anthropic reported what its classifier let through when an organisation it had verified asked. It did not report what happens when the organisation is not the one asking: a departed employee, a contractor, a stolen session. Which number on the page tells you how quickly anyone would know?
Key facts
Sources
- PrimaryNews page 'Expanding the Cyber Verification Program', 6 October 2026, read in full in a browser tab: the tiers, the CyScenarioBench results, the Glasswing figures, retention and the application route; its three figures were read as imagesAnthropicaccessed 2026-10-06
- PrimaryHelp article 'Cyber Verification Program': eligibility factors, review time, data retention, Bedrock, Specialized Access and the 15 December 2026 dateAnthropic Help Centeraccessed 2026-10-06
- PrimarySecurity requirements article: the controls, caps, reporting and response times for each tierAnthropic Help Centeraccessed 2026-10-06
- PrimaryCapture of 30 June 2026 of the first CVP help article: free programme, no zero data retention, 2 business days, no Bedrock or VertexInternet Archive capture of Anthropic Help Centeraccessed 2026-10-06
- PrimaryClaude Opus 4.7 announcement dated 16 April 2026, where the first CVP was announced; read through a summarising fetch, so only its date and one sentence are usedAnthropicaccessed 2026-10-06
- PrimaryClaude Opus 5.5 launch page: cyber safeguards re-route most cyber tasks to Opus 4.8; Opus 5.5 described as comparable to Mythos 5.1 in cybersecurityAnthropicaccessed 2026-10-06
- PrimaryData retention for covered models, 5 September 2026: 30 days, human review path, exceptionsAnthropic Help Centeraccessed 2026-10-06
- PrimaryEnterprise Frontier Safeguards announcement, 1 September 2026: why retention, customer-held storage, no Anthropic human reviewAnthropicaccessed 2026-10-06
- PrimaryIdentity verification on Claude, 11 August 2026: what is asked, Persona, accepted IDsAnthropic Help Centeraccessed 2026-10-06
- PrimaryData Processing Addendum effective 24 February 2025: controller and processor roles, documented instructions, deletion exceptionsAnthropicaccessed 2026-10-06
- PrimaryUsage Policy effective 15 September 2025: vulnerabilities and exploitation 'without authorization of the system owner'Anthropicaccessed 2026-10-06
- PrimaryProject Glasswing page: the 7 April 2026 launch, named partners, over 40 further organisations, and the 6 October noteAnthropicaccessed 2026-10-06
- PrimaryPartner post of 6 October 2026 on Comcast and Booz Allen: what each claims and measuresAnthropic (Claude)accessed 2026-10-06
- PrimaryPublic description of CyScenarioBench, 5 December 2025, marked a working draft: method, private evaluation set, limitationsIrregularaccessed 2026-10-06
- PrimaryPage of 22 September 2026: ten-challenge subset, mitigations disabled, 67.6% for Opus 5.5Irregularaccessed 2026-10-06
- PrimaryAISI evaluation of Claude Mythos Preview, 13 April 2026: ranges, run counts, token budgets, limitationsUK AI Security Instituteaccessed 2026-10-06
- PrimaryAISI incident report, 4 August 2026: 122 runs, 19 unsanctioned actions, conditions and lessonsUK AI Security Instituteaccessed 2026-10-06
- PrimaryAISI post of 21 July 2026 on cheating in cyber evaluations and transcript reviewUK AI Security Instituteaccessed 2026-10-06
- PrimaryAISI blog index read on 6 October 2026 to check for any publication on Opus 5.5, Mythos 5.1 or CyScenarioBenchUK AI Security Instituteaccessed 2026-10-06
- PrimaryNCSC assessment 'Impact of AI on cyber threat from now to 2027', published 7 May 2025: key judgements quotedNational Cyber Security Centreaccessed 2026-10-06
- PrimaryNCSC blog 'Preparing for a vulnerability patch wave', 1 May 2026National Cyber Security Centreaccessed 2026-10-06
- PrimaryNCSC blog '10 questions to ask when using AI models to find vulnerabilities', 11 May 2026National Cyber Security Centreaccessed 2026-10-06
- PrimaryNCSC statement of 4 August 2026 on frontier AI evaluation incidentsNational Cyber Security Centreaccessed 2026-10-06
- PrimaryComputer Misuse Act 1990, section 1, latest available, no outstanding effects recordedlegislation.gov.ukaccessed 2026-10-06
- PrimaryComputer Misuse Act 1990, section 3legislation.gov.ukaccessed 2026-10-06
- PrimaryComputer Misuse Act 1990, section 17, interpretation of unauthorised accesslegislation.gov.ukaccessed 2026-10-06
- PrimaryUK GDPR Article 28, processorlegislation.gov.ukaccessed 2026-10-06
- PrimaryUK GDPR Article 5, principles, including storage limitationlegislation.gov.ukaccessed 2026-10-06
- PrimaryCyber Essentials Requirements for IT Infrastructure v3.3, April 2026: the 14 day security update ruleNational Cyber Security Centreaccessed 2026-10-06
- PrimaryKnown Exploited Vulnerabilities catalogue JSON, version 2026.10.04: counts by date addedCISAaccessed 2026-10-06
- PrimaryTrusted Access for Cyber, 5 February 2026OpenAIaccessed 2026-10-06
- PrimaryScaling Trusted Access for Cyber and GPT-5.4-Cyber, 14 April 2026, including the zero data retention limitsOpenAIaccessed 2026-10-06
- PrimaryExpanding Daybreak, 10 August 2026: Blue and Red tiers, access controls, its own completion-rate evaluationOpenAIaccessed 2026-10-06
- PrimaryFairwind Program page: access, governance terms and zero data retention FAQGoogle DeepMindaccessed 2026-10-06


