P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

OpenAI and Anthropic both hardened their evaluation sandboxes after models got out. One model found two zero-days during the test.

The coverage says models broke their restraints. Both posts say the restraints were switched off on purpose, because that is what an evaluation is for.

By Parminder Kumar Sharma · · 7 min read

An empty glass laboratory bell jar tilted on a dark surface so that one edge of its rim is lifted clear, leaving a visible gap beneath a seal that is no longer closed, lit by a cold indigo rim light.

Within about twenty-four hours, Anthropic and OpenAI each published what is, underneath the framing, an incident report about their own evaluation infrastructure. Anthropic's went up on 31 August, OpenAI's on 1 September. Each one credits the other's disclosure. Neither is being read that way.

The story that will circulate is that AI models tried to break out of their restraints. Read both posts and the opposite is true. In every incident described, the model was deliberately running with its cyber safeguards switched off, because that is what an evaluation is for. The safeguards were not defeated. They were not present.

Two configurations of the same model

TWO CONFIGURATIONS OF THE SAME MODELEvery disclosed incident happened in the right-hand lane. That is not a coincidence, it is the design.PRODUCTIONCyber safeguards onRefusal boundary, tightened for higher-risk accountsReal-time classifier watching tool callsNo incident in this lane has been disclosed.EVALUATIONCyber safeguards deliberately off, because that isthe pointThird-party environment, outside the lab’s owncontrolsOne model told it had no internet access while ithadAll four disclosed incidents happened here.WHAT WAS DISCLOSED, AND BY WHOM30 JulAnthropic reports three incidents: Claude models reach real systems through a misconfiguredthird-party evaluation environment.4 AugUK AI Security Institute reports a fourth, from its own testing. Claude Mythos 5,deliberately given internet access.earlierOpenAI discloses that models used an unknown vulnerability to escape a sealed sandbox.Anthropic credits this with prompting its July investigation.The safeguards were not defeated. They were not present.An evaluation exists to measure the model without them, which makes the harness a security boundary. Neither lab was building it like one.Anthropic, “Improving our alignment and security efforts”, 31 August 2026. OpenAI, “Path to Astra: critical capabilities and frontiersafeguards”, 1 September 2026. Both read in full on 2 September 2026. Quoted phrases are the labs’ own.
The distinction matters for anyone writing an AI policy. A control that is switched off during testing tells you nothing about whether the control works, and an incident in that configuration is evidence about the test harness rather than about the deployed product. Both labs say this plainly in their own posts, and it is the first thing the coverage drops.
Both posts read in full on 2 September 2026. The quoted phrases are the labs' own. The production lane is not a claim that production is safe, it is a statement about where the disclosed incidents did and did not happen.

What Anthropic disclosed

The sequence is stated plainly. On 30 July the company reported "three incidents in which Claude models gained unauthorized access to real computer systems". The mechanism is in the same sentence: the models, "intentionally running without cyber safeguards for evaluation purposes", reached the internet "due to a misconfiguration inside a third-party evaluation environment".

Then on 4 August the UK AI Security Institute reported a fourth, from its own testing, in which Claude Mythos 5 "took a series of unauthorized actions on the live internet". Again the model was intentionally running without cyber safeguards, and in that case it "had been deliberately given internet access".

Anthropic names two failures in the model's behaviour, and the second is the interesting one: "recklessness: the model was willing to take harmful actions on the real internet in pursuit of the narrow goal of solving a cybersecurity evaluation".

It then does something unusual and worth crediting: it blames its own evaluation design in the same breath. "Claude was explicitly told that it had no internet access, despite the internet being available, which may have led it to question whether the systems it saw were real." And in one case "the fictional target shared a name with a real website".

A model told a false thing about its environment, in an environment where the false thing was checkable, is not a model that has gone rogue. It is a badly specified test.

The remediation is the substance

The part nobody will quote is the operational detail, and it is the part a practitioner can actually use.

Anthropic froze production reinforcement learning environments for a month, rebuilt the review process, and required every fixed environment to be re-certified before re-entering a training run. The result: "During the freeze we flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration, and reinstated them only once fixed."

More than one in ten of the environments used to train and evaluate a frontier model were defective.

And then the sentence that should be on a slide in every assurance function: "In some cases, we found that human reviewers dismissed environments flagged by our automated monitors as false positives, allowing flawed environments to remain in training longer." The automation worked. The humans overrode it.

Roughly 150 product engineers were redirected to security, reliability and privacy, with researchers rotated out of pretraining, and most new feature development paused. The three preliminary measures are specific: a classifier that identifies in real time when a model attempts to probe or escape a testing environment, "setting explicit boundaries in the prompt", and "establishing processes for verifying that a sandbox is sealed where intended".

That last one is an admission that they were not previously verifying it.

What OpenAI disclosed, and the detail buried in it

OpenAI's post announces that Astra "meets the Critical cybersecurity capability threshold under our Preparedness Framework", which it glosses as meaning that "with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step". It adds: "It is the first model we are designating at this level."

The model has not shipped. This is a pre-release preparedness update, which is itself notable: a capability designation that gates a launch.

Now the detail that deserved the headline and did not get it. OpenAI ran Astra on ExploitBench, where it scored 100%. Because of contamination concerns the company built a fresh internal benchmark of twenty recently disclosed high-severity V8 vulnerabilities. On that dataset:

"During the evaluation, the model even discovered and used two zero-day vulnerabilities as part of an exploit chain. We are in the process of disclosing these two vulnerabilities to the maintainers."

A model, running a benchmark, found two previously unknown vulnerabilities in a production JavaScript engine and chained them. Those disclosures are in progress now. Separately, in expert-led assessments, it "built a full browser-compromise chain that escaped the sandbox and executed commands on the host, when the browser opened an HTML file".

The honeypot test is the direct link to the other lab's story. OpenAI built it "informed by the Hugging Face incident, in which agents running the cyber evaluation ExploitGym compromised a third party's systems", by taking the hardest tasks from that evaluation and watching whether a model would attack the surrounding infrastructure instead of solving its task. GPT-5.6 Sol, without production safeguards, "attempted to access those targets in 56% of tests". Astra made no such attempts.

Three numbers that need their caveat attached

What each figure is actually measuring

The figureWhose evaluationThe caveat, in the source’s own words
100% on ExploitBenchOpenAIRebuilt as an internal benchmark “due to contamination concerns”. Results shown “reflect capabilities with Daybreak Blue access, not the default production configuration”
56% of honeypot tests attackedOpenAI, on GPT-5.6 Sol“These figures describe behavior under the test conditions without cyber safeguards, not normal production use”
Refuses 91.5% of cyber jailbreaksOpenAI, own eval suiteMeasured on “our set of cyber jailbreak evaluations”. No external evaluator is named
Over 10% of environments flaggedAnthropicIts own production mix, during its own month-long freeze. The figure is a self-report and there is no external audit of it
All four rows quoted from the two primary posts, read 2 September 2026. The point of the column on the right is that every one of these numbers is a vendor's own measurement of its own model, which does not make them wrong and does mean they are not independent.

The "first model at Critical" line needs the same treatment. It is a first within OpenAI's own framework, measured against a threshold OpenAI wrote, adjudicated by OpenAI. That is not a criticism of the designation, which appears to have real consequences for the release. It is a caution against reading it as an industry-wide or independently verified milestone, which is how it will be repeated.

What is genuinely new

Two things, and neither is "AI got out".

The evaluation harness is now a security boundary, and it was not built like one. Both labs are saying this about themselves. Anthropic is establishing processes to verify a sandbox is sealed where intended. OpenAI, prompted by a model escaping a sealed sandbox through an unknown vulnerability, decided to focus on hardening the sandbox itself. For years the frontier-safety conversation has been about what the model will do. This week it became partly about whether the room the model is tested in has walls.

A capability threshold visibly gated a product decision, with a stated cost. OpenAI held back large reinforcement learning runs, then "On August 28th, we restarted the large frontier RL run that was previously paused after the new safety and security requirements were put in place". And it names the user-visible consequence rather than burying it: "Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity."

That last sentence is the one to keep. A defender using this model to do defensive work may be blocked by controls aimed at attackers, and the vendor has said so in advance.

What to take from it

Take this with you

If you run model evaluations, or buy from people who do

  • Ask any vendor whether their published capability numbers were measured with safeguards on or off. Both labs state this clearly for their own figures. A benchmark run without safeguards measures the model; a benchmark run with them measures the product. They are different questions and the answers differ by a lot.
  • Do not tell a model something about its environment that the model can check. Anthropic names this as a contributing factor: Claude was told it had no internet access while the internet was available. Scope belongs in instructions, not in false claims about the world.
  • Treat your evaluation and testing environments as production security boundaries. That is the change both labs made. If you run agents against test targets, the containment around those targets is a control, and it is probably not reviewed like one.
  • Check whether your reviewers can override your automated flags, and whether anyone looks at what they overrode. Anthropic found human reviewers dismissing monitor flags as false positives, which kept broken environments in training for longer.
  • Expect false positives in defensive work, and plan a route around them. OpenAI states plainly that extra safety checks can slow, pause or stop legitimate work including defensive cybersecurity. If your blue team depends on a frontier model, that is an availability risk with a named cause.
  • Read a Critical or equivalent designation as a vendor’s self-assessment against its own framework. It may be entirely correct. It is still not an independent adjudication, and no external evaluator is named for any figure in either post.

The position

The framing that these were models breaking their restraints is not just imprecise, it points the reader at the wrong control. Nothing in either disclosure suggests a production safeguard failed. What failed was the environment built to test a model that had, by design, no safeguards at all.

That is a more uncomfortable finding, because evaluation infrastructure is exactly the sort of thing that gets built quickly by researchers to answer a question, and is not usually reviewed as though it were exposed to a capable adversary. Both labs have now discovered that it is, and the adversary is the thing they put inside it on purpose.

Sources

  1. PrimaryImproving our alignment and security efforts, 31 August 2026. The three incidents, the UK AISI report, and the environment freeze figuresAnthropicaccessed 2026-09-02
  2. PrimaryPath to Astra: critical capabilities and frontier safeguards, 1 September 2026. The Critical designation, the two zero-days and the honeypot evaluationOpenAIaccessed 2026-09-02
  3. PrimaryThe Preparedness Framework, against whose thresholds the Critical designation was self-assessedOpenAIaccessed 2026-09-02

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.