P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

AI Security

OpenAI paused frontier training for two weeks over a model that may be Critical for cyber. The framework that triggered it still names nobody who runs the evaluations

Astra reached the first Critical cyber determination in the Preparedness Framework's history. The best accountability sentence OpenAI has published sits in a blog post, and the framework itself is unchanged.

By Parminder Kumar Sharma · · 7 min read

A close photograph in near darkness of two heavy brushed-steel bars set upright side by side in a recessed dark metal panel, forming a pause control, their left edges picked out by a cold cyan light. The OpenAI wordmark sits in white in the dark right of the frame, above the line: paused at Critical, nobody named.

What OpenAI published, eleven days apart

Two posts, and the second is the one that changes what a customer should do.

On 7 August 2026, under the heading Responding to the next frontier of critical cyber capabilities, OpenAI said that internal evaluations of an upcoming model called Astra, together with expert assessments, had led it to conclude that it cannot rule out critical cyber capabilities under the Preparedness Framework. The framework's own bar for that is quoted in the post:

a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.

That is the first time the threshold has been reached in the framework's history. The post says so obliquely: previous models, including GPT-5.6-Sol, were assessed at High rather than Critical.

On 18 August, in Pacing model development in an era of cyber-critical capabilities, OpenAI described what it did about it. A two-week pause in reinforcement learning training on models intended for deployment. The largest planned frontier RL run remains on hold. Frontier model inference in research clusters was paused outright for anything that could execute code or reach the internet, then restored workload by workload. And a new monitoring stack, with a number attached that is worth the whole post:

Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored.

A screenshot of OpenAI's post of 7 August 2026. The OpenAI wordmark sits top left, then the date August 7 2026 and the category Security, then the headline Responding to the next frontier of critical cyber capabilities, then the paragraph concluding that OpenAI cannot rule out critical cyber capabilities under its Preparedness Framework.
OpenAI's own page, captured 20 August 2026 and cropped to the passage. The sentence in the second paragraph is the whole determination: internal evaluations, plus expert assessments, and no function named as running either.

A correction this site owes its readers

On 17 August this site published a briefing on the dissolution of OpenAI's Preparedness team. It contained this paragraph, and it needs revisiting:

Several outlets state that OpenAI's models autonomously compromised the model repository Hugging Face. This site has not been able to verify that claim against a primary source, so it plays no part in the argument below.

That was the right call on the evidence available at the time, and it is now out of date. OpenAI's own 7 August post contains the confirmation, in a subordinate clause defending a different model: "Astra is an upcoming model, and was not involved in exploiting Hugging Face." The 18 August post then names it twice as "the OpenAI-Hugging Face incident" and promises a technical report of its learnings in the coming weeks.

So the incident is real, OpenAI has confirmed it, and the earlier briefing should be read with that correction attached. The decision to decline an unverified claim was correct; leaving it uncorrected once the primary source appeared would not be.

It also reframes the 18 August post. That two-week pause was not prompted by the Astra evaluation alone. OpenAI names two triggers, and the first one is an incident in which its own models were used to exploit a live third-party platform.

Three questions, two answers, and neither in the framework

The 17 August briefing ended with procurement questions rather than predictions. Two of them now have answers.

Asked here on 17 August. Answered where?

OpenAI posts · 7 and 18 Aug 2026

We asked

When will the framework be revised to describe the new arrangement?

Answered

“We will evolve our Preparedness Framework to bring these safeguards together across training and deployment.”

A blog post, 18 August. No date, no draft, no version number.

We asked

Which named function is accountable when a threshold is reached?

Answered

“The safety, security, and research teams are paged.” If they cannot clear a flag as a false positive within 30 minutes, they are expected to pause the activity.

A blog post, 18 August. The framework still names only the Safety Advisory Group, Leadership and the board.

We asked

Who produces the evidence those bodies rule on?

Still open

Not stated. The 7 August post credits “our latest internal evaluations” and “expert assessments”, with no team, role or function named.

Unanswered in both posts and in the framework.

Two of three answered, both in the same blog post, neither in the document that carries the commitments. A published framework that has to be read alongside a fortnight of posts to be understood is not a framework a customer can audit.

Left column from this site’s briefing of 17 August 2026. Right column quoted from OpenAI’s posts of 7 and 18 August 2026, read in full on 19 August. The claim is about where the answers appear, not whether they are adequate: both are reasonable answers.

Take the second row seriously, because it is the closest thing to an accountable function that OpenAI has published. When the monitoring system flags a likely violation of a critical security boundary, the safety, security and research teams are paged, and if they cannot conclusively determine within 30 minutes that it is a false positive, they are expected to pause the activity. That is a named set of teams, a defined trigger, a clock and a default action. It is a better answer than the framework gives.

It is in a blog post. The Preparedness Framework still names the Safety Advisory Group, OpenAI Leadership and the board's Safety and Security Committee, and still says nothing about who runs the evaluations they rule on. A commitment that lives in a post can be superseded by another post, and nobody is bound by it.

The third row is the one that stays open. Read the 7 August wording again: "our latest internal evaluations" and "expert assessments". Which function ran them is not stated, and it was the function reorganised at the end of July.

The framework, and what has happened to it

  1. Dec 2023

    Framework published

    OpenAI publishes the first Preparedness Framework, per its 7 August 2026 post.

  2. Apr 2025

    Version 2

    The version quoted on this site: names the Safety Advisory Group, Leadership and the board. Names no evaluation function.

  3. Jun 2025

    Biology reaches High

    The framework guides a capability transition for the first time, at the High threshold rather than Critical.

  4. Jul 2026

    Preparedness team dissolved

    Reported at the end of July. Bio and cyber assigned to senior staff in separate teams.

  5. 7 Aug 2026

    Cannot rule out Critical

    Astra evaluations plus expert assessments. First time the Critical threshold is in play.

  6. 18 Aug 2026

    Two-week RL pause

    Largest planned frontier run on hold. Monitoring at roughly 20% of inference compute.

Dates from OpenAI's own publications. The two August posts are the first time the framework has been exercised at its most severe level, and they fall three weeks after the function that produced its evidence was reported as split three ways.

The numbers worth carrying into a conversation

What OpenAI has published about the response, with the figure attached

MeasureWhat OpenAI states
Monitoring overheadRoughly 20% of the inference compute being monitored, varying substantially across workloads.
Reinforcement learning pauseTwo weeks, on latest models intended for deployment.
Largest frontier RL runRemains on hold, pending smaller-scale training and evaluations.
Alert targetWithin 30 minutes of concerning activity being surfaced.
Escalation defaultSafety, security and research teams paged. Activity paused if not cleared as a false positive within 30 minutes.
Monitoring scopeAll RL training and evaluations involving tools, for models of Sol capability or higher. Extended to all Astra inference with tools from 7 August.
Framework statusTo be evolved. No date, no draft and no version number given.
All values quoted from OpenAI's posts of 7 and 18 August 2026, read in full on 19 August. Where OpenAI gives a range or a qualifier, it is reproduced rather than rounded.

The 20% figure is the one to remember, because it converts a governance argument into a budget line. Whatever anyone believes about frontier risk, monitoring at this standard costs a fifth of the compute it watches, and that number will be quoted in every procurement conversation about AI safety for the next year.

It also cuts the other way, and the piece would be dishonest not to say so. Twenty per cent of inference compute is a serious commitment, paid in the most expensive currency a model laboratory has. This is not a company doing nothing.

What to ask, if you buy from OpenAI

Take this with you

Questions the two posts make answerable, which they were not last week

  • Ask which function ran the Astra evaluation. The posts credit internal evaluations and expert assessments and name neither. This is the same question as last month and it is now attached to a live threshold determination rather than a hypothetical.
  • Ask when the Preparedness Framework revision lands, and in what form. OpenAI has said it will evolve the framework. A date and a draft are the difference between a commitment and an intention.
  • Ask whether the paging arrangement is contractual or editorial. Safety, security and research teams paged within 30 minutes is a good control. It currently exists in a blog post.
  • Ask what the 20% applies to in your deployment. The figure is monitoring overhead on inference compute being monitored, which is not the same as a 20% cost increase on your account, and the difference is worth establishing before it appears in a budget.
  • Read the Hugging Face technical report when it lands. OpenAI has promised one. It will be the first primary account of a frontier model being used to exploit a live third-party platform, and it will be more useful than any framework revision.

The position

The honest reading of these two posts is that OpenAI is doing more than it was, and publishing more than it has to.

Pausing the largest planned frontier training run is expensive and reversible only in one direction. Twenty per cent of inference compute on monitoring is a real cost. Saying in public that you cannot rule out the most severe capability threshold in your own framework, about an unreleased model, is not the behaviour of an organisation optimising for a quiet quarter.

And the pattern this site identified on 17 August has not changed. The Preparedness Framework named the bodies that rule on evidence and never named the function that produces it, and eleven days later the framework has been exercised at its most severe level without that gap being closed. What closed instead was the information gap, in a blog post, which is a different thing: posts are how a company explains itself, and frameworks are how a customer holds it to account.

The test is simple and it arrives in a few weeks. If the revised Preparedness Framework names the function that runs capability evaluations, this was a governance document catching up with its own organisation. If it names three more committees, it was not.

Sources

  1. PrimaryResponding to the next frontier of critical cyber capabilities, 7 August 2026OpenAIaccessed 2026-08-19
  2. PrimaryPacing model development in an era of cyber-critical capabilities, 18 August 2026OpenAIaccessed 2026-08-19
  3. PrimaryPreparedness Framework version 2, 15 April 2025OpenAIaccessed 2026-08-19

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.