P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

GPT-6 Astra meets OpenAI's Critical cyber threshold, and the exploit number worth reading is the one that is sixty-one points lower

OpenAI built a contamination-free benchmark to undercut their own headline score and published the worse result beside it. Both numbers are theirs. Only one will be quoted.

By Parminder Kumar Sharma · · 11 min read

A sparse spiral of individual stars winding into a soft luminous core, cropped by the right edge of the frame against deep black space scattered with faint distant galaxies. The left half of the frame is empty sky.

OpenAI shipped GPT-6 Astra today, and in their own words it "meets the Critical threshold in cybersecurity under our Preparedness Framework".

That sentence is the news. Everything else in a very long announcement is downstream of it.

On 19 August this site asked three questions about what would happen when a model actually reached that threshold. Two had been answered in a blog post. The third, which was who produces the evidence the framework's committees rule on, had not. It still has not. The announcement credits "internal and third-party expert evaluations" and "expert-led assessments", and names no function.

The two exploit numbers

Two exploit numbers, drawn on the same axis

TWO EXPLOIT NUMBERS, DRAWN ON THE SAME AXISOpenAI published both. One of them is the headline.ExploitBenchknown vulnerabilities, some of which predatethe training cut-offGPT-6 Astra100.0%GPT-5.6 Sol78.5%Claude Opus 570.0%a perfect score, and the sentence everyone will useExploitBench, June to August 202620 high-severity V8 flaws from the previousthree months, across 13 stable Chrome releasesGPT-6 Astra39.0%GPT-5.6 Sol11.5%built because OpenAI raised the contaminationobjection against their own headline numberOpenAI’s own footnote: some included vulnerabilities may not permitarbitrary code execution under the evaluation’s constraints, so 100%may not be achievable here. 39.0% is not a ceiling measurement.On SRE-Bench, reverse engineering binaries without source,Astra solved 88.0% first try and 99.2% within four attempts,against 55.9% and 68.7% for Sol, and 12.5% for Opus 5.Both numbers are OpenAI’s. The one that is harder to contaminate is the one that is 61 points lower.
Building a contamination-free benchmark and publishing a much lower score on it, in the same post as the perfect one, is the opposite of the behaviour this site spent a fortnight documenting. It deserves saying before anything else. The observation here is only that a reader who takes away “100% on ExploitBench” has taken the number OpenAI themselves flagged as the less trustworthy of the two.
Both figures are from OpenAI's announcement of 3 September 2026 and its published tables, read in full. Drawn on one 0 to 100 axis, because two panels at different scales would let 39.0 read as comparable to 100.

On ExploitBench, which measures turning known vulnerabilities into working exploits, Astra scored 100.0%. GPT-5.6 Sol scored 78.5%. Claude Opus 5 scored 70%.

That is the sentence that will travel, and OpenAI supplied the reason not to trust it. Concerned that "exposure to historical software vulnerabilities may have affected benchmark results", they built a second version from twenty high-severity V8 vulnerabilities disclosed in the previous three months, across thirteen stable Chrome releases.

On that one Astra scored 39.0%, against 11.5% for Sol.

Both numbers are real, both are theirs, and the honest reading needs both. The perfect score tells you very little, because the model may have read the write-ups. The 39.0% is the harder number to contaminate and it is sixty-one points lower.

It also comes with a qualification OpenAI put in a footnote and which cuts the other way: some of those twenty vulnerabilities may not permit arbitrary code execution at all under the evaluation's constraints, so 100% may not be achievable on it. 39.0% is not a measurement against a reachable ceiling. Anyone quoting either figure without its pair has taken half the evidence.

The result that impressed me most is neither. On SRE-Bench, which tests reverse engineering compiled binaries with no access to source, Astra solved 88.0% of tasks first try and 99.2% within four attempts, against 55.9% and 68.7% for Sol, and 12.5% for Opus 5. That is not a benchmark you can contaminate by reading disclosure blogs, and it is the capability that changes what an attacker can do with a binary nobody has published source for.

And during the contamination-free evaluation, Astra found and used two previously unknown zero-day vulnerabilities. OpenAI say they are disclosing both to the maintainers.

Three labs, one week, three different answers

This is now a pattern rather than an event, and this site has covered all three as they landed.

On 1 September, Anthropic's Fable 5.1 shipped with a capability boundary: it can find vulnerabilities but not develop exploits for them, enforced by the vendor, generally available.

On 2 September, Google's Gemini 3.8 Flash Cyber shipped with more permissive cyber mitigations and no stated capability line, gated instead by an access programme whose enforcement is the customer's own joiners and leavers process.

Today OpenAI ships a model at the Critical threshold with a staged relaxation. Astra as launched "will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits". Then: "Through OpenAI Daybreak, we plan to expand access and roll out less restrictive safeguards in the coming weeks."

Three frontier labs, three different compensating controls, one direction of travel. If your AI usage policy assumes vendors tighten cyber controls over time, it was written against the wrong trend, and it is now three data points wrong rather than one.

The change you will actually feel

The benchmark table is not what most people will notice first. The turn structure is.

Where the question lands

Where the question lands

the same underspecified request, both lanes

GPT-5.6 Sol

one turn: work, then answer

Request

underspecified, as most are

Works

gaps filled silently

Answers

assumptions baked into the output

GPT-6 Astra

the question moves to the front, and the independent work does not stop for it

Request

the same request

Asks

only where the answer changes the outcome

Keeps working

on the parts that do not depend on the reply

Answers

or proceeds on sensible assumptions if you skip

The claim is about ordering, not politeness. A model that asks after it has finished has already chosen; a model that asks first has moved the decision back to you while it was still cheap to change. OpenAI say Astra asks only where the answer could change the outcome, proceeds on assumptions if you skip, and waits where the decision is consequential. Which of those three it picks is the model’s judgement, and it is not a setting you control.

Drawn from OpenAI’s description rather than from their screenshots. The behaviour they document is conditional in three ways at once, so the honest version of it is a branch, not a promise: it asks when the answer could change the outcome, it continues on work that does not depend on your reply, and it proceeds without you where it judges that acceptable.
Drawn from OpenAI's description of the behaviour rather than from their product screenshots. The three conditions are theirs: it asks where the answer could change the outcome, continues on work that does not depend on the reply, and proceeds on assumptions where it judges that acceptable.

OpenAI say Astra "uses context to fill in routine gaps and asks focused questions when the answer could change the outcome", and that in Codex "it can ask asynchronously while continuing work that doesn't depend on your reply".

This is a better default than silently guessing, and it is worth being clear about why. A model that asks after it has finished has already made the choice; a model that asks first has handed the decision back while changing it was still cheap.

The part to hold onto is that all three branches are the model's judgement. Whether a gap is routine, whether an answer would change the outcome, whether a decision is consequential enough to wait for you. None of those is a setting you configure, and the failure mode is quiet: a question not asked looks exactly like a task that did not need one.

What OpenAI published that went the wrong way

Two things in this announcement point against the launch, and both are OpenAI's own findings.

The first: "Our evaluations found Astra's written reasoning harder to monitor than GPT-5.6 Sol's, based on tests that explicitly asked it to evade monitoring." They attribute it to Astra having greater control over its written reasoning and solving problems in fewer written steps, they say it still struggles to conceal reasoning on complex tasks, and they say they take the decline seriously.

Publishing a monitorability regression on launch day, in the same post as a Critical cyber threshold, is not what a marketing document does. It should be said plainly before anything else in this section.

The second is quieter and sits in the comparison tables. On Humanity's Last Exam with tools, Astra scores 57.2% against 65.0% for Claude Fable 5.1. On the Artificial Analysis Intelligence Index, Astra is 61.2 against Fable 5.1's 65.7. On the Artificial Analysis Coding Agent Index, Astra is 67.0 against Claude Opus 5's 68.1. On FrontierCode 1.1 Main, Astra's 53.3 sits behind Fable 5's 53.5 and Opus 5's 53.4.

The post opens by calling Astra "the world's most intelligent and aligned model". Four rows in OpenAI's own tables say something narrower than that, and they published all four.

Two footnotes worth reading

Comparison details that only appear in the footnotes

What the table showsWhat the footnote saysWhy it matters
Claude scores on ScreenSpot-Pro and ExploitGymThose figures come from Mythos, described as “Fable with fewer safeguards”The comparator is not the model a customer receives, so on those two rows Astra is measured against a less restricted variant than the shipping one
Fable 5 and 5.1 absent from LifeSciBench, GeneBench Pro and MedChemBenchThey “refuse the majority of questions in these evaluations”A blank is not a low score. On those three rows the competitor declined rather than failed, which is a different result
FrontierMath Tier 4 “saturated with a 98% score”The comparison table reports 97.6%Rounding, and small, but the prose figure and the table figure are not the same number
FrontierCode resultsAstra was run with a developer message resembling its Codex prompt, which OpenAI say was not optimised for the evalDisclosed, and worth knowing when comparing against models run without one
From the footnotes to OpenAI's GPT-6 Astra announcement, 3 September 2026, read in full. Neither invalidates the results. Both change how a specific number should be read.

The Mythos row is the one I would raise with a vendor. On two benchmarks the published Claude comparison is against a variant with fewer safeguards than the shipping product. OpenAI disclose it, in footnote seventeen, and it does not appear anywhere near the chart.

Beyond the terminal

For readers whose interest is not security: the professional-work claims are the substantive half of this release. Astra reaches 95.9% on BenchCAD against 83.3% for Sol, 64.6% on Terminal-Bench Science against 52.6% for Fable 5.1, and 41.4% on AutomationBench against 31.4%. On long context, OpenAI report 96.3% on the 512K to 1M eight-needle retrieval test, against 73.8% for Sol.

OpenAI also demonstrate the model modelling a house in Blender and turning it into a walkable Unreal Engine 5 scene, and doing PCB layout in KiCad from a schematic. Both are edited demonstration clips rather than benchmarks, and OpenAI label them as such in a footnote. They are worth watching anyway, because they show the shape of the claim better than the tables do: not a chatbot that writes code, but an agent driving the desktop software the work actually lives in.

What to do about it

Take this with you

This week, in order

  • Note that Enterprise access is off by default at launch and requires an administrator to enable it. That is your decision point, and it exists once. Decide before somebody asks.
  • Do not quote 100% on ExploitBench without the 39.0%. OpenAI published both and built the second one specifically to qualify the first.
  • Diary the Daybreak relaxation. Astra refuses proof-of-concept exploit creation today; OpenAI say less restrictive safeguards arrive in the coming weeks. Controls that pin to a model name will not notice.
  • Update any policy that assumes cyber safeguards tighten over time. Three frontier labs relaxed or staged a relaxation of the same class of control inside four days, and all three documented it.
  • If you are comparing vendors on these tables, read footnote 17. Two of the Claude comparisons are against Mythos, which OpenAI describe as Fable with fewer safeguards, not the shipping model.
  • Treat the SRE-Bench result as the operationally significant one. 88.0% first-attempt reverse engineering of binaries without source is the capability that changes what an attacker does with your compiled artefacts, and it is not contaminable by published write-ups.
  • Budget check if you use the API: $10 per million input tokens and $50 per million output, with Fast mode at twice the speed and twice the price.

The position

This is the most checkable frontier launch I have read this year, and I want to say that before the criticism, because the criticism is only possible because of it.

OpenAI built a benchmark designed to undercut their own headline number and published the much worse result next to it. They published a monitorability regression on launch day. They published four tables in which a competitor beats them. They footnoted the fact that two of their Claude comparisons use a less restricted variant. None of that was required.

The gap, as with Google yesterday, is between the prose and the tables. "The world's most intelligent and aligned model" and "a perfect score of 100%" are the sentences that will be repeated. "39.0%", "57.2% against 65.0%", and "harder to monitor than GPT-5.6 Sol's" are in the same document and will not be.

For a security team the practical position is narrow. A model that reverse engineers binaries at 88% first try, that found two live zero-days while being evaluated, and whose remaining cyber refusals are scheduled to be relaxed within weeks, is a capability change you should have an opinion about before your administrator is asked to tick the box. The opinion does not have to be no. It has to exist.

Sources

  1. PrimaryGPT-6 Astra: A new generation of intelligence, 3 September 2026. The Critical threshold statement, the full benchmark tables, the monitorability finding, the availability and pricing terms and the footnotesOpenAIaccessed 2026-09-03
  2. PrimaryOpenAI's demonstration of GPT-6 Astra modelling a house in Blender and converting it into a walkable Unreal Engine 5 scene, labelled by OpenAI as an edited excerpt of a demonstration runOpenAIaccessed 2026-09-03
  3. PrimaryThe Preparedness Framework, whose Critical cybersecurity threshold Astra is stated to meet, and which still does not name the function that produces the evidence its committees rule onOpenAIaccessed 2026-09-03

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.