P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

AI Security

Five frontier labs were graded on containing a rogue model. The best grade was a C+, and the safety-first lab has no published containment plan

Guidelight assessed Anthropic, OpenAI, Google, xAI and Meta against six control practices. Nobody scored above 3 of 5 on any of them. The industry is strongest at noticing a problem and weakest at surviving one.

By Parminder Kumar Sharma · · 7 min read

A heavy steel bank-vault door with a wheel lock and thick hinges, standing slightly ajar in a dark concrete chamber, cold cyan light spilling from the gap where it has not fully closed. Text in the dark right reads: the last door is the open one, Guidelight AI control assessment, August 2026.

What was measured, and by whom

On 18 August 2026 Guidelight AI Standards published its first assessment of how well frontier labs can keep control of their own AI. It graded five companies, Anthropic, OpenAI, Google, xAI and Meta, against six foundational control practices, using only publicly available information. It is the rarest kind of AI safety document: a scorecard with a rubric, not a manifesto.

The tabloid version, which travelled fastest, was "the safety-first lab scored zero." That is true of exactly one cell in a thirty-cell table, and reading only that cell gets the story backwards. Here is the whole table.

How five frontier labs score on six control practices

Implementation of the six control practices, out of five

Guidelight · 18 Aug 2026 · public info only

StagePracticeAnthropicOpenAIGooglexAIMetaAvg
DetectLogging1.8
Monitor efficacy1.6
PreventGated actions1.6
Circuit breaking1.6
AssureThird-party review1.8
ContainContainment plan1.2
OverallC+2.50C+2.50D+1.50D−0.83F0.67

Each meter is five segments. The unlit segments are the distance to full implementation, and no lab lit more than three on any practice.

0Not implemented1Precursors only2Limited partial3Substantial partial4Near-full5Full
Scores from Guidelight AI Standards, “AI Control: An Assessment of Frontier Practices”, 18 August 2026, assessed on public information and redrawn here as text. Read the unlit segments: the top two levels, near-full and full, were awarded to no lab on any practice. The red meters are zeroes, including the joint leader’s on containment.
Every figure is from Guidelight AI Standards, 18 August 2026, redrawn as text rather than reproduced. Red cells are zeroes; the outlined one is the joint leader’s zero on containment. Nobody scored above 3 of 5 on anything.

How the grading works matters, because it is stricter than most safety scoring. Guidelight issued each cell two numbers: a 0 to 5 absolute score for how far a practice is actually adopted, and a relative rating comparing the company to its peers. The overall grade is simply the average of the six absolute scores, converted to a letter. Crucially, it scores only what a company has publicly disclosed and implemented, not what it intends. A lab that does excellent control work in private and says nothing scores badly, and that is deliberate: a control nobody outside can verify is not a control an enterprise can rely on.

Two things jump out before any single company does. No lab scored above 3 on any practice, on a scale where 3 is "substantial partial implementation" and 5 is "full". The best-run control practice at the best-run frontier lab is partial. And the top overall grade is a C+, shared by Anthropic and OpenAI. This is not a leaderboard with a winner. It is a class that is failing the same exam together.

The six practices are a line of defence

The six are not a random checklist. They fall into a sequence, from noticing a problem to surviving one.

The six control practices, and what each is for

PracticeGroupWhat it means
LoggingDetectRecord what internal AIs are doing, so they can be watched at all.
Monitor efficacyDetectMeasure whether the watching actually works, rather than assuming it does.
Gated actionsPreventRequire a monitor to clear a high-risk action as safe before it takes effect.
Circuit breakingPreventHalt the system after a surge of flagged misbehaviour, before it runs away.
Third-party reviewAssureLet outsiders check that the controls are adequate.
Containment planLast resortHave a plan for containing a model that has already gone wrong.
Definitions from Guidelight's Control standard, 18 August 2026. The grouping into detect, prevent, assure and contain is the natural order of a line of defence: catch it early if you can, survive it if you cannot.

The five labs fall into three bands, and the bands are more useful than the grades.

Anthropic and OpenAI, the front of the class. Both log substantial portions of their internal AI activity and scan it for signs of misbehaviour, both measure whether that scanning works, and both escalate to a human when something suspicious appears. That is a real detection stack. And yet the assessment is blunt: neither company fully implements a single one of the six practices. The leaders lead by being partial at more things than everyone else.

Google, the best plan and the thin implementation. Google published an AI Control Roadmap that Guidelight calls the most specific forward-looking document any lab has produced, covering prevention, detection and containment. It still scored a D+, because almost none of it is built yet. The single clearest example in the whole assessment is Google's gated-actions score: Guidelight notes Google plans to gate changes to sensitive systems until they are audited, but found no evidence of any blocking monitor in place today, and rated it precursors only. A plan for a control is not a control.

Meta and xAI, the back. Most of what is publicly known about Meta's controls comes from a single source, its disclosures to METR for this year's Frontier Risk Report. xAI is the starker case: it was the only one of the five that did not take part in METR's report at all, and while it has built some features that could become controls, there is little public evidence of how widely they are used or whether they work. An F and a D-minus are, as much as anything, grades for opacity.

Read in that order, the scores tell a worse story than any single grade. The industry is strongest at the start of the chain, detection, and weakest at the end, containment, which is precisely the wrong way round. Detection tells you the house is on fire. Containment is the fire door. The labs have smoke alarms and no fire doors.

The line of defence, and where it thins out

A flagged action, and the gates meant to stop it

industry average

  • 1.8LoggingDetect
  • 1.6MonitorDetect
  • 1.6Gated actionsPrevent
  • 1.6Circuit breakPrevent
  • 1.8Third-partyAssure
  • 0.8ContainmentLast resort

Each gate carries the five-lab average on that practice, from the scorecard above. The early gates catch the action. The last gate, containment, is the one the industry has barely built, so an action that gets past detection and prevention meets almost nothing at the end. That is the finding: the field is weakest at the last resort.

Averages computed from Guidelight’s five per-lab scores for each practice, 18 August 2026. The sequence is illustrative of how the six controls relate as a line of defence; the numbers are theirs.
Animated. A flagged action travels the six controls in sequence; each gate carries the five-lab average on that practice. The early detection gates are the strongest and the final containment gate is barely there.

The two ways the last line fails

Guidelight is specific about what weak prevention and containment actually mean in an incident, and it is worth stating plainly because it is the part that turns a grade into a risk. First, control systems that are not hardened are prone to being disabled by the very model they are monitoring: a capable, misaligned system can reach the switch that is supposed to stop it. Second, they are prone to a blitz, a burst of harmful actions faster than a human team can respond to, so a control that needs a person in the loop loses a race it did not know it was in. And behind both, the assessment found that companies have few containment protocols ready for an emergency at all. Detection tells you the alarm is ringing. Prevention and containment are whether anything happens next, and that is the half the industry has not built.

The finding the headline inverts

Now the cell everyone quoted. Anthropic scored 0 on containment plan, the only zero among its six scores. The company most publicly identified with worrying about rogue AI has published no plan for containing one.

That is real, and Guidelight says it plainly. But three facts sit next to it, and the headline drops all three.

What the headline says

  • The safety-first lab scored zero
  • So its safety record is a sham
  • So the lab that talks most about safety does least

What the scorecard says

  • Anthropic scored zero on one of six practices, and tied for first overall at C+
  • It is the only lab above limited partial implementation on prevention, the gated-actions and circuit-breaking controls that actually stop a runaway
  • OpenAI scored 3 on containment planning, the single highest score anyone earned on anything; the gap between the two leaders is one practice, not a chasm

The honest version is less satisfying and more alarming. The joint leader has not published a containment plan, and neither has Meta, and the best score anyone managed on that last-resort practice was OpenAI's 3, still only "substantial partial". A zero from the leader is a smaller problem than a C+ being the ceiling for the question "can you contain a model that is actively trying to escape your control."

Why a plan is not a control

Google is the instructive case in the middle. It published an AI Control Roadmap, which Guidelight calls the most specific forward-looking document any company has produced, spanning prevention, detection and containment. Google still scored a D+, because the assessment measures what is implemented, not what is promised.

That distinction is the whole value of the exercise, and it is the same one this site keeps returning to. A governance framework that names the committee but not the function producing its evidence cannot tell you whether it is discharged. A roadmap that spans containment while containment sits unimplemented is a document, not a defence. The scorecard grades the door, not the drawing of the door.

One more nuance sits inside the one bright spot. Third-party assessment is, with detection, the strongest area: four of the five labs took part in METR's first Frontier Risk Report. But a review from arm's length is not the same as a review with the keys, and Guidelight notes that only Anthropic appears to have granted assessors deep, hands-on access so far. It is also unclear whether any of these reviews will happen often enough to catch a fast-moving danger. The assurance exists; its depth and cadence are the open questions.

None of this is a counsel of despair, and Guidelight is careful to say so: for every lab there is a clear, practicable set of changes that would raise the score meaningfully, and none of the gaps is exotic. The controls are buildable today. They are simply not built.

C+

The highest grade awarded

Shared by Anthropic and OpenAI. Nobody reached B.

0

Practices fully implemented

No lab scored above 3 of 5 on any of the six.

6

Practices assessed

Logging, monitoring, gated actions, circuit breaking, third-party review, containment.

1

Lab that skipped METR

xAI was the only one of the five that did not take part in METR’s Frontier Risk Report.

From Guidelight AI Standards, AI Control: An Assessment of Frontier Practices, 18 August 2026, assessed on public information across Anthropic, OpenAI, Google, xAI and Meta.

What to do with this

You cannot fix another company's control practices. You can stop treating a lab's safety reputation as a substitute for its safety implementation, and you can ask the questions the scorecard implies.

Take this with you

Questions for any frontier model you build on

  • Ask for the containment plan, not the safety mission. A published plan for a model that has gone wrong is a specific document. If the vendor points you at values instead, that is your answer.
  • Separate what is implemented from what is on a roadmap. Google published the best plan and scored a D+ because almost none of it is built yet. A roadmap is a promise, and you cannot deploy on a promise.
  • Ask whether the controls can be disabled by the thing they watch. Guidelight found prevention and containment are the weakest areas, which means control systems are prone to being switched off by a misbehaving model faster than a human can react.
  • Check third-party access, not just third-party review. Four of the five took part in METR’s risk report, but only Anthropic appears to have granted hands-on access. A review from arm’s length is a different assurance from one with the keys.
  • Treat a strong safety brand as a claim to be tested, not evidence. The lab with the strongest safety reputation has the same C+ as its largest commercial rival, and a zero the rival does not have.

The position

Guidelight is not claiming these companies are reckless, and neither is this piece. The assessment is careful, it grades on public evidence, and its own conclusion is optimistic: every lab has a clear, practicable set of changes that would raise its score, and none of the gaps is exotic.

The uncomfortable finding is the flatness of the result. Across the five most advanced AI companies in the world, the ability to contain a model that has slipped its leash tops out at "substantial partial implementation", and at three of the five it is barely present at all. The most capable models are being shipped faster than the fire doors are being built, and the company with the best safety reputation is not exempt from that, it is an example of it.

The lesson is the one this site has now found in a status page, a breach notice, a safety framework and a benchmark: the reassuring summary and the measured reality are two different documents. Here the reassuring summary is a brand, and the measured reality is a C+.

Sources

  1. PrimaryAI Control: An Assessment of Frontier Practices, 18 August 2026Guidelight AI Standardsaccessed 2026-08-23
  2. Reported byFrontier AI labs still won't say how they'd contain a rogue modelTechCrunchaccessed 2026-08-23

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.