P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

AI Security

Anthropic ran 133 million contractor exchanges with its biological filters off. The same flag switched off the logging

Eleven months, roughly 50,000 people, no customer traffic, and no alert at any point. The control that failed was wired to the same switch as the telemetry that would have reported it, so the failure looked exactly like success.

By Parminder Kumar Sharma · · 8 min read

An industrial toggle switch with a knurled metal bezel set into a dark brushed panel, thrown to one side, macro close-up, with a small unlit indicator pin beside it and a cyan edge light down the left of the panel.

What Anthropic disclosed about itself

Buried in section 4.5.8.2.2 of Anthropic's August 2026 Risk Report, under a heading that does not soften anything, is this:

Since our last Risk Report, we discovered that our bio classifiers were not running for a substantial amount of time on our human feedback vendor traffic.

The scale of "a substantial amount of time" follows a paragraph later, and the numbers are worth setting out in full because most coverage has quoted only the largest one.

The gap, as Anthropic describes it

MeasureValue
Period the biological classifiers did not blockMay 2025 to April 2026, roughly eleven months
What May 2025 marksTheir first deployment of models carrying CB safeguards
People with accessRoughly 50,000 contractors
Who vetted themTheir vendors alone, many of which by Anthropic’s account could not screen out even CB-1 threat actors
ExchangesRoughly 133 million
Access typeThe vast majority had general open-ended conversation, not fixed rating interfaces
Customers affectedNone. This was internal human feedback collection, not the production API
Transcripts retainedAlmost all, excepting one platform at roughly 1% where the user never pressed submit
All figures from the August 2026 Risk Report, section 4.5.8.2.2, read 17 August 2026. CB is Anthropic's abbreviation for chemical and biological.

Two things are worth saying immediately, because they will otherwise get lost. No customer traffic was involved, and Anthropic found and published this themselves. Very few laboratories publish a document like this at all, and the ones that do not are not safer.

The finding is the second clause of one sentence

The number everyone is quoting is 133 million. The number is not the finding. This is:

This traffic was controlled by a flag meant only for internal use, which disabled not only the blocking behavior of the classifiers but also the logging of their flags; as such, traffic that would otherwise have been flagged was not recorded or propagated to any review mechanisms.

Blocking and logging were wired to the same switch.

One flag, two switches

Anthropic Risk Report · Aug 2026

  1. What was intended

    Blocking on
    Logging on

    Zero flags, and the zero is evidence

    Classifiers run on the traffic. A quiet dashboard means the queries were clean.

  2. May 2025 to April 2026

    Blocking off
    Logging off

    Zero flags, and the zero means nothing

    One flag disabled both. Traffic that would have been flagged was never recorded or sent to any review mechanism.

  3. If only blocking had been disabled

    Blocking off
    Logging on

    Flags accumulating, unactioned

    Still a serious gap, but a visible one. Somebody looking at the queue would have asked why nothing was being blocked.

The first two columns look the same from outside, and that is the entire failure. Eleven months of a disabled safety control produced exactly the dashboard a working one produces.

Drawn from Anthropic’s own description in the August 2026 Risk Report. The third column is not what happened; it is included because the difference between it and the second is the design lesson.

That coupling is the whole story. A safety control that fails loudly gets fixed on a Tuesday. A safety control that fails silently is indistinguishable from one that is working, and it stays that way for as long as nobody thinks to test the thing that would have raised the alarm.

Eleven months of a disabled classifier produced exactly the dashboard a functioning classifier produces: nothing. There was no alert to miss, no queue backing up, no metric bending. The absence of flags meant the queries were clean, right up until it meant the opposite.

What the review found, and what it does not settle

Anthropic did the work rather than asserting the conclusion. They ran Claude Sonnet 5 as a prompted classifier over every human turn sent in the affected period, scoring against a constitution for CB-1 biological harm.

1,197 transcripts flagged high for biological harm, by who sent them

Internal Anthropic teams, on the same infrastructure757 transcripts
Commissioned red-teaming exercises378 transcripts
Everything else, all manually reviewed62 transcripts
From the same section. The residual category, 62 transcripts, is the only part of this that was neither Anthropic staff nor a commissioned red-teaming exercise. All 62 were reviewed by hand.

They reviewed all 62 by hand, plus a random 30 of the red-teaming transcripts, and did not observe clearly concerning misuse that would have given a threat actor meaningful uplift. They found "a handful of potentially dual-use conversations". The longest external transcript containing any flagged query ran to 240,000 tokens; the 75th percentile was 2,102 tokens, which is to say most of these exchanges were short.

Their reasoning for assessing the risk as low is sound and worth repeating: anyone seriously trying to build a bioweapon through this channel could not have sent more than a few dozen detailed high-harm queries without human review very likely picking up at least one conversation.

Then there is footnote 75, and it does more damage than the body text it is attached to.

Read those two things together. The reassurance is that 378 of the flags came from red-teamers rather than from anyone dangerous. The footnote is that a determined threat actor could probably have become a red-teamer. That does not make the conclusion wrong, and Anthropic clearly know it, which is why the footnote is there at all. It does mean the category that absorbs most of the flags is not as clean a category as it looks in a chart.

The other incident in the same report

A few paragraphs earlier there is a second disclosure, and reading them together is what turns an incident into a pattern.

Two failures in one document, both through the vendor layer

  1. May 2025

    Biological classifiers deployed

    First models shipped with CB safeguards. The human feedback platforms are not covered, and the flag that excludes them also silences their logging.

  2. Undated

    Contractors obtain an API key

    A small number of data-labelling contractors exploit a flaw in a collection platform to reach models outside their assigned tasks, including Mythos Preview. The path is open for several weeks, about two of them with Mythos Preview reachable.

  3. Within 90 minutes

    That one is contained

    Systems offline, vendor credentials disabled, the vector closed the day it was reported. No model weights, no customer data, no breach of core systems.

  4. April 2026

    The classifier gap ends

    Eleven months after it began. Remediated, reviewed, and disclosed.

  5. Aug 2026

    Both published

    In a report that also states the discovery raises the likelihood of similar unknown issues.

Both from the August 2026 Risk Report. Neither touched the production API or customer data, and both ran through the same third-party contractor environment.

The contrast between them is instructive. The API key incident was contained within ninety minutes of somebody learning about it, which is genuinely fast. The classifier gap ran for eleven months. The difference is not competence or urgency. It is that one of them generated a report and the other generated silence.

The sentence that should have led

Anthropic wrote the conclusion themselves, and it is a more serious statement than any headline about 133 million:

The discovery of this gap, however, leads us to believe that there is an increased likelihood of other, similar issues unknown to us.

That is a laboratory saying, in a published document, that its confidence in its own safety instrumentation has gone down. It is the correct inference, it is uncomfortable, and it generalises well past Anthropic. If a control can be disabled by a flag that also disables its telemetry, then the number of such controls you have is unknown by construction, and no amount of reviewing the ones you know about changes that.

What to check this week

Take this with you

This generalises past model safety, and the first item is the one with teeth

  • Find every control in your estate where the same setting governs both enforcement and logging. Feature flags, WAF rules in monitor mode, DLP policies, agent guardrails and EDR exclusions are the usual places. If turning one off also turns off its evidence, that control cannot be audited and its silence proves nothing.
  • Write down what a working control is supposed to produce, then alert on the absence of it. A queue that should receive flags and receives none for a week is an incident, not a quiet week. Almost no one alerts on zero.
  • Treat contractor and vendor platforms as production for safety purposes. Both failures in this report ran through the third-party environment used for model development, which is exactly where the controls applied to customer traffic were assumed not to be needed.
  • Ask how your vendors screen the people they assign to you, and specifically whether that screening would stop somebody who wanted the access rather than the work. Anthropic say plainly that theirs would not have, prior to April 2026.
  • Periodically prove a control fires rather than assuming it. A synthetic input that should trip the classifier, the filter or the rule, run on a schedule, would have surfaced this in days rather than months.
  • Read the footnotes of any safety report you are relying on, including this one. The load-bearing qualification in this disclosure is in a footnote, not in the summary.

The second item is the cheapest and the one almost nobody has. Most monitoring is built to alert on events, and this failure mode is defined by the absence of them.

The position

Publishing this was the right call and it should be said clearly, because the incentive to bury it was enormous and the reporting will punish them for candour. A laboratory that discloses an eleven-month failure in its own biosecurity instrumentation is behaving better than one with a clean public record, and anyone drawing the opposite lesson is encouraging the next lab to keep quiet.

But the substance is not reassuring, and the reason is not the 133 million exchanges. It is that the mechanism which failed was the mechanism that would have reported the failure, and Anthropic's own conclusion is that they now expect more of these.

This is the fifth briefing in a week where a control worked exactly as specified while doing none of the thing people believed it was doing. An antivirus named ransomware and could not remove it. Passkeys stop phishing and were heard to stop theft. A zero-day rated Important was the only one under active exploitation.

The version here is the purest of them, because there was no gap between specification and marketing at all. The classifiers were specified correctly, built correctly, and simply not running, and the only reason anybody found out is that Anthropic went looking without being prompted by an incident.

Ask yourself the honest question that follows. In your own estate, which control would you discover was switched off only because you went looking, rather than because it told you?

Sources

  1. PrimaryRisk Report, August 2026, redacted, section 4.5.8.2.2Anthropicaccessed 2026-08-17
  2. Reported byAnthropic ran 133 million contractor chats with its bioweapon filters offThe Next Webaccessed 2026-08-17
  3. Reported byAnthropic's bio-weapons filter was down for nearly a yearThe Decoderaccessed 2026-08-17

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.