Allen AI audited 16 benchmarks across 100 models. Three filed as safety turned out to be measuring reasoning.
One of the three is WMDP, the benchmark people reach for when they need a number for hazardous dual-use knowledge.
By Parminder Kumar Sharma · · 8 min read

If you cite a safety benchmark score in a governance document, this is the most useful thing published this week.
The Allen Institute for AI released BenchMIRT on 1 September, a method for auditing benchmarks at the level of individual questions. They ran it across 100 large language models, 16 benchmarks and more than 34,000 questions, and deliberately did not tell it which benchmark was supposed to measure what.
It recovered two dominant dimensions on its own: safety and general reasoning. Repeating the analysis from scratch produced the same two, which is the check that makes the result worth reading.
Most benchmarks landed where they are filed. Three did not, and one of the three is routinely cited as a proxy for dangerous knowledge.
Where each benchmark is filed, and what it loaded on
The three that did not land where filed
BBQ, which tests whether models rely on social stereotypes and is commonly grouped with safety benchmarks, "aligned much more strongly with general reasoning". Allen AI gives a concrete example: a question about a grandson and grandfather booking a taxi probes age bias, but it also requires the model to track who is who and reason from the evidence rather than from assumptions. So "a low BBQ score may partly reflect difficulty understanding or reasoning through certain questions, rather than safety behavior alone".
WMDP is the one that matters most for governance work, because it is the benchmark people reach for when they want a number for hazardous dual-use knowledge in biology, chemistry and cyber security. BenchMIRT "found that WMDP scores were more strongly associated with general reasoning than with safety".
The direction is counterintuitive and worth quoting exactly rather than paraphrasing: "Stronger general reasoning, however, was associated with lower WMDP scores, because the benchmark counts refusing or failing to provide the dangerous knowledge as the desired response."
HarmBench splits internally, which is the subtlest of the three. Its standard prompts and its contextual prompts both aligned with safety, as you would expect. But its copyright questions, requests to reproduce song lyrics and similar, "were more closely associated with general reasoning". One benchmark, one headline score, two different things being measured inside it.
What this does and does not mean
Allen AI is careful, and the caution belongs in any summary of the work: "These findings don't necessarily mean the benchmarks are flawed or incomplete. Rather, they show that a single benchmark score can combine several different signals."
That is right, and it is not a reason to relax. A score that combines several signals is fine as long as everyone reading it knows which signals. The problem arises the moment that score is lifted out of the paper and put into a risk register, a model card comparison or a procurement matrix, where it appears as a single number with a category name attached.
Two specific things now need care.
Any statement of the form "this model scores better on WMDP, therefore it holds less hazardous knowledge" is doing more work than the benchmark supports, because a meaningful part of what separates models on that benchmark is general reasoning rather than safety training or unlearning.
And any comparison of two models on BBQ is partly a comparison of their reasoning, so a bias claim resting on it is weaker than it looks.
The part that is immediately useful
Two results from the same work are practical rather than cautionary.
BenchMIRT ranked questions by how well they distinguished stronger from weaker models. Keeping only 10% of the questions across those 16 benchmarks "generally preserved nearly the same picture of which models were stronger or weaker on the underlying safety or reasoning capability as using the full set". Keeping 50% "often matched the full benchmark's measure of those capabilities even more closely".
If you run internal evaluations and the cost of a full benchmark sweep is why you run them quarterly rather than per release, that is a real finding: a well-chosen tenth of the questions may be enough to keep the ranking.
Second, it predicts unseen results. Given what it has learned about a model and about a question, BenchMIRT "correctly predicted whether a model would answer a held-out question correctly 79% of the time", against 70% for the naive approach of assuming a model performs on each question about as well as it does on the benchmark overall.
Nine points is a modest-sounding gain and it is the difference between a benchmark score and a model of why the score came out that way.
What the study actually did
| Element | Detail |
|---|---|
| Method | Multidimensional item response theory, applied at the level of individual questions rather than whole benchmarks |
| Corpus | 100 large language models, 16 benchmarks, more than 34,000 questions |
| Benchmark mix | Six measuring general reasoning, including MMLU-Pro, GPQA, MATH and BBH. Ten from the Olmo 3 safety suite, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP and XSTest |
| Blinding | The method was not told which benchmarks measured which capabilities. It recovered safety and general reasoning as the two dominant dimensions independently, and the same two re-emerged on repeat analysis |
| Compression result | Keeping 10% of questions generally preserved nearly the same ranking of models; keeping 50% often matched the full benchmark more closely |
| Prediction result | 79% accuracy predicting whether a model answers a held-out question correctly, against a 70% baseline |
Why this belongs on a security site
Because benchmark scores are becoming compliance artefacts.
Under the EU AI Act, under procurement questionnaires, and inside internal model-approval processes, teams are being asked to evidence that a model is safe enough for a use. The easiest available evidence is a benchmark number, and WMDP is one of the few that sounds like it is about hazard rather than about helpfulness.
This work says that number is partly measuring something else. Not that it is worthless, and not that the benchmark is badly built, but that its output carries a mixture, and the mixture is not what the name implies.
The honest response is not to stop using benchmarks. It is to stop citing a single score as though it were a measurement of one property, and to say which dimension a number is sensitive to when you put it in front of a decision.
Take this with you
If a benchmark score appears in your governance documents
- Stop using WMDP on its own as a dangerous-knowledge proxy. BenchMIRT found its scores more strongly associated with general reasoning than with safety, so a difference between two models is partly a difference in reasoning.
- Do not rest a bias claim on BBQ alone. It aligned more strongly with general reasoning, so a low score may reflect difficulty with the question rather than reliance on a stereotype.
- Treat a single benchmark as potentially several measurements. HarmBench’s copyright items loaded differently from its standard and contextual items, inside one score.
- Record which dimension a number is sensitive to whenever you cite it. That sentence is short, it is defensible, and it survives the question a regulator or an auditor will actually ask.
- Consider question-level selection for internal evaluation. Keeping 10% of questions generally preserved the ranking, which changes what you can afford to run per release rather than per quarter.
- Read this one at source rather than from a summary. The WMDP direction is easy to state backwards, and I did exactly that before reading it.
The position
The uncomfortable implication is not about these three benchmarks. It is that nobody was checking.
Sixteen widely used benchmarks, ten of them labelled as safety, and it took a purpose-built statistical method to discover that three of them are sensitive to something other than their label. That check was available at any point. The field simply took the names at face value, and so did every downstream document that cited the scores.
The lesson generalises past the specific findings. If a number is going to be treated as evidence, somebody has to have established what it measures, and "it is called a safety benchmark" is not that.
Sources
- PrimaryBenchMIRT, 1 September 2026. The method, the corpus, and the BBQ, WMDP and HarmBench findings, read in full rather than via a summaryAllen Institute for AIaccessed 2026-09-02
- PrimaryBenchMIRT code repository, linked from the announcementAllen Institute for AIaccessed 2026-09-02


