GLM-5.3 beat Anthropic at finding bugs by 0.7 points. Z.ai's own table shows it losing the exploit benchmark by 23.6
The vendor published four benchmarks and the coverage used one. Finding a flaw and weaponising it are different capabilities, and the gap between them is the number that matters to anyone defending an estate.
By Parminder Kumar Sharma · · 7 min read

What Z.ai actually published
Z.ai released GLM-5.3 on 14 August 2026 and the coverage settled quickly on one sentence: a Chinese open-weight model is now a better bug-finder than Anthropic's and OpenAI's. The Register ran it. So did most of the trade press.
The claim comes from a benchmark table Z.ai published in its own model documentation. That table has four rows, and the coverage used one of them.
GLM-5.3 against the frontier, from Z.ai's own documentation
| Benchmark | GLM-5.3 | GLM-5.2 | Mythos 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| CyberGym, finding vulnerabilities | 84.5% | 77.2% | 83.8% | 83.6% |
| ExploitBench, building working exploits | 54.4% | 24.4% | 78.0% | 76.5% |
| ExploitGym, tasks completed in 2 hours | 105 | 29 | 181 | not published |
| ExploitGym, tasks completed in 6 hours | 130 | 39 | 247 | not published |
GLM-5.3 leads on the first row by 0.7 percentage points. It trails on the second by 23.6 points, and completes 47 per cent fewer ExploitGym tasks than Mythos 5 over six hours.
Every one of those numbers is Z.ai's. This is the unusual case where a vendor published the material that undercuts its own press coverage, in the same table, on the same page.
Finding is not exploiting
Finding is not exploiting
Z.ai’s own figures · GLM-5.3
The benchmark every headline used
CyberGym
Can the model find the flaw
GLM-5.3 84.5%Mythos 5 83.8%+0.7 points
The three in the same table that nobody quoted
ExploitBench
Can it turn the flaw into a working exploit
GLM-5.3 54.4%Mythos 5 78.0%23.6 points behind
ExploitGym, 2 hours
How many it completes under time pressure
GLM-5.3 105 tasksMythos 5 181 tasks42% fewer
ExploitGym, 6 hours
Whether more time closes the gap
GLM-5.3 130 tasksMythos 5 247 tasks47% fewer
A model that finds flaws at the frontier and cannot reliably weaponise them is a different threat from one that does both. The headline claim is true of the left-hand column only.
The distinction is not pedantry, and it is the one a defender should care about most.
CyberGym measures whether a model can locate a flaw in real code. ExploitBench and ExploitGym measure whether it can turn that flaw into something that actually works. Those are different problems: the first is comprehension, the second is engineering under constraints, and the gap between them is where most vulnerability research dies.
A model at the frontier for discovery and well behind it for weaponisation changes your world in a specific way. It means the volume of reported flaws goes up, your triage queue grows, and the proportion of them that anyone can turn into a working attack does not move nearly as fast. That is a resourcing problem before it is a threat.
The number that should worry you is not on the chart the headlines used
ExploitBench: GLM-5.2 to GLM-5.3, same base model
Twenty-four point four to fifty-four point four, in one release, without retraining the base model. ExploitGym over two hours went from 29 tasks to 105, which is a factor of 3.6.
Argue about whether 54.4 beats 78.0 all you like. The interesting quantity is the derivative, and Z.ai got it by adding vulnerability-discovery data to a post-training mix. That is a cheap, repeatable intervention that any lab with the data can run, and it is a far better predictor of where this sits in six months than today's ranking.
Z.ai's own account of the result is worth quoting, because it is candid about not having aimed at what they hit:
As part of post-training, we introduced vulnerability discovery data and environments into the training mix.
They expected a model better at finding and reasoning about vulnerabilities. What they describe getting is a model that reasons across multi-stage exploitation chains. The capability arrived ahead of the intention, which is a governance fact rather than a technical one.
Nobody can check any of this yet
Every figure above is vendor-reported. The open weights are not out; the expected release is around 28 August 2026, and until then no independent party can reproduce a single row of that table.
That is not an accusation. Z.ai publishing the benchmarks it loses is the behaviour of a lab that expects to be checked. But there is a standing rule for this site and it applies without exception: a benchmark you cannot reproduce is a claim, not a measurement.
The real-world figures need the same reading. Z.ai reports 2,436 vulnerabilities found across 269 open-source projects, of which 1,097 were medium-to-high severity, from collaborations with Chinese security teams. The Register's Simon Sharwood makes the right observation about that: the finding rate may say as much about which codebases were pointed at as about the model doing the pointing. Nobody has run Mythos 5 over the same 269 projects.
What this changes
Take this with you
Practical, and mostly about volume rather than threat
- Plan for more reported vulnerabilities, not more exploited ones. Discovery has reached the frontier in an open-weight model, and weaponisation has not. Your triage queue is the thing under pressure first.
- Ask whoever sends you vulnerability reports whether a model found them, and treat that as neutral rather than disqualifying. The relevant question is whether the report reproduces, which was always the question.
- Watch the 28 August weights release rather than this week’s headline. The moment the weights are public, three things become possible at once: independent verification, fine-tuning by anybody, and running it offline with no provider log of what you asked it.
- Note that an open-weight model with frontier-adjacent discovery capability has no refusal policy you can rely on, because whoever holds the weights sets it. This is the same point made here about the Daybreak refusal rate being a dial, except that here the dial ships with the download.
- Do not restructure a security programme around a 0.7 point benchmark margin, in either direction. The trajectory from 24.4 to 54.4 in one post-training run is the number worth briefing upward.
The position
The headline is defensible and almost useless. GLM-5.3 does lead CyberGym, by a margin narrow enough that a different sample would probably reverse it, and the same table shows it losing the other three comparisons by margins nothing would reverse.
What is actually true is more interesting than the headline and worse for anyone hoping this plateaus. An open-weight model closed most of the discovery gap to the frontier, and it did so through post-training data rather than a new base model. That is the cheap path, and cheap paths get walked repeatedly.
It also lands in the same week as Anthropic disclosing that a safety control ran with its logging disabled for eleven months, and both stories point at the same uncomfortable place. Capability is arriving faster than the instrumentation that is supposed to observe it, and in one of these cases the lab said so about itself.
The line I would give a board is short. Nothing about your defensive posture should change because of a 0.7 point benchmark lead. Quite a lot should change on the day those weights are downloadable, and that date is 28 August, not today.
Sources
- PrimaryGLM-5.3 overview and benchmark results, developer documentationZ.aiaccessed 2026-08-17
- Reported byChinese AI company Zhipu claims its new model is a better bug-finder than Anthropic, OpenAIThe Registeraccessed 2026-08-17


