P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

AI Security

GLM-5.3 beat Anthropic at finding bugs by 0.7 points. Z.ai's own table shows it losing the exploit benchmark by 23.6

The vendor published four benchmarks and the coverage used one. Finding a flaw and weaponising it are different capabilities, and the gap between them is the number that matters to anyone defending an estate.

By Parminder Kumar Sharma · · 7 min read

A slender steel lock pick lying diagonally across the closed brass shackle of a padlock, macro close-up on a near-black surface, the pick edge-lit in cyan and a small crimson highlight on the lock body.

What Z.ai actually published

Z.ai released GLM-5.3 on 14 August 2026 and the coverage settled quickly on one sentence: a Chinese open-weight model is now a better bug-finder than Anthropic's and OpenAI's. The Register ran it. So did most of the trade press.

The claim comes from a benchmark table Z.ai published in its own model documentation. That table has four rows, and the coverage used one of them.

GLM-5.3 against the frontier, from Z.ai's own documentation

BenchmarkGLM-5.3GLM-5.2Mythos 5GPT-5.6 Sol
CyberGym, finding vulnerabilities84.5%77.2%83.8%83.6%
ExploitBench, building working exploits54.4%24.4%78.0%76.5%
ExploitGym, tasks completed in 2 hours10529181not published
ExploitGym, tasks completed in 6 hours13039247not published
Read from Z.ai's published GLM-5.3 documentation on 17 August 2026. Mythos 5 is Anthropic's model and GPT-5.6 Sol is OpenAI's. GLM-5.2 is included because the same base model underlies both, with the difference coming from post-training alone.

GLM-5.3 leads on the first row by 0.7 percentage points. It trails on the second by 23.6 points, and completes 47 per cent fewer ExploitGym tasks than Mythos 5 over six hours.

Every one of those numbers is Z.ai's. This is the unusual case where a vendor published the material that undercuts its own press coverage, in the same table, on the same page.

Finding is not exploiting

Finding is not exploiting

Z.ai’s own figures · GLM-5.3

The benchmark every headline used

  • CyberGym

    Can the model find the flaw

    GLM-5.3 84.5%Mythos 5 83.8%

    +0.7 points

The three in the same table that nobody quoted

  • ExploitBench

    Can it turn the flaw into a working exploit

    GLM-5.3 54.4%Mythos 5 78.0%

    23.6 points behind

  • ExploitGym, 2 hours

    How many it completes under time pressure

    GLM-5.3 105 tasksMythos 5 181 tasks

    42% fewer

  • ExploitGym, 6 hours

    Whether more time closes the gap

    GLM-5.3 130 tasksMythos 5 247 tasks

    47% fewer

A model that finds flaws at the frontier and cannot reliably weaponise them is a different threat from one that does both. The headline claim is true of the left-hand column only.

All figures published by Z.ai in its own GLM-5.3 documentation, read 17 August 2026. Mythos 5 is Anthropic’s model; GPT-5.6 Sol scores 83.6% on CyberGym and 76.5% on ExploitBench, and is omitted here only to keep the comparison to two columns.

The distinction is not pedantry, and it is the one a defender should care about most.

CyberGym measures whether a model can locate a flaw in real code. ExploitBench and ExploitGym measure whether it can turn that flaw into something that actually works. Those are different problems: the first is comprehension, the second is engineering under constraints, and the gap between them is where most vulnerability research dies.

A model at the frontier for discovery and well behind it for weaponisation changes your world in a specific way. It means the volume of reported flaws goes up, your triage queue grows, and the proportion of them that anyone can turn into a working attack does not move nearly as fast. That is a resourcing problem before it is a threat.

The number that should worry you is not on the chart the headlines used

ExploitBench: GLM-5.2 to GLM-5.3, same base model

GLM-5.224.4%
GLM-5.354.4%
GPT-5.6 Sol76.5%
Mythos 578%
The jump came from post-training alone. Z.ai states that it introduced vulnerability discovery data and environments into the training mix. It is still 23.6 points behind Mythos 5, and it more than doubled in one release.

Twenty-four point four to fifty-four point four, in one release, without retraining the base model. ExploitGym over two hours went from 29 tasks to 105, which is a factor of 3.6.

Argue about whether 54.4 beats 78.0 all you like. The interesting quantity is the derivative, and Z.ai got it by adding vulnerability-discovery data to a post-training mix. That is a cheap, repeatable intervention that any lab with the data can run, and it is a far better predictor of where this sits in six months than today's ranking.

Z.ai's own account of the result is worth quoting, because it is candid about not having aimed at what they hit:

As part of post-training, we introduced vulnerability discovery data and environments into the training mix.

They expected a model better at finding and reasoning about vulnerabilities. What they describe getting is a model that reasons across multi-stage exploitation chains. The capability arrived ahead of the intention, which is a governance fact rather than a technical one.

Nobody can check any of this yet

Every figure above is vendor-reported. The open weights are not out; the expected release is around 28 August 2026, and until then no independent party can reproduce a single row of that table.

That is not an accusation. Z.ai publishing the benchmarks it loses is the behaviour of a lab that expects to be checked. But there is a standing rule for this site and it applies without exception: a benchmark you cannot reproduce is a claim, not a measurement.

The real-world figures need the same reading. Z.ai reports 2,436 vulnerabilities found across 269 open-source projects, of which 1,097 were medium-to-high severity, from collaborations with Chinese security teams. The Register's Simon Sharwood makes the right observation about that: the finding rate may say as much about which codebases were pointed at as about the model doing the pointing. Nobody has run Mythos 5 over the same 269 projects.

What this changes

Take this with you

Practical, and mostly about volume rather than threat

  • Plan for more reported vulnerabilities, not more exploited ones. Discovery has reached the frontier in an open-weight model, and weaponisation has not. Your triage queue is the thing under pressure first.
  • Ask whoever sends you vulnerability reports whether a model found them, and treat that as neutral rather than disqualifying. The relevant question is whether the report reproduces, which was always the question.
  • Watch the 28 August weights release rather than this week’s headline. The moment the weights are public, three things become possible at once: independent verification, fine-tuning by anybody, and running it offline with no provider log of what you asked it.
  • Note that an open-weight model with frontier-adjacent discovery capability has no refusal policy you can rely on, because whoever holds the weights sets it. This is the same point made here about the Daybreak refusal rate being a dial, except that here the dial ships with the download.
  • Do not restructure a security programme around a 0.7 point benchmark margin, in either direction. The trajectory from 24.4 to 54.4 in one post-training run is the number worth briefing upward.

The position

The headline is defensible and almost useless. GLM-5.3 does lead CyberGym, by a margin narrow enough that a different sample would probably reverse it, and the same table shows it losing the other three comparisons by margins nothing would reverse.

What is actually true is more interesting than the headline and worse for anyone hoping this plateaus. An open-weight model closed most of the discovery gap to the frontier, and it did so through post-training data rather than a new base model. That is the cheap path, and cheap paths get walked repeatedly.

It also lands in the same week as Anthropic disclosing that a safety control ran with its logging disabled for eleven months, and both stories point at the same uncomfortable place. Capability is arriving faster than the instrumentation that is supposed to observe it, and in one of these cases the lab said so about itself.

The line I would give a board is short. Nothing about your defensive posture should change because of a 0.7 point benchmark lead. Quite a lot should change on the day those weights are downloadable, and that date is 28 August, not today.

Sources

  1. PrimaryGLM-5.3 overview and benchmark results, developer documentationZ.aiaccessed 2026-08-17
  2. Reported byChinese AI company Zhipu claims its new model is a better bug-finder than Anthropic, OpenAIThe Registeraccessed 2026-08-17

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.