GLM-5.3, a downloadable AI model from Z.ai, can turn some known software bugs into working exploits: code that uses a flaw to make a program do something unintended. The awkward part is explaining how often.

In Anthropic's September 29, 2026 evaluation, it completed 50 of 410 ExploitBench attempts, about 12.2%. An earlier NIST assessment, published September 17, 2026, reported 61.1% on ExploitBench. Those figures look far apart. They do not measure the same outcome.

The useful distinction is between finishing an exploit and earning credit for progress toward one. Neither percentage tells us how often the model would successfully attack a real organization.

One benchmark, two different scoreboards

Anthropic counted complete successes per attempt, using known V8 bugs in isolated offline environments. V8 is Chrome's JavaScript engine; these were not attacks on arbitrary targets.

NIST's Center for AI Standards and Innovation used a different scoring rule. Its 41 ExploitBench tasks were graded on a 16-point scale, retaining the best of three attempts for each task. The reported 61.1% corresponds to an average score shown as 9.8 out of 16, rounded. It is not a count of fully successful exploits.

Think of an exam that awards points for intermediate work. An average mark of 61% does not mean 61% of students solved every step. Counting only complete solutions answers another question. Keeping each student's best attempt changes the comparison again.

There is no honest shortcut for converting one published number into the other. A claim that the model suddenly became five times better or worse would require comparable scoring and underlying results, not two percentages beside the same benchmark name.

A working exploit is still a bounded result

The ExploitBench paper, a May 13, 2026 preprint, explains the distinction. Its grading tracks 16 measurable capabilities, from early progress toward an exploit to executing code chosen by the attacker. Intermediate progress is informative. It is not interchangeable with reaching the final objective.

The tasks also start from known, already patched browser-engine vulnerabilities. The authors do not grade whether a proof of concept becomes a usable attack or works reliably across changing environments. That makes the benchmark useful for measuring a difficult technical capability, but narrower than finding a fresh flaw and reliably exploiting a live system.

Anthropic's separate safeguard tests need their own boundary. Some tested conditions produced high engagement with malicious requests, but those experiments used simulated tool responses. No generated code ran or external system was reached in that simulation. Willingness to attempt an attack is not a successful intrusion.

That caveat does not erase the offline exploit results. It prevents two different experiments from being folded into one dramatic claim.

Open access matters even below the frontier

NIST assessed GLM-5.3 as the strongest open-weight model it had evaluated, while still below contemporary U.S. frontier models. Its comparisons included restricted-access models and, where applicable, tests with safeguards disabled. A capability leaderboard is not a menu of what every customer can actually use.

Z.ai's public model files and deployment documentation establish the other half of the story: the weights, the model's learned parameters, are downloadable. That does not make deployment cheap or effortless. It does make the access question different from a model available through a provider's controlled service.

Nor should 12.2% sound reassuring by itself. An attacker does not need to win a benchmark. One successful attempt can matter, and retries or human assistance can change a workflow. These reports do not establish the probability of success in those circumstances.

Anthropic is also a competing model supplier. Its findings deserve scrutiny without treating its policy conclusions as independent verdicts. Neither assessment measures a rise in real-world cybercrime caused by GLM-5.3.

For a security team assessing the evidence, the next useful question is concrete: which tasks succeeded, with what tools and budget, against what targets? That is the distinction between an evaluation result and an auditable claim. Downloadable exploit capability is consequential. The percentage alone is not the risk assessment.