A model whose weights anyone can download built end-to-end exploits in 50 of 410 attempts on a browser-vulnerability benchmark. Anthropic’s Claude Mythos Preview succeeded in 56 attempts under the same test. That narrow gap is specific to the benchmark; it does not show that either model can reliably compromise a real browser.
Anthropic’s GLM-5.3 evaluation goes beyond scores. It argues that a capability Anthropic restricted to vetted defenders through its own Mythos release has appeared in a publicly downloadable model whose refusals can be bypassed or removed. An earlier, independent assessment by NIST also found GLM-5.3 unusually capable for an open-weight model, while placing it below the current U.S. frontier on NIST’s broader set of tests.
The Exploit Result Is Specific, and Significant
Anthropic tested GLM-5.3 on ExploitBench, which asks a model to develop exploits from known vulnerabilities in V8, the JavaScript engine used by Chrome. Its headline measure was whether an attempt produced an end-to-end exploit. GLM-5.3 succeeded 50 times in 410 attempts, compared with 56 for Claude Mythos Preview, or roughly 12% and 14% of attempts respectively.
The runs were isolated and sandboxed, with offline targets set up for evaluation. The models were not attacking arbitrary websites or users. A known vulnerability and a defined testing environment remove obstacles that a real attacker would still have to address. GLM-5.3 sometimes completed a demanding exploit-development task under those conditions, but the result does not establish a 12% success rate for real-world attacks.
A second Anthropic test points in the same direction, with a smaller success rate. On 100 selected tasks from an internal benchmark involving vulnerable open-source software, GLM-5.3 achieved the top outcome, full control-flow hijack, in 4% of trials. Mythos Preview reached 6%. Anthropic says two earlier models it tested, GLM-5.2 and Claude Opus 4.6, recorded no such successes on those tasks.
Most attempts failed, and GLM-5.3 did not match Mythos Preview across the board. Still, in Anthropic’s setup, an open-weight model repeatedly converted software flaws into working exploits. It went beyond locating bugs or producing code that crashes a program.
Public Weights Make Refusals a Different Kind of Defense
GLM-5.3 does have built-in refusals, according to Anthropic. How much they constrain someone who can prompt the model repeatedly or alter a local copy is another question.
In Anthropic’s simulated harmful-request tests, deceptive cover stories led GLM-5.3 to engage with the requests 64% of the time. Prefilled reasoning raised that figure to 92%. An altered version of the model engaged 100% of the time. Anthropic says safeguarded Claude models recorded zero engagement in those same tests.
“Engagement” is the important qualification. These percentages measure how often the models proceeded with simulated requests using fake tooling, not how often they completed a real intrusion. They reflect Anthropic’s chosen prompts and evaluation conditions, so they are evidence about safeguards in that setup, not a universal bypass rate.
The altered model illustrates a limitation specific to public weights. Anthropic used abliteration, a technique that edits a model to reduce its tendency to refuse requests. The company reports that, across three public harmful-request benchmarks, GLM-5.3’s refusal rate fell from above 90% to about 3%, 2% and 12%, depending on the benchmark. Anthropic says the edit took roughly 2,200 GPU hours and cost about $4,400 in compute. That is no one-click trick. Once a capable user has downloaded the weights, though, the model provider cannot prevent that user from making the edit.
Anthropic also tested whether the edit damaged useful performance. It reports no change on GPQA-Diamond, a general scientific-knowledge evaluation, and a decline of a few percentage points on the tested subset of CyberGym tasks. Every offensive capability may not survive unchanged. Even so, refusals embedded in downloadable weights are a weaker control than restrictions on who can access a model in the first place.
NIST Finds the Same Direction of Travel, Not Frontier Parity
NIST’s Center for AI Standards and Innovation, or CAISI, published its assessment on September 17, before Anthropic’s September 29 report. It called GLM-5.3 the most cyber-capable open-weight model it had evaluated. NIST also found its capabilities “significantly lower” than those of current U.S. frontier models and estimated an aggregate lag of about four months across its cyber benchmarks.
Both findings can be true. “Most capable open-weight model” describes GLM-5.3’s position among publicly released weights in NIST’s testing. The frontier comparison includes stronger U.S. models, some available only through trusted access. NIST tested U.S. models with cyber safeguards disabled where applicable, making that comparison about underlying capability rather than what an ordinary user could obtain from a guarded service.
NIST’s individual results cannot be collapsed into Anthropic’s 50-out-of-410 figure. On NIST’s ExploitGym userspace tasks, GLM-5.3 succeeded on 9.4% of evaluated tasks, compared with 44.4% for the best U.S. frontier result. On NIST’s private OSS-Fuzz evaluation, the figures were 7.7% and 23.2%. These tasks, harnesses and success criteria differ from Anthropic’s ExploitBench runs.

Even the two organizations’ ExploitBench figures measure different things. Anthropic counted end-to-end successes across attempts. NIST reported a score based on a 16-point grading scale, using each task’s best of three attempts. Its reported 61.1% for GLM-5.3 is not a claim that the model completed exploits in 61.1% of attempts.
NIST independently strengthens the case that GLM-5.3 marks a step up for accessible models. It does not validate Anthropic’s safeguard-bypass percentages, which come from Anthropic’s separate experiments.
Z.ai Built Cyber Training Into the Release
The result is consistent with how Z.ai presents the model. In its GLM-5.3 release announcement, the developer says it used vulnerability-discovery environments in training and reports substantial gains on CyberGym and ExploitBench. Those are Z.ai’s claims about its own model, not independent measurements of attack success.
NIST says Z.ai first released GLM-5.3 on August 14, 2026, and published its weights two weeks later. That sequence matters more to the security implications than the model’s position on any one chart. An API provider can change access rules for a hosted model. Published weights let others run copies independently, including copies with modified refusal behavior.
Frequently Asked Questions
4 questions
1Did GLM-5.3 Perform as Well as Claude Mythos Preview?
GLM-5.3 came close on one measure in Anthropic’s sandboxed ExploitBench test: it completed end-to-end exploits in 50 of 410 attempts, compared with 56 for Claude Mythos Preview. That does not mean the models are equally capable overall. GLM-5.3 also trailed Mythos Preview on Anthropic’s selected open-source exploitation tasks, and NIST placed GLM-5.3 below the current U.S. frontier across its broader cyber assessment.
Sources
- GLM-5.3 evaluationanthropic.com
- assessment by NISTnist.gov
- GLM-5.3 release announcementz.ai





