On 29 September 2026 Anthropic’s Frontier Red Team published a report saying that GLM-5.3, an open-weight model from Zhipu AI (known outside China as Z.ai), can build working software exploits about as well as Anthropic’s own gated model from April, and that its built-in refusals are easy to remove. The capability finding is well supported, including by Z.ai itself. The harder questions are how the comparison is framed, what the safeguard test does and does not show, and who benefits from each reading. This report takes those in turn.
Facts at a glance:
- The report: “GLM-5.3 and the spread of advanced cyber capabilities”, Anthropic Frontier Red Team, 29 September 2026.
- The model: GLM-5.3, released on Z.ai’s application programming interface (API) on 14 August 2026. Weights were published about two weeks later.
- Size: about 753 billion parameters by the Hugging Face file metadata. A smaller GLM-5.3-Flash has about 321 billion.
- Licence: a custom GLM-5.3 licence, close to MIT, with one added condition for very large hosting companies. GLM-5.3-Flash is MIT.
- Independent test: the Center for AI Standards and Innovation (CAISI), part of the United States National Institute of Standards and Technology (NIST), published its assessment on 17 September 2026.
- Headline numbers: 50 full exploits in 410 attempts for GLM-5.3 against 56 for Claude Mythos Preview on one benchmark, and refusals bypassed in 64% to 100% of simulated trials.
- Z.ai’s reply: none that we could find as of 3 October 2026.
What Anthropic claims
The report makes four claims.
One: GLM-5.3 builds exploits end to end. On ExploitBench, a public benchmark built on 41 known bugs in V8, the JavaScript engine in Chrome, GLM-5.3 produced a full exploit in 50 of 410 attempts. Claude Mythos Preview, with safeguards switched off, did so in 56 of 410. On Anthropic’s internal binary exploitation benchmark, GLM-5.3 took full control of the program in 4 of 100 randomly chosen tasks and Mythos Preview in 6. Earlier models, including Claude Opus 4.6 and GLM-5.2, did not succeed on any of those tasks.
Two: it works in a researcher’s hands. In about a day, with under an hour of human attention, a researcher used GLM-5.3 to find several unknown flaws in a popular browser’s JavaScript engine and chain them into a web page that reads files from a visitor’s machine. In a second session, GLM-5.3-Flash turned a published Chrome flaw tracked as CVE-2026-11645 and one other known bug into a reliable exploit chain for 64-bit Arm processors that gets past pointer authentication. That took 20 minutes of human time, eight hours of model time, and $20.40 at Z.ai’s API prices.
Three: the safeguards come off. Anthropic removed the model’s refusal behaviour with a published technique called abliteration. Refusal rates on three public benchmarks fell from above 90% to about 3%, 2% and 12%. General science scores did not change and cyber scores fell by a few percent. The work took about 2,200 graphics processing unit (GPU) hours, roughly $4,400, for a team that had not done it before.
Four: weaker tricks also work. In a simulated environment the model was given openly malicious orders against critical systems.
| Condition | GLM-5.3 engaged | Claude through the API |
|---|---|---|
| Direct malicious order | 0% | 0% |
| False cover story | 64% | 0% |
| Prefilled reasoning | 92% | Not possible |
| Abliterated weights | 100% | Not possible |
Each cell is 50 trials. From this Anthropic concludes that GLM-5.3 is a step change in what attackers can freely use, that defenders need wider access to frontier models, and that governments should test capable models, including GLM-5.3’s successors.
The sequence matters, so here it is in one place. All dates are 2026.
| Date | Event |
|---|---|
| 7 April | Anthropic announces Claude Mythos Preview, released only to vetted partners |
| 13 May | The ExploitBench paper is posted |
| 22 May | Anthropic’s Project Glasswing update reports more than 10,000 vulnerabilities found |
| 14 August | Z.ai releases GLM-5.3 on its API and holds the weights |
| 25 to 28 August | The weights appear on Hugging Face |
| 26 August | The first uncensored copies we found appear |
| 17 September | CAISI publishes its assessment |
| 29 September | Anthropic publishes its report |
| About 1 October | Hugging Face disables one repository built for offensive use |
What holds up
Three parties with different interests agree on the central point.
CAISI ran its own four benchmarks and called GLM-5.3 “the most cyber-capable open-weight model released to date”. The previous best from labs in the People’s Republic of China (PRC), which was Kimi K3 on most of the benchmarks, was well behind on every one.
| CAISI benchmark | GLM-5.3 | Best US model | Previous PRC best |
|---|---|---|---|
| SEC-Bench Pro (183 tasks) | 40.4% | 90.2% | 27.3% |
| ExploitBench (41 tasks) | 61.1% | 100.0% | 32.2% |
| ExploitGym userspace (502 tasks) | 9.4% | 44.4% | 2.6% |
| CAISI OSS-Fuzz (297 tasks) | 7.7% | 23.2% | 2.4% |
Z.ai says the same thing in its own model card. It states that cyber capability grew faster than the team expected as post-training scaled, and that the gains were largest in exploitation, where GLM-5.3 more than doubled GLM-5.2. Z.ai held the weights for about two weeks for safety evaluation on those grounds. A developer warning about its own product is strong evidence.
The removal of safeguards is the least contestable claim of all. The technique dates from a 2024 paper by Arditi and co-authors, which found that refusal in 13 open chat models runs through a single direction in the model’s activations, and that erasing it stops refusals with little effect on anything else. We checked Hugging Face on 3 October 2026. The Flash weights repository was created on 25 August. A repository labelled uncensored and abliterated was created on 26 August. An abliterated build of the full model tagged for cybersecurity followed on 30 August and showed about 81,000 downloads on Hugging Face’s public counter. Another uncensored Flash build showed about 222,000. A search for abliterated GLM-5.3 returns at least twenty repositories.
So the core of the report stands. An open-weight model crossed the line into building real exploits, and nothing that ships inside the weights keeps it from being used that way.
Where the framing is stronger than the evidence
These points are our reading of the sources.
The comparison is with a five-month-old model. Anthropic says GLM-5.3 is similar to Mythos Preview, which was announced on 7 April 2026. CAISI’s second key finding is that GLM-5.3 is significantly below current US frontier models and about four months behind on its composite index. Both statements are true and they describe the same gap. Z.ai’s own table agrees: on ExploitBench it reports 54.4 for GLM-5.3 against 78.0 for Claude Fable 5 and 76.5 for GPT-5.6 Sol. A reader who sees only the Mythos comparison could miss that the frontier has moved.
The samples are small. Fifty against 56 out of 410 is not a measurable difference. By our calculation the two-proportion z statistic is about 0.6, and 410 attempts on 41 bugs are probably ten tries per bug and so not independent. Four against six out of 100 has 95% intervals of roughly 2% to 10% and 3% to 13%. The honest summary is “same class”, which is what Anthropic says, and no more precise than that. CAISI publishes confidence intervals. The text of Anthropic’s post gives none.
One benchmark gives three different numbers. ExploitBench reads 12% in Anthropic’s report, 61.1% in CAISI’s and 54.4 in Z.ai’s. The scoring differs: Anthropic counts only attempts that reach a complete exploit, CAISI averages a 16-point ladder over the best of three attempts, and Z.ai does not say. None is wrong. It does mean that a bare percentage from this benchmark is meaningless without the scoring rule.
The safeguard test is a simulation. Anthropic’s footnote says no model code ran and the shell was fake, with a second language model inventing the command output. “Engaged” means the model tried to connect to the target. Fifty trials per cell is five orders, two targets and five repeats. The direction of the result is believable because abliteration is well understood. The exact figures of 64% and 92% deserve less weight than the headlines gave them.
The zero for Claude is partly structural. Two of the three attacks cannot be attempted against a closed API, so two of the zeros are true by construction. The third, a cover story, was tested by the model’s own maker, with 50 trials, which rules out a success rate above about 7% for that prompt and says nothing about others. Anthropic’s threat report for the same month describes real operations that used Claude Haiku, Sonnet and Opus, including one group that targeted about 50 organisations and found around a dozen previously unknown vulnerabilities. Those operations were detected and shut down, which is the real advantage of a hosted model: the vendor can see misuse and revoke access. That is a stronger and more accurate argument than a refusal rate of zero.
The capability did not need a bypass. A second footnote says the exploit-building tasks in the first half of the report did not trigger GLM-5.3’s refusals at all. Refusals appeared only for malware and for attacks on remote targets. So the two halves of the report measure different things. Vulnerability research looks the same whether the person doing it is a defender or an attacker, and no refusal policy can tell them apart from the prompt alone.
Running it is not free. The full model is about 753 billion parameters. The independent benchmarker Artificial Analysis lists eight NVIDIA B200 GPUs for its reference deployment. It is accurate that anyone can download the weights, but prefilling and abliteration need either that hardware or a host willing to serve a modified copy. The cover story needs neither. The cheaper Flash model, which built the $20.40 exploit, is the more practical risk and gets less attention in the report.
What Z.ai did is left out. The report says the model shipped without meaningful safeguards. It does not mention the two-week hold on the weights, Z.ai’s public statement about the capability, or its claim of 2,436 vulnerability findings across 269 open source projects with most still under embargo. A reader can think those steps were inadequate, and the bypass rates suggest they were, but they are part of the record.
Closed labs have containment failures too. On 30 July 2026 Anthropic disclosed three incidents in which its own cyber evaluations reached real systems because an evaluation environment had live internet access. One run extracted credentials from a real company and another published a malicious package that was downloaded by 15 machines. Anthropic attributed these to harness and operations errors. The lesson for this debate is that control depends on operations, and operations fail under every release model.
Incentives
Every party here has an interest, and none of that makes their numbers wrong.
- Anthropic sells closed models and presents closed weights as a security property in this report. It runs paid and gated programmes for defenders. Secondary reporting says it filed confidentially for a public listing in June 2026. Its report asks governments to test capable models.
- Z.ai has been on the US Commerce Department’s Entity List since January 2025 and listed in Hong Kong in January 2026. It markets GLM-5.3’s cyber strength as a feature for defenders.
- CAISI is a US government body whose published comparison sets US models against PRC models.
- Open-weight users depend on these models staying available. On Hacker News the thread on the report drew 251 points and 234 comments, and the common objections were conflict of interest and fear that testing rules would shut out open competitors. The Decoder raised the same regulatory capture concern while accepting the findings.
The practical test is whether the claims are confirmed by someone with the opposite interest. Here the capability claim is confirmed by Z.ai, and the safeguard claim by public repositories anyone can count.
First principles
A safeguard is a control only for whoever holds the lever. On a hosted model the vendor holds it: classifiers, logs, account bans. With open weights the user holds it, and a user who wants it gone can remove it. Refusal training in open weights sets a default for honest users and a small cost for dishonest ones. It should be judged as that, and not as a barrier.
So the decisions that matter come before release. After weights are public nothing can be recalled. Hugging Face disabled one repository that advertised itself for offensive use, and a renamed copy from the same account stayed up. What a developer can still choose is whether to release, when, and what defenders get first.
The lag is short. By CAISI’s measure, open weights reached this level about four months after the closed frontier. Any plan that relies on a capability staying scarce has roughly that long.
Withholding only helps if it slows attackers more than defenders. A July 2026 paper by Daniel Commey models this directly. Well-resourced attackers can find substitutes through theft, distillation or their own work, while small defenders often cannot. In that case restriction leaves the weakest defenders furthest behind. The answer depends on numbers nobody has measured well, which is an argument for humility on both sides.
Patching is the bottleneck, and finding bugs is no longer scarce. Anthropic’s Glasswing update in May reported 530 high or critical bugs disclosed to maintainers and 75 patched. Z.ai reports 53 of 2,436 findings public. Both sides can now find flaws faster than maintainers can fix them. More access for defenders does not help much unless fixes ship faster.
Testing cannot substitute for release decisions. Anthropic asks for government testing. CAISI did test GLM-5.3 and published about three weeks after the weights appeared. The test was accurate and it changed nothing about who could download the model. Testing informs. It does not control.
What this means if you build with models
- Assume exploit-building capability is available to anyone, at about $20 for a known bug. Plan patch windows on that basis.
- Do not treat a model’s refusals as a security boundary, in either direction. They will not stop an attacker using open weights, and on closed models they will sometimes stop your own security work. Anthropic’s support pages describe a verification programme for that case.
- Sandbox agents as if the model will do what it is told. Anthropic’s July incidents came from a network setting, not from the model’s intent.
- When you compare cyber numbers, check which model version, whether safeguards were on, the scoring rule, and the sample size.
- If you host open weights for others, your terms and logging are now the only control between that capability and the open internet.
What we could not verify
- We could not read Z.ai’s own release article, which sits behind a paid wall on X. We relied on its model card, its licence text, and reporting by The Batch and VentureBeat.
- We found no response from Z.ai to the Anthropic report.
- We did not reproduce any benchmark or attempt any bypass.
- The date Hugging Face disabled the repository comes from secondary reports. We confirmed only that it is flagged disabled and that a renamed copy is live.
- Reports of Anthropic’s listing plans are secondary. We found no primary document.
- Some reports give 744 billion parameters. The Hugging Face metadata gives about 753 billion.
- Why Z.ai’s ExploitBench figure differs from CAISI’s is our inference from the two stated scoring rules.
- We could not open the South China Morning Post, Tom’s Hardware or OpenAI pages cited by others.
Disclosure
Velofy uses Anthropic’s models in its daily work, and its coding harness Kestrel lists GLM-5.3-Flash among the models it can run. We have a stake on both sides of this argument.
Sources
All read on 3 October 2026.
- GLM-5.3 and the spread of advanced cyber capabilities, Anthropic, 29 September 2026
- CAISI’s assessment of Z.ai’s GLM-5.3 cyber capabilities, NIST, 17 September 2026
- GLM-5.3 model card and licence, Z.ai on Hugging Face
- GLM-5.3-Flash model card, Z.ai on Hugging Face
- ExploitBench: a capability ladder benchmark for LLM cybersecurity agents, Lee and Brumley, 13 May 2026
- Refusal in language models is mediated by a single direction, Arditi and co-authors, 2024
- Who does withholding delay?, Commey, 24 July 2026
- Claude Mythos Preview: cybersecurity capabilities, Anthropic, 7 April 2026
- Project Glasswing: an initial update, Anthropic, 22 May 2026
- Threat intelligence report, September 2026, Anthropic
- Investigating three incidents in our cybersecurity evaluations, Anthropic, 30 July 2026
- Real-time cyber safeguards on Claude, Anthropic support
- Z.ai delayed weights for GLM-5.3 due to cybersecurity risk, The Batch
- GLM-5.3 is here with advanced cyber capabilities, VentureBeat, 14 August 2026
- Anthropic says Zhipu’s open-weight GLM-5.3 nearly matches Claude Mythos Preview, The Decoder, 30 September 2026
- Hacker News discussion of the report
- GLM-5.3 cybersecurity benchmarks: what the 84.5 score hides, D-Central
- Why Hugging Face disabled the offensive cyber GLM-5.3 model, AI IDE List
- GLM-5.3 providers, Artificial Analysis
- Z.ai, Wikipedia
- Anthropic IPO timeline, Forge Global, secondary reporting