Anthropic's GLM-5.3 Cyber Report: The Numbers, the Backlash, and What Developers Should Take From It

TL;DR
Anthropic's Frontier Red Team says GLM-5.3 builds working exploits at rates close to Claude Mythos Preview, and that its safeguards can be stripped in days.
Anthropic's Frontier Red Team published a study on September 29 arguing that Z.ai's GLM-5.3 is the first freely downloadable model that can autonomously build end-to-end cyber exploits, and that its safeguards can be stripped with standard techniques. Its own numbers put GLM-5.3 within two points of Claude Mythos Preview, Anthropic's gated cyber model: 12 percent versus 14 percent on end-to-end exploit development, against near zero for earlier open models.
The useful part is not the warning. The gap between gated frontier capability and downloadable capability is now measured in months, and refusal training stops being a control surface once weights are public: the open-weights fight Anthropic's CEO entered in July with his position on open-weights models, reopened with data.
The numbers Anthropic published#
Five months ago Anthropic previewed Claude Mythos through Project Glasswing behind a trusted-access program. The new post argues GLM-5.3 crossed the same threshold without access controls:
- On ExploitBench, which measures end-to-end exploits against known V8 bugs, GLM-5.3 hit 50 of 410 attempts (12 percent) and Mythos Preview 56 of 410 (14 percent). Claude Opus 4.6, GLM-5.2, Kimi K3, and DeepSeek V4.1-Flash scored at or near zero.
- On Anthropic's internal Binary Exploitation benchmark (100 OSS-Fuzz tasks), GLM-5.3 produced full control-flow hijacks in 4 percent of trials versus Mythos's 6. Earlier models scored zero.
- With researchers in the loop, GLM-5.3 spent about a day (under an hour of human attention) on a sandboxed browser build, found unknown JavaScript-engine flaws, and chained them into a drive-by exploit that reads arbitrary files. Findings were disclosed to the maintainer.
- GLM-5.3-Flash turned public details of CVE-2026-11645, a recently disclosed Chrome flaw, into a reliable ARM64 exploit chain that bypassed pointer-authentication hardening: 20 minutes of human attention plus eight hours of model time, about $20.40 at Zhipu's API prices.
The safeguard findings are the policy payload. GLM-5.3 refused every overt malicious request by default, but engaged 64 percent of the time with a false cover story, 92 percent with prefilled reasoning, and 100 percent after abliteration, the standard edit that removes refusal behavior. That took about 2,200 GPU hours and $4,400 on a first attempt (600 hours, $1,200 for an experienced team) and left capability largely intact: GPQA-Diamond unchanged, CyberGym a few points lower. Refusals fell from above 90 percent to 3, 2, and 12 percent on JailbreakBench, HarmBench, and StrongREJECT.
The government numbers that came first#
NIST's Center for AI Standards and Innovation published its evaluation on September 17, calling GLM-5.3 "the most cyber-capable open-weight model released to date" while measuring it about four months behind the US frontier. US models were tested with cyber safeguards disabled where applicable.
| Benchmark | GLM-5.3 | US frontier best |
|---|---|---|
| SEC-Bench Pro (find and trigger) | 40.4% | 90.2% |
| ExploitBench (build the exploit) | 61.1% | 100% |
| ExploitGym, userspace | 9.4% | 44.4% |
| CAISI OSS-Fuzz | 7.7% | 23.2% |
The two assessments agree on capability and differ on framing: CAISI measures the gap, Anthropic measures how easily the safeguards come off.
Who wins, who loses, and the second-order effect#
This is a Stratechery story about incentives, not about malware.
The winner is Z.ai. The report is the strongest validation a competitor could provide: Anthropic's red team published numbers showing the downloadable model lands in the same class as its own gated one, and named it. The community priced that in immediately; the top r/LocalLLaMA thread is titled "Anthropic just dropped the greatest advertisement for GLM ever."
The loser is gating as protection. Anthropic says defenders should use the best available tools and that vetted defenders should get frontier cyber access. That concedes the report's own point: once capability is downloadable, restriction is a distribution problem, not a safety boundary. The abliteration math undercuts refusal training as a moat: days and a few thousand dollars, not a research lab.
The second-order effect is procurement. If closed frontier APIs refuse legitimate defensive work while open weights answer, security teams migrate to the tools that work, and the buyer segment the safety narrative targets moves away from the labs making the argument. Multiple practitioners in the Hacker News thread already describe doing this. A refusal layer that blocks paying defenders is a competitive disadvantage dressed as a feature, the same logic as our GLM 5.2 margin collapse thesis, applied to security work.
The sharpest counter-case from the thread: a flat 0 percent red-team rate against your own models is not evidence of safety, it is evidence about which jailbreaks you tested.
What people are actually saying#
The Hacker News thread (233 points, 224 comments) and the Reddit crossposts are largely hostile to Anthropic's framing, with a real minority taking the risk seriously.
- The conflict-of-interest read dominates. The top-voted comment calls the paper a research conflict of interest with a direct competitor; several others argue it sets a narrative for restricting open weights, and one r/LocalLLaMA crosspost is framed as marketing.
- The counter-case is present. One commenter says open models create risks that are hard to contain and asks to be corrected. Others attack the methodology rather than the motive: the 0 percent result against Claude models could reflect cherry-picked jailbreak attempts.
- Practitioner reports are the most useful signal. A cybersecurity worker says closed models refuse legitimate security work while GLM-5.3 and Flash stay reliable; another used an open model for malware forensics after a Claude refusal; a third had a buffer-overflow question answered elsewhere minutes after a refusal.
- Local serving is real, if not casual. One commenter runs GLM-5.3-Flash at Q4 on two 128GB M2 Ultra Mac Studios at about 50 tokens per second; another estimates the full 753B model needs a four-node DGX Spark cluster at NVFP4.
What this means for your team#
- Stop treating refusals as a security boundary. Any model you can download can be made to answer. Sandbox agent egress, scope credentials, and audit reach: the security models comparison and runtime security checklist cover the practical layer.
- If you do security work, budget for open-weight tooling. Anthropic itself says defenders should use the best tool for the job. Verified prices and setup are in Where to Run GLM-5.3 Free and Cheap; the full model is datacenter-class, Flash is the practical runner.
FAQ#
Is GLM-5.3 as capable as Claude at cyber tasks?#
Not quite. Anthropic's numbers put Mythos Preview ahead (14 versus 12 percent on ExploitBench, 6 versus 4 percent internally), and NIST CAISI measures GLM-5.3 about four months behind the US frontier. The change is that it is the first open-weight model close to that class.
Did Anthropic prove that open weights are dangerous?#
It proved refusal training is removable: abliteration cost under $5,000 in compute, took days, and left capability nearly unchanged. CAISI independently found similar capability. That is a property of releasing weights, not of where they were trained.
Should teams stop using GLM-5.3 for coding?#
No. The report's capabilities are vulnerability discovery and exploit development, also core defensive skills. Anthropic's own post argues defenders should use the best available tools. Review permissions before pointing it at production.
Sources#
- Anthropic Frontier Red Team: GLM-5.3 and the spread of advanced cyber capabilities - September 29, 2026
- NIST CAISI: Assessment of Z.ai's GLM-5.3 Cyber Capabilities - September 17, 2026
- Hacker News thread: GLM-5.3 and the spread of advanced cyber capabilities - 233 points, 224 comments, accessed September 30, 2026
- r/LocalLLaMA: Anthropic just dropped the greatest advertisement for GLM ever - accessed September 30, 2026
- r/LocalLLaMA: GLM-5.3 crosspost - accessed September 30, 2026
- r/Anthropic: GLM-5.3 and the spread of advanced cyber capabilities - accessed September 30, 2026
- Hugging Face: zai-org/GLM-5.3 - open weights and serving paths
Continue Reading#
- Where to Run GLM-5.3 Free and Cheap - verified provider prices, the Coding Plan, and self-host routes
- Anthropic CEO Dario Amodei on Open-Weights Models - the July position this report lands inside
- OpenAI Ships GPT-5.6-Cyber Through Daybreak Red - the gated alternative to open cyber capability
- GLM 5.2 and the AI Margin Collapse Thesis - why free and cheap capability keeps winning procurement
- Cybersecurity Skills for AI Agents Are Becoming Runtime Infrastructure - the defensive layer that does not depend on refusals
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on local and open-weight models
Where to Run GLM-5.3 Free and Cheap: Providers Compared
GLM-5.3 is Z.ai's open-weights coding model. Run it through Z.ai's API at $1.40 per million input tokens, cheaper OpenRouter hosts, a Coding Plan, or self-host.
7 min readAnthropic CEO Dario Amodei on open-weights models: the position, the pushback, and what it means for developers
Dario Amodei published Anthropic's stance on open-weights models this week - no total ban, but support for chip export controls, distillation crackdowns, and mandatory safety testing. HN responded with 800+ comments calling it regulatory capture. Here is what the CEO said, what the thread argued, and why the debate matters for every developer deploying AI.
8 min readOpenAI Ships GPT-5.6-Cyber Through Daybreak Red: The Numbers, the Chrome CVE, and What Access Looks Like
GPT-5.6-Cyber is OpenAI's gated model for authorized vulnerability research and exploit validation, with a 95% completion rate on sensitive security queries versus 1.5% for the base model. It already produced a fixed Chrome CVE. Here is what actually shipped and who gets it.
8 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








