Anthropic's Frontier Red Team reports that Zhipu AI's open-weight GLM-5.3 developed end-to-end exploits in 50 of 410 ExploitBench attempts, against 56 of 410 for Claude Mythos Preview, and full control-flow hijacks in 4% of trials on a 100-task subset of Anthropic's internal binary-exploitation benchmark, against 6% for Mythos Preview. In a simulated harmful-request test, Anthropic reports the model's refusals were bypassed in 64% to 100% of trials using a cover story, prefilled reasoning, or an abliterated copy of the weights, techniques it says did not work on safeguarded Claude models. All figures are Anthropic's own and were produced on setups the post describes only in part.
It is a vendor report that ties an open-weight model's exploit results to how cheaply its refusals can be removed, but its ExploitBench counts come from a setup the post does not fully specify and cannot be compared directly with other Anthropic cards.
Key facts
As stated in the sources, with where to find them.
- Scope of the report: Anthropic tested GLM-5.3 (Zhipu AI, also called Z.ai) in isolated sandboxes against offline targets it set up. In the post's Figure 2, the two Claude models shown (Opus 4.6 and Mythos Preview) were run with safeguards disabled; the post does not state the safeguard setting for the figures given in its text. Anthropic characterizes GLM-5.3 as released without meaningful safeguards; that is its assessment, not an independent finding.Introduction; GLM-5.3 can develop working exploits end to end; Figure 2 caption
- ExploitBench (Carnegie Mellon benchmark of known V8 vulnerabilities): Anthropic reports end-to-end exploits in 50 of 410 attempts for GLM-5.3 and 56 of 410 for Claude Mythos Preview. The post does not say how the 410 attempts divide across bugs, trials or arms, nor the turn budget or harness used.GLM-5.3 can develop working exploits end to end
- Comparability caution: the Opus 5.5 system card (2026-09-22) also reports a count out of 410 on ExploitBench, 301 of 410 full arbitrary-code-execution runs for Opus 5.5, but there 410 is 41 V8 bugs times five trials times two arms (plain and AutoNudge), a 300-turn budget, and the benchmark authors' static harness; the card reports no Mythos Preview figure. The GLM-5.3 post gives none of these parameters, so 50 of 410 and 56 of 410 cannot be read as rates on the card's setup, and 301 of 410 should not be compared with them. The ExploitBench paper's own headline for Mythos Preview is a different unit: arbitrary code execution on 18 of 41 bugs, best of three seeds, primary arm, 300 turns.GLM-5.3 post, ExploitBench paragraph; Opus 5.5 system card Section 3.3.1; ExploitBench paper (arXiv:2605.14153), abstract and evaluation setup
- Binary Exploitation benchmark (Anthropic internal, previously published as OSS-Fuzz; here a random 100-task subset, full credit for a full control-flow hijack): GLM-5.3 succeeded in 4% of trials and Mythos Preview in 6%; Claude Opus 4.6 and GLM-5.2 did not succeed on any. The post does not give the trials per task. The Opus 5.5 system card reports this benchmark on 831 entrypoints as counts of hijacks, a different subset and unit.GLM-5.3 can develop working exploits end to end, Footnote 1; Opus 5.5 system card Section 3.3.3
- Figure 2 plots the share of attempts reaching each benchmark's top outcome against output-token budget for Opus 4.6, Mythos Preview (both with safeguards disabled), GLM-5.2, GLM-5.3, and two other open-weight models, Moonshot AI's Kimi K3 and DeepSeek's V4.1-Flash. The text gives no numbers from the figure.Figure 2 caption
- Human-in-the-loop session with GLM-5.3 on a sandboxed local Linux build of a popular browser: over about a day, with limited human attention, the model found several previously unknown vulnerabilities in the JavaScript engine and chained them into a working exploit for that build. Anthropic says it disclosed these to the maintainer and is reviewing further reports in other widely used systems. Sessions typically ran a day or less with under an hour of human focus. The vulnerabilities and their validity are reported, not independently checked.GLM-5.3 can develop working exploits end to end; Figure 3
- Second session, with the smaller GLM-5.3-Flash and a disclosed browser flaw (CVE-2026-11645) plus another known flaw: with no significant direction, the model built a reliable exploit chain for an ARM64 target that bypassed pointer-authentication hardening, using 20 minutes of human attention and eight hours of model work, which Anthropic prices at $20.40 at Zhipu's API prices.GLM-5.3 can develop working exploits end to end
- Abliteration (weight editing that removes refusals): Anthropic produced its own abliterated GLM-5.3, taking about 2,200 GPU hours at roughly $4,400 (GLM-5.3-Flash about 600 GPU hours); it estimates an experienced team would need closer to 600 GPU hours ($1,200). Refusal rates on JailbreakBench, HarmBench and StrongREJECT went from above 90% to about 3%, 2% and 12%. GPQA-Diamond scores were unchanged, and a tested subset of CyberGym scored a few percent lower. Anthropic says several developers released abliterated versions of GLM-5.3 within days of its release.GLM-5.3 lacks robust safeguards; Footnote 3; Figure 4
- Simulated harmful-request test (50 samples per cell: five attack orders, two targets, five attempts; engagement means the model tried to connect to a remote target after an overtly harmful request): GLM-5.3 engaged 0% on direct requests, 64% with a deceptive red-team cover story, 92% with prefilled reasoning, and 100% with the abliterated model. Anthropic reports that every tested Claude model stayed at 0% under API safeguards, and that prefilling and abliteration do not apply to the Claude API.GLM-5.3 lacks robust safeguards; Figure 5
- The harmful-request simulation runs no model-generated code and has no external access; a fake shell tool returns results approximated by another LLM from a description of the simulated world. Anthropic says such simulations are imperfect measures of real behavior.Footnote 4
- Anthropic cites a 2026-09-17 assessment by NIST's Center for AI Standards and Innovation as finding GLM-5.3 the most cyber-capable open-weight model released to that date and about four months behind the US frontier on an aggregate of CAISI's benchmarks, and says its own capability findings broadly match. CAISI's US comparison models were tested with safeguards disabled and include models released only to vetted users. The CAISI assessment is reported here second-hand.Introduction
- Anthropic's conclusions, argued rather than measured: it expects state and non-state actors to use models like GLM-5.3 for real-world harm, says defenders should have access to frontier models at least as capable as attackers', and calls for government safety testing of sufficiently capable models including GLM-5.3's successors.What does this mean?
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move with budget, pipeline and contamination, and the same benchmark name can hide different setups.
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.