Anthropic's Claude Opus 5.5 system card reports that, evaluated through the API with its cyber safeguards off, the model meets or exceeds Claude Mythos 5.1 and Claude Opus 5 on every cyber evaluation it includes (ExploitBench, CyScenarioBench, a rewritten Binary Exploitation Benchmark and ExploitGym). Anthropic places the model in the lower category of its Frontier Compliance Framework and says it sees no indication of novel offensive capability. It deployed the model with a three-stage cyber safeguard (an activation probe and two classifiers) that blocks vulnerability discovery in compiled binaries, and reports no evidence of a critical-severity jailbreak from internal and contracted external red teaming (the card gives no CAISI results).
It pairs a lab's own exploit-capability numbers with the safeguard design and red-team results it relied on for general deployment, all reported by the vendor and run on different setups from earlier cards.
Key facts
As stated in the sources, with where to find them.
- Anthropic states Claude Opus 5.5 falls within the lower category of its Frontier Compliance Framework's two cyber tiers (Tier 1: meaningful assistance for operations using known techniques, still dependent on humans; Tier 2: autonomous operations with novel offensive capability) and that it sees no indication of novel offensive capability. All capability results were run with cyber classifiers disabled and a new escape-attempt classifier enabled; many harnesses were rewritten and older models re-evaluated, so numbers differ from earlier system cards.Section 3.1
- ExploitBench (41 V8 vulnerabilities from 2024 onward, mitigations on, five trials each, 300-turn budget, the authors' static harness): Claude Opus 5.5 captured a mean of 13.99 capability flags (Cap% 91%) in the plain arm and 14.15 (Cap% 91%) in the AutoNudge arm, which re-prompts a model that stops early.Section 3.3.1, Figure 3.3.1.A
- ExploitBench, same setup: Opus 5.5 reached full arbitrary code execution in 301 of 410 runs (73.4%) across the plain and AutoNudge arms. Anthropic cautions the results may not be comparable to public leaderboard entries produced under vendors' deployed conditions.Section 3.3.1
- CyScenarioBench (Irregular; private 10-challenge subset of multi-stage operations in containerized networks; cyber mitigations off): overall solve rate 67.6% for Opus 5.5, 61.7% for Claude Mythos 5.1, 53.0% for Claude Opus 5, and under 1% for Claude Sonnet 5.Section 3.3.2, Figure 3.3.2.A
- Binary Exploitation Benchmark, formerly called OSS-Fuzz (831 fuzz-harness entrypoints across 228 projects, patched snapshots, no hints; harness rewritten and older models rerun; the text gives no turn or time budget): control-flow hijacks numbered 106 for Opus 5.5, 81 for Mythos 5.1, 66 for Opus 5 and 3 for Sonnet 5. Anthropic says the numbers are not comparable to previous system cards.Section 3.3.3, Figure 3.3.3.A
- ExploitGym (869 instances: 502 userspace, 181 V8, 186 Linux kernel; all mitigations disabled; success verified by an agent judge as using the intended vulnerability): Opus 5.5 exploited 289 instances within 2 hours and 300 within 6 hours; Mythos 5.1 exploited 203 and 264. Anthropic retired CyberGym from the card as saturated.Section 3.3.4, Figure 3.3.4.A
- Anthropic reports the generally available model uses the same cyber policy as Claude Opus 5 but safeguards it describes as comparably robust to those on Claude Fable 5 and 5.1: an activation probe screens all traffic, a lightweight classifier running on Opus 5.5 scores escalated traffic, and a trained LLM classifier decides with the probe's verdict. Source-code vulnerability discovery is allowed at all access levels; vulnerability discovery in compiled binaries is blocked. Blocked requests fall back to Claude Opus 4.8 (automatic in Anthropic's apps, opt-in on the API), and Anthropic adopted a temporarily wider safety margin against jailbreaks while working to reduce false positives.Sections 3.2 and 3.4
- The card says classifier coverage recall is comparable to prior models and that blockrates on defensive vulnerability discovery fell against prior Opus and Fable models; the numeric values appear only in figures and are not restated here.Sections 3.4.1 and 3.4.2, Figures 3.4.1.A and 3.4.2.A
- Internal rewind-attacker evaluation (a helpful-only attacker model with a 400-call limit and rewinds, against a Claude Code agent with production safeguards on tasks such as ransomware, scaled exfiltration, CVE exploitation, command-and-control and self-replication): Anthropic reports Opus 5.5 comparably robust to the Fable models and more robust than Opus 5, shown in a figure only. It reports no evidence of a critical-severity jailbreak for Opus 5.5, Fable 5.1, Fable 5 or Opus 5, testing an earlier Opus 5.5 snapshot.Sections 3.5 and 3.5.1, Figure 3.5.1.A
- Trajectory Labs spent about 95 hours and sent over 29,000 requests against exploit-reproduction tasks, reporting 13 candidate breaks across seven tasks and no universal jailbreak; on one task, after roughly five hours, a working privilege-escalation chain was produced with the work split across over 100 separate contexts. 10a Labs spent about 56 hours on 82 multi-turn conversations and none advanced past proof of concept.Section 3.5.3
- Gray Swan's Shade attacker: across 61 critical-infrastructure scenarios (about 3,300 attempts of up to 30 turns) the safeguards refused over 90% of attempts outright and none reached the objective; across six exploit-reproduction and ransomware-staging scenarios (about 1,700 attempts) about a quarter were refused outright and no working exploit was produced. Gray Swan's state judge recorded no breaks. CAISI also collaborated on cyber and biological safeguard measurement; the card gives no results for that work.Sections 3.5.2 and 3.5.3
- Malicious use of Claude Code, without additional safeguards (61 malicious and 61 dual-use or benign cyber prompts, each run 10 times): Opus 5.5 refused malicious requests 79.8% of the time (Opus 5 83.6%, Mythos 5.1 90.3%, Sonnet 5 90.7%) and assisted with dual-use and benign requests 99.8% of the time. Anthropic says refusals are a secondary protection behind the blocking classifier.Section 5.1.1, Table 5.1.1.A
- Anthropic's launch post says Opus 5.5 is comparable to Claude Mythos 5.1 in cybersecurity and is deployed with safeguards similar to those on Fable 5.1, and that access for verified practitioners through its Cyber Verification Program would expand in the weeks after the announcement.Launch post, Safety
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move with budget, pipeline and contamination, and the same benchmark name can hide different setups.
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.