Organizations/evaluator

Trajectory Labs

Third-party AI security evaluator that red-teams model safeguards for labs.

2 records1 attack1 capability
Sep 22, 2026
Claude Opus 5.5 system card reports cyber results at or above Mythos 5.1 and a three-stage cyber safeguard
CapabilitySystem cardAnthropic, Irregular, US Center for AI Standards and Innovation

Anthropic's Claude Opus 5.5 system card reports that, evaluated through the API with its cyber safeguards off, the model meets or exceeds Claude Mythos 5.1 and Claude Opus 5 on every cyber evaluation it includes (ExploitBench, CyScenarioBench, a rewritten Binary Exploitation Benchmark and ExploitGym). Anthropic places the model in the lower category of its Frontier Compliance Framework and says it sees no indication of novel offensive capability. It deployed the model with a three-stage cyber safeguard (an activation probe and two classifiers) that blocks vulnerability discovery in compiled binaries, and reports no evidence of a critical-severity jailbreak from internal and contracted external red teaming (the card gives no CAISI results).

Aug 26, 2026
Researcher reports a prompt-injection chain that gets code execution in Claude Code's Auto Mode; Anthropic closes it as working as designed
AttackVulnerability disclosureJohann Rehberger (Embrace The Red), Anthropic, Trajectory Labs

Security researcher Johann Rehberger (Embrace The Red) reports that a request to summarize a web page led Claude Code with Opus 5 in Auto Mode to run attacker-controlled code in his lab setup, using a multi-step chain of individually benign-looking actions, with success in 3 of 5 to 4 of 5 trials per variant. He says Anthropic closed his report as Informative and working as designed, and relays that Anthropic's position is that Auto Mode is a best-effort classifier for convenience, not a security boundary. He contrasts this with a third-party evaluation, commissioned by Anthropic and described in a post he cites, that showed 0.00% prompt-injection success for Opus 5 in Auto Mode on a fixed scenario set.