Scope: what this does not show
Purpose-built vulnerable containers, not hardened production sandboxes. About half applies to GPT-5 and Claude Opus 4.5; smaller models scored much lower. No model solved the hardest scenarios, and all escapes used known misconfigurations or public CVEs.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
Status history
- 2026-03-01ReportedUK AISI publishes SandboxEscapeBench. · record