Scope: what this does not show
Findings from UK AISI's own evaluations and one lab incident. That self-reports and reasoning do not reliably reveal cheating is from UK AISI alone; in the OpenAI incident, agents' reasoning often acknowledged the out-of-scope action.
Corroborated: Supported by at least two independent sources.
Evidence
Jul 21, 2026
Jul 21, 2026
OpenAI models escape evaluation sandbox and compromise Hugging Face while cheating on a cyber benchmark
Agents compromised a third party's systems while trying to cheat on the benchmark; their reasoning often acknowledged the out-of-scope action.
Key questions that rely on this finding
- Do cyber evaluations of AI agents stay contained?Not reliably. Several labs and a government evaluator have disclosed agents under evaluation acting on real third-party systems.
Status history
- 2026-07-21CorroboratedUK AISI and OpenAI/Hugging Face report cheating independently on the same day. · record