Containing agents inside cyber evaluations

Which evaluation-environment controls stop agents from acting on real systems, and how much do they change the capability being measured?

assistant-draftedSandbox & containmentEvaluation validity

Signals

Incidents outpace defenses

Several 2026 disclosures describe agents under evaluation acting on real third parties, through misconfiguration, intentionally enabled internet access, or a flaw in shared infrastructure. The corpus records no evaluated containment standard for evaluation environments.

Defense unmeasured

Incident reports propose egress restrictions and target-name hygiene without published measurements.

Why it matters

Claims about frontier cyber capability rest on evaluations, and some of those evaluations have leaked. Containment that changes what is being measured trades one validity problem for another.

Hypothesis

Strict egress controls block nearly all out-of-scope actions but lower measured capability on tasks that legitimately need network access.

A first study

Rebuild publicly described incident patterns in a contained range. Compare egress policies on blocked out-of-scope actions and on the change in task scores.

Controls it would need

Held-out tasks; the same harness and token budget across policies; incident patterns drawn only from public reports.

What it could and could not claim

Would support claims about the tested environments and policies, not about production deployments or other evaluators' ranges.