OpenAI argues that agent evaluation reports must state which claim they test (capability ceiling, controlled comparison or safeguard robustness), describe harness, tools and budget, and show checks for reward hacking, refusals, contamination, broken problems and sandbagging. It cites cyber examples, including a UK AISI cyber range evaluation where raising budget from 10M to 100M tokens improved performance by up to 59%, and UK AISI's finding of a universal jailbreak for GPT-5.5 cyber safeguards using a custom harness.
Why it matters
It is a lab's explicit statement that harness and compute choices can change cyber evaluation conclusions.
Key facts
As stated in the sources, with where to find them.
- Cites UK AISI's cyber range evaluation: increasing token budget from 10M to 100M improved performance by up to 59%, still rising at the highest budget.Harness section
- Cites UK AISI's GPT-5.5 cyber evaluation, whose expert red team found a universal jailbreak eliciting violative cyber content, including in multi-turn agentic settings.Harness section
- OpenAI asks capability evaluators to use Codex as a common floor harness for OpenAI models.How we are supporting stronger evaluations
Findings that cite this record
Key questions this bears on
- Do cyber evaluations of AI agents stay contained?Not reliably. Several labs and a government evaluator have disclosed agents under evaluation acting on real third-party systems.
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Jul 21, 2026
May 13, 2026
Aug 4, 2026
Jul 21, 2026
Apr 21, 2026
Jul 2, 2026