Chronicle/Defense & research

OpenAI publishes a playbook on harness choice and validity checks for third-party evaluations

DefenseGuidanceSignificance assistant-drafted

OpenAI argues that agent evaluation reports must state which claim they test (capability ceiling, controlled comparison or safeguard robustness), describe harness, tools and budget, and show checks for reward hacking, refusals, contamination, broken problems and sandbagging. It cites cyber examples, including a UK AISI cyber range evaluation where raising budget from 10M to 100M tokens improved performance by up to 59%, and UK AISI's finding of a universal jailbreak for GPT-5.5 cyber safeguards using a custom harness.

Why it matters

It is a lab's explicit statement that harness and compute choices can change cyber evaluation conclusions.

Key facts

As stated in the sources, with where to find them.

  • Cites UK AISI's cyber range evaluation: increasing token budget from 10M to 100M improved performance by up to 59%, still rising at the highest budget.Harness section
  • Cites UK AISI's GPT-5.5 cyber evaluation, whose expert red team found a universal jailbreak eliciting violative cyber content, including in multi-turn agentic settings.Harness section
  • OpenAI asks capability evaluators to use Codex as a common floor harness for OpenAI models.How we are supporting stronger evaluations

Findings that cite this record

Key questions this bears on

Sources

Related records