Methods/Evaluation method

Compute-scaled evaluation

Measuring capability across token and compute budgets instead of at one fixed budget.

5 records5 defense3 findings (3 measured)First recorded 2026-02assistant-drafted

How it works

The same tasks are run at several budgets; capability that appears only at higher budgets would be missed by a single low cap.

What we know

2 reported, 1 qualified

Records over time

RangeLanes
5 of 5 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
2026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

5 records · newest first
Jul 2026
Jul 16, 2026
Cost-aware evaluation finds defensive SOC agents do not scale with compute like offensive CTF agents
DefensePaper

Researchers evaluate security agents at fixed cost levels on offensive Cybench challenges and defensive Splunk BOTS v1 investigations, splitting spend into inference and tool use. They find offensive success rises with test-time compute, while defensive investigation depends more on disciplined tool use and telemetry navigation, and argue benchmarks should report cost and operational fit alongside success.

Jul 2, 2026
UK AISI finds agent evaluations understate cyber capability without accounting for test-time compute
DefenseEvaluation reportUK AI Security Institute

UK AISI's Science of Evaluation team measured how agent success changes with token budget across software, academic and cyber tasks. About 8% of cyber tasks were solved only at budgets of 10M tokens or more, and the frontier cyber time-horizon trend was about 60% steeper at a 50M budget than at 2.5M; AISI recommends reporting capability curves rather than single scores.

May 2026
May 29, 2026
OpenAI publishes a playbook on harness choice and validity checks for third-party evaluations
DefenseGuidanceOpenAI, UK AI Security Institute, METR

OpenAI argues that agent evaluation reports must state which claim they test (capability ceiling, controlled comparison or safeguard robustness), describe harness, tools and budget, and show checks for reward hacking, refusals, contamination, broken problems and sandbagging. It cites cyber examples, including a UK AISI cyber range evaluation where raising budget from 10M to 100M tokens improved performance by up to 59%, and UK AISI's finding of a universal jailbreak for GPT-5.5 cyber safeguards using a custom harness.

Mar 2026
Mar 1, 2026
UK AISI's SandboxEscapeBench measures frontier models' ability to escape container sandboxes
DefenseBenchmarkUK AI Security Institute, University of Oxford

Researchers at the UK AI Security Institute and the University of Oxford built SandboxEscapeBench, 18 nested-sandbox capture-the-flag scenarios covering orchestration, runtime and kernel weaknesses rated difficulty 1 to 5. In the March 2026 version, which tested nine models, GPT-5 and Claude Opus 4.5 succeeded on about half of all runs, difficulty 1-2 scenarios were solved at high rates, no model succeeded at difficulty 4 or 5, and success rose roughly log-linearly with token budget. An August 2026 revision reports that newer models did markedly better, with Claude Mythos Preview saturating the public benchmark.

Feb 2026
Feb 7, 2026
AIxCC SoK finds stability decided results and many validated AI patches were still semantically wrong
DefensePaperGeorgia Institute of Technology, Texas A&M University, DARPA

A systematization-of-knowledge paper by organizers and competitors analyzes AIxCC's design, the seven finalist architectures and results beyond the scoreboard. It reports that system stability and accuracy penalties decided rankings, that LLM-based systems found vulnerabilities a fuzzing baseline missed, and that among patches passing all automatic validation, manual review found semantic errors in 38-46% from baseline agents; the top two systems had 83.8% and 79.2% competition-scored patch accuracy.

All records