Chronicle/Defense & research

UK AISI finds agent evaluations understate cyber capability without accounting for test-time compute

DefenseEvaluation reportSignificance assistant-drafted

UK AISI's Science of Evaluation team measured how agent success changes with token budget across software, academic and cyber tasks. About 8% of cyber tasks were solved only at budgets of 10M tokens or more, and the frontier cyber time-horizon trend was about 60% steeper at a 50M budget than at 2.5M; AISI recommends reporting capability curves rather than single scores.

Why it matters

Single-budget cyber evaluation scores can miss capability that appears at higher, attacker-affordable compute.

Key facts

As stated in the sources, with where to find them.

  • About 8% of cyber tasks were solved only at 10M+ token budgets, some requiring up to 50M tokens.Our findings
  • Cyber time horizons doubled every 4.7 months at a 2.5M budget; the trend is about 60% steeper at 50M; one frontier model's horizon rose from about 40 minutes (2.5M) to about 4 hours (50M).Our findings
  • Human task time predicted agent compute need via a power law (exponent about 0.7-1.0) across 211 software and 78 cyber tasks.Our findings

Findings that cite this record

Key questions this bears on

Sources

Related records