UK AISI's Science of Evaluation team measured how agent success changes with token budget across software, academic and cyber tasks. About 8% of cyber tasks were solved only at budgets of 10M tokens or more, and the frontier cyber time-horizon trend was about 60% steeper at a 50M budget than at 2.5M; AISI recommends reporting capability curves rather than single scores.
Why it matters
Single-budget cyber evaluation scores can miss capability that appears at higher, attacker-affordable compute.
Key facts
As stated in the sources, with where to find them.
- About 8% of cyber tasks were solved only at 10M+ token budgets, some requiring up to 50M tokens.Our findings
- Cyber time horizons doubled every 4.7 months at a 2.5M budget; the trend is about 60% steeper at 50M; one frontier model's horizon rose from about 40 minutes (2.5M) to about 4 hours (50M).Our findings
- Human task time predicted agent compute need via a power law (exponent about 0.7-1.0) across 211 software and 78 cyber tasks.Our findings
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Jul 21, 2026
May 29, 2026
May 13, 2026
Jul 30, 2026
Jul 16, 2026
May 13, 2026