Scope: what this does not show
UK AISI measurements on its task suite; the size of the effect varies by model and task.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
May 29, 2026
Jul 2, 2026
UK AISI finds agent evaluations understate cyber capability without accounting for test-time compute
Raising the budget from 2.5M to 50M tokens moved one model's cyber time horizon from about 40 minutes to about 4 hours.
Mar 1, 2026
May 13, 2026
UK AISI says frontier cyber task horizons doubled every 4.7 months, with Mythos Preview and GPT-5.5 above trend
UK AISI's May post already said the 2.5M-token cap understates frontier capability.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies Under a 2.5M-token cap, frontier cyber task time horizons doubled on the order of months between late 2024 and early 2026.The doubling estimate was measured at a fixed low budget.
Key questions that rely on this finding
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Status history
- 2026-05-29ReportedOpenAI's evaluation playbook warns that unreported budgets understate capability. · record
- 2026-07-02CorroboratedUK AISI measures the effect directly. · record
- 2026-09-25ReportedcorrectionThe OpenAI playbook cites UK AISI's own measurements, so all evidence comes from one evaluator. · record