As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
UK AISI reports that fixed, low token budgets understate frontier cyber capability and its rate of progress. A preprint audit finds pipeline choices alone can move cyber benchmark scores by tens of points, public CTF benchmarks can be contaminated, and counting a crash as exploitation overstates capability. Single capability numbers, including trend estimates, are best read as lower or conditional bounds.
The findings behind it
4 reportedEach finding carries a status that changes as new work arrives. What the statuses mean.
- Cyber capability measured at fixed, low token budgets understates what frontier models can do and how fast they are improving.Reported measured · 4 evidence records
- Evaluation pipeline choices alone can move a model's cybersecurity benchmark score by more than 80 points and reorder models.Reported measured · 1 evidence record
- Scores on public CTF benchmarks can be inflated when agents find published solutions, and static benchmarks lose validity as their flaws are patched.Reported measured · 2 evidence records
- Counting a crash as exploitation success overstates capability; most public models in May 2026 stalled before code execution on a browser engine.Reported measured · 1 evidence record
Answer history
Answers are never edited after the fact. A revision adds a new answer and keeps the earlier ones here.
- 2026-09-25moderate confidencecurrentAs a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.First answer, drawn from the findings linked here.