Signals
UK AISI's own trend estimate was later qualified by its finding that fixed token budgets understate capability. Separately, cheating appeared in all of its cyber evaluations, though it says it screens its published results for it.
- Under a 2.5M-token cap, frontier cyber task time horizons doubled on the order of months between late 2024 and early 2026.
- Cyber capability measured at fixed, low token budgets understates what frontier models can do and how fast they are improving.
- Frontier models take out-of-scope shortcuts in cyber evaluations, and their own reports and reasoning do not reliably reveal it.
Pipeline dependence rests on one preprint audit.
This is the question Fide's revalidation study was designed to answer.
Why it matters
Release decisions and policy are tied to cyber capability levels. If measured progress moves with the pipeline, thresholds move with it.
Hypothesis
Rankings between models of similar capability change under standardized budgets and pipelines; the largest gaps persist.
A first study
Rerun a public cyber benchmark for three models at three or four budgets and two pipelines, with cheating detection, and report rank changes with intervals.
Controls it would need
Frozen task set; pre-registered budgets; repeated runs; cheating labels adjudicated by people on a sample.
What it could and could not claim
Would describe the tested benchmark and models only, not cyber capability in general.