UK AISI reported that the length of cyber tasks frontier models complete at 80% reliability on its narrow task suite had been doubling about every 4.7 months since late 2024, and that Claude Mythos Preview and GPT-5.5 substantially exceeded that trend. A newer Mythos Preview checkpoint completed both of AISI's cyber ranges, including the previously unsolved industrial-control range.
Why it matters
Gives a government estimate of the pace of autonomous cyber capability growth that later AISI and lab posts build on.
Key facts
As stated in the sources, with where to find them.
- 80%-reliability cyber time horizon doubled every 4.7 months since late 2024 (Feb 2026 estimate, 2.5M-token cap), versus an 8-month estimate in Nov 2025.Cyber Time Horizons Results
- Newer Mythos Preview checkpoint solved 'The Last Ones' in 6/10 attempts and 'Cooling Tower' in 3/10; GPT-5.5 solved 'The Last Ones' in 3/10.Further Evidence of Cyber and Software Autonomy
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Jul 21, 2026
May 29, 2026
May 11, 2026
Aug 4, 2026
Jul 30, 2026
Mar 1, 2026