Capability frontier
For each offensive and defensive task, the strongest level of AI autonomy that public evidence supports, quarter by quarter. Each level is a judgment against a written rubric, and every cell links to the records behind it.
Benchmark results and system-card cyber evaluations from before 2026 are thinly recorded. The benchmarks themselves are catalogued, but most per-model results are not yet records, so several offense rows on the frontier start late and show as not yet assessed. Reconnaissance and intrusion start in late 2025 from observed attacks rather than measurements. All known gaps.
One tab stop: use the arrow keys to move between quarters and rows, Home and End to jump to the first and last quarter.
Agents under evaluation identified and acted against real third-party systems outside their test environments.
Rubric
A level records the strongest credible public evidence, not the typical case. Public evidence lags private capability, and labs disclose selectively, so a level is a floor on what exists, not an estimate of it. Offense rows follow MITRE ATT&CK tactics and defense rows follow NIST CSF 2.0 functions.
| Level | Label | What must be true |
|---|---|---|
| 0 | No public evidence | No credible public source shows AI doing this task in a meaningful way. |
| 1 | Assists a human | AI speeds up a person who does the task; the person remains the operator. |
| 2 | Completes benchmark tasks | Agents complete the task autonomously on benchmarks or CTF-style challenges. |
| 3 | End to end in realistic settings | Agents complete the task autonomously against realistic targets, ranges, or real software, under test. |
| 4 | Observed on real systems | Agents have done the task autonomously against real systems outside a test, per a credible source. |
Thresholds, gating, and access
Release decisions, capability thresholds, and access programs tied to cyber capability.
Measurements and benchmarks
Evaluations of offensive capability. See benchmarks and tools for the instruments themselves.
Open-weight diffusion
The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.
NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.
NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.