Locate

Capability frontier

For each offensive and defensive task, the strongest level of AI autonomy that public evidence supports, quarter by quarter. Each level is a judgment against a written rubric, and every cell links to the records behind it.

assistant-draftedLevels assessed as of Sep 25, 2026. Read the rubric before citing a level.
Coverage note

Benchmark results and system-card cyber evaluations from before 2026 are thinly recorded. The benchmarks themselves are catalogued, but most per-model results are not yet records, so several offense rows on the frontier start late and show as not yet assessed. Reconnaissance and intrusion start in late 2025 from observed attacks rather than measurements. All known gaps.

as of Q3 2026

One tab stop: use the arrow keys to move between quarters and rows, Home and End to jump to the first and last quarter.

2023
2024
2025
2026
2026
Offense · MITRE ATT&CK tactics
Recon & targeting
Vulnerability discovery
Exploit development
Intrusion & movement
Evasion & persistence
Defense · NIST CSF 2.0 functions
Detect & triage
Patch & repair
Contain & respond
Recover
Oversee other agents
0 · No public evidence1 · Assists a human2 · Completes benchmark tasks3 · End to end in realistic settings4 · Observed on real systemsnot yet assessed● level changed this quarter
Recon & targeting · Q3 2026
Level 4: Observed on real systems

Agents under evaluation identified and acted against real third-party systems outside their test environments.

Evidence · level set from Q3 2026
Average assessed level, offense vs defense
0123420232026
— offense 3.4— defense 3.3

Rubric

A level records the strongest credible public evidence, not the typical case. Public evidence lags private capability, and labs disclose selectively, so a level is a floor on what exists, not an estimate of it. Offense rows follow MITRE ATT&CK tactics and defense rows follow NIST CSF 2.0 functions.

LevelLabelWhat must be true
0No public evidenceNo credible public source shows AI doing this task in a meaningful way.
1Assists a humanAI speeds up a person who does the task; the person remains the operator.
2Completes benchmark tasksAgents complete the task autonomously on benchmarks or CTF-style challenges.
3End to end in realistic settingsAgents complete the task autonomously against realistic targets, ranges, or real software, under test.
4Observed on real systemsAgents have done the task autonomously against real systems outside a test, per a credible source.

Thresholds, gating, and access

Release decisions, capability thresholds, and access programs tied to cyber capability.

Measurements and benchmarks

Evaluations of offensive capability. See benchmarks and tools for the instruments themselves.

Open-weight diffusion

Jul 23, 2026
UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models
CapabilityEvaluation reportUK AI Security Institute, US Center for AI Standards and Innovation, Moonshot AI

The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.

May 1, 2026
CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months
CapabilityEvaluation reportUS Center for AI Standards and Innovation, DeepSeek, NIST

NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.

Sep 30, 2025
CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack
CapabilityEvaluation reportUS Center for AI Standards and Innovation, NIST, DeepSeek

NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.