NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.
Why it matters
Tracks how quickly open-weight models approach frontier cyber capability, which governs how long closed-model safeguards buy defenders.
Key facts
As stated in the sources, with where to find them.
- On CTF-Archive-Diamond, DeepSeek V4 Pro scored 32% (imputed via item response theory), versus 71% for GPT-5.5 and 46% for Opus 4.6.Cyber results
- CAISI estimates V4 Pro lags the frontier by about 8 months, performing similarly to GPT-5; IRT-estimated Elo 800 plus or minus 28 vs 1260 plus or minus 28 for GPT-5.5.Key findings
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Sep 30, 2025
Mar 16, 2026
Apr 21, 2026
Mar 13, 2026
Jul 23, 2026
Jul 21, 2026