Chronicle/Capability & gating

CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months

CapabilityEvaluation reportSignificance assistant-drafted

NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.

Why it matters

Tracks how quickly open-weight models approach frontier cyber capability, which governs how long closed-model safeguards buy defenders.

Key facts

As stated in the sources, with where to find them.

  • On CTF-Archive-Diamond, DeepSeek V4 Pro scored 32% (imputed via item response theory), versus 71% for GPT-5.5 and 46% for Opus 4.6.Cyber results
  • CAISI estimates V4 Pro lags the frontier by about 8 months, performing similarly to GPT-5; IRT-estimated Elo 800 plus or minus 28 vs 1260 plus or minus 28 for GPT-5.5.Key findings

Findings that cite this record

Key questions this bears on

Sources

Related records