DeepSeek

AI lab that releases open-weight models.

2 records2 capabilityWebsite
May 1, 2026
CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months
CapabilityEvaluation reportUS Center for AI Standards and Innovation, DeepSeek, NIST

NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.

Sep 30, 2025
CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack
CapabilityEvaluation reportUS Center for AI Standards and Innovation, NIST, DeepSeek

NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.