NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.
Why it matters
It is a government evaluation that treats agent hijacking susceptibility as a national security property of foreign models.
Key facts
As stated in the sources, with where to find them.
- On software engineering and cyber tasks, the best US model evaluated solves over 20% more tasks than the best DeepSeek model.NIST news release, key findings
- Agents built on DeepSeek R1-0528 were on average 12 times more likely than evaluated US frontier models to follow malicious hijacking instructions in simulated environments.NIST news release, security findings
- With a common jailbreak, R1-0528 responded to 94% of overtly malicious requests versus 8% for US reference models.NIST news release, security findings
- US reference models: GPT-5, GPT-5-mini, gpt-oss (OpenAI) and Opus 4 (Anthropic).NIST news release
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
May 1, 2026
Mar 16, 2026
Nov 24, 2025
Mar 13, 2026
Mar 10, 2026
Mar 1, 2026