Scope: what this does not show
One evaluation of one model on non-public benchmarks.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
May 1, 2026
CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months
CAISI found weaker results on non-public agentic and reasoning benchmarks (ARC-AGI-2 semi-private, PortBench, CTF-Archive-Diamond) than DeepSeek self-reported.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding supports Evaluation pipeline choices alone can move a model's cybersecurity benchmark score by more than 80 points and reorder models.Both show headline scores depend on evaluation choices; here the difference is which benchmarks are run, not how the same benchmark is run.
Status history
- 2026-05-01ReportedCAISI evaluation of DeepSeek V4 Pro. · record