Scope: what this does not show
One preprint audit of public cyber benchmarks.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Independent government testing can find a model weaker on agentic tasks than its developer's self-reported benchmarks suggest. supports this findingBoth show headline scores depend on evaluation choices; here the difference is which benchmarks are run, not how the same benchmark is run.
Key questions that rely on this finding
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Status history
- 2026-09-08ReportedAudit of pipeline dependence in cyber benchmarks. · record