Findings/fixed-budgets-understate-cyber-capability

Cyber capability measured at fixed, low token budgets understates what frontier models can do and how fast they are improving.

Reportedmeasured4 evidence records from 2 independent sourcesassistant-drafted
Scope: what this does not show

UK AISI measurements on its task suite; the size of the effect varies by model and task.

Reported: Stated by one source and not yet corroborated or challenged.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2026-05-29ReportedOpenAI's evaluation playbook warns that unreported budgets understate capability. · record
  2. 2026-07-02CorroboratedUK AISI measures the effect directly. · record
  3. 2026-09-25ReportedcorrectionThe OpenAI playbook cites UK AISI's own measurements, so all evidence comes from one evaluator. · record