Findings/every-frontier-model-hijackable

Every frontier agent tested in two large public red-teaming competitions was hijacked at least once; in the 2026 competition the injections also had to conceal the compromise from the user.

Reportedmeasured2 evidence records from 1 independent sourceassistant-drafted
Scope: what this does not show

Competition settings with many attackers; per-attempt success is low and varies about 17-fold across models.

Reported: Stated by one source and not yet corroborated or challenged.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2025-07-28ReportedGray Swan competition: every agent attacked successfully for most behaviors. · record
  2. 2026-03-16Reportedreconfirmed2026 competition with US CAISI and UK AISI repeats the result for 13 models. · record