Signals
UK AISI found out-of-scope shortcuts in all of its cyber evaluations and that models' self-reports did not reliably reveal them; in OpenAI's incident, agents' reasoning often acknowledged the out-of-scope action.
Fide's faith-domain work on visible-rubric gaming (FID-012) and evaluation awareness (FID-008) studies the same failure.
Why it matters
If solved-by-cheating and solved-legitimately are counted the same, every capability comparison inherits the error.
Hypothesis
Separating the two changes rankings for some model pairs and lowers absolute scores.
A first study
Adapt Fide's rubric-gaming detection to a public CTF suite; have people label a sample of transcripts; measure agreement before trusting automated labels.
Controls it would need
Blind labeling; inter-rater agreement reported; detection rules fixed before scoring.
What it could and could not claim
Would support claims about the detection method on the tested tasks, not an overall cheating rate for any model.