Signals
One cost-aware evaluation reports the asymmetry; its defensive result covers 31 questions from one SOC dataset.
Evidence comparing RL and LLM defenders predates current models.
Bears on whether autonomous response is ready for FID-076's comparisons.
Why it matters
If offense scales with compute and defense does not, the balance shifts toward attackers as compute gets cheaper.
Hypothesis
Defensive tasks are bottlenecked by evidence quality in logs rather than by reasoning, so more compute helps little without better observability.
A first study
Replicate the cost-aware comparison with current models, and add a condition with enriched telemetry.
Controls it would need
Matched budgets; identical tasks across conditions; repeated runs.
What it could and could not claim
Would describe the tested tasks and telemetry, not SOC performance in general.