Key questions/can-agents-defend-autonomously

Can AI agents defend and oversee systems on their own?

moderate confidenceAnswered Sep 25, 2026Reviewed assistant-drafted
Current answer

Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.

LLM agents perform poorly on realistic threat-hunting and investigation benchmarks, while vendor-run trials report assistants speeding up analysts who remain in charge. Agent monitors can be driven to miss covert actions under optimized attack, so catch rates measured against ordinary mistakes say little about adversarial settings.

The findings behind it

2 corroborated, 1 reported, 1 qualified

Each finding carries a status that changes as new work arrives. What the statuses mean.

Answer history

Answers are never edited after the fact. A revision adds a new answer and keeps the earlier ones here.

  1. 2026-09-25moderate confidencecurrentNot yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.First answer, drawn from the findings linked here.