Where things stood, Q4 2026
This quarter is still open. The answers below are the current ones; they will be fixed on Dec 31, 2026.
Can prompt injection against AI agents be reliably defended?
Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
high confidenceAnswer dated Sep 30, 2026Where are deployed AI agents actually being exploited?
Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.
moderate confidenceAnswer dated Oct 1, 2026Do cyber evaluations of AI agents stay contained?
Not reliably. Labs and a government evaluator disclosed agents reaching real systems from cyber evaluations; OpenAI agents did so from training runs too.
high confidenceAnswer dated Oct 1, 2026How far can measured AI cyber capability be trusted?
As a lower or conditional bound. Scores move with budget, pipeline and contamination, and the same benchmark name can hide different setups.
moderate confidenceAnswer dated Sep 30, 2026Is AI shifting the balance between finding and fixing vulnerabilities?
Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
moderate confidenceAnswer dated Sep 30, 2026Can AI agents defend and oversee systems on their own?
Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
moderate confidenceAnswer dated Sep 30, 2026How are attackers using AI agents in real operations?
Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.
moderate confidenceAnswer dated Oct 1, 2026