Where things stood, Q4 2026

This quarter is still open. The answers below are the current ones; they will be fixed on Dec 31, 2026.

  1. Can prompt injection against AI agents be reliably defended?

    Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.

    high confidenceAnswer dated Sep 30, 2026
  2. Where are deployed AI agents actually being exploited?

    Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.

    moderate confidenceAnswer dated Oct 1, 2026
  3. Do cyber evaluations of AI agents stay contained?

    Not reliably. Labs and a government evaluator disclosed agents reaching real systems from cyber evaluations; OpenAI agents did so from training runs too.

    high confidenceAnswer dated Oct 1, 2026
  4. How far can measured AI cyber capability be trusted?

    As a lower or conditional bound. Scores move with budget, pipeline and contamination, and the same benchmark name can hide different setups.

    moderate confidenceAnswer dated Sep 30, 2026
  5. Is AI shifting the balance between finding and fixing vulnerabilities?

    Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.

    moderate confidenceAnswer dated Sep 30, 2026
  6. Can AI agents defend and oversee systems on their own?

    Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.

    moderate confidenceAnswer dated Sep 30, 2026
  7. How are attackers using AI agents in real operations?

    Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.

    moderate confidenceAnswer dated Oct 1, 2026