How to read this
The explorer connects what happened (records) to what we can conclude (findings) to what it means (key questions), and keeps each step linked to the evidence. This page explains where to start and what every label means.
In one minute
- 1Key questions
Start with the lasting questions about the field. Each has a current answer, a confidence level, and the findings behind it.
- 2Findings
Each finding is one checkable claim with its evidence, its limits, and a status that changes as new work arrives.
- 3Records
Records say what happened: a paper, incident, report, or policy, summarized in our words and linked to its sources.
- 4Openings
Where the evidence runs out, openings describe the research that would settle the question.
Everything else is a way into the same evidence: the methods an attack or defense uses, the part of an agent system it hits (atlas), how far AI has reached on a task (frontier), and what changed each week (desk). Search anything with ⌘K (Ctrl K on Windows and Linux).
Lanes
Every record belongs to one of four lanes, color-coded throughout.
- Attacks & incidents
- Real-world misuse, AI-enabled intrusions, AI malware, attacks on deployed agents, and agents acting outside their authority.
- Capability & gating
- Measurements of offensive cyber capability, thresholds crossed, release and access decisions, and the spread of capability to open models.
- Defense & research
- Defensive agents, vulnerability discovery and repair, agent-security defenses, and research on how to evaluate all of it.
- Policy & standards
- Regulation, government guidance, standards, taxonomies, and lab safety frameworks.
Significance
An editorial judgment of how much a record matters to the field, from one to five.
- Changes how the field works or thinks.
- A solid, representative contribution.
- Minor, kept for completeness.
Finding statuses
A finding’s status changes as later work supports, narrows, disputes, or replaces it. Every change is kept.
- Reported
- Stated by one source and not yet corroborated or challenged.
- Corroborated
- Supported by at least two independent sources.
- Qualified
- Still standing, but later work narrows how far it applies.
- Contested
- Later work directly disputes it.
- Superseded
- Replaced by a newer measurement of the same thing. Kept for the trend.
- Revalidate
- Older than its half-life with no newer evidence. May no longer hold.
Evidence kinds
How the evidence behind a finding was produced.
- measured
- An experiment or evaluation produced the number.
- observed
- Seen in a real incident or operation.
- reported
- Stated by a party without shareable evidence.
- argued
- A position, framework, or analysis.
Confidence
How strongly the findings support a key question’s current answer.
- low confidence
- The evidence is thin, one-sided, or limited by a known gap in the corpus.
- moderate confidence
- The evidence is consistent but limited in scope, independence, or measurement.
- high confidence
- Independent, consistent findings support the answer.
Review flags on key questions
A key question is flagged for an editor when the evidence behind its answer moves. The flag clears when the answer is revised or confirmed.
- new since review
- Records bearing on the answer were added, or one of its findings changed status, after the last review.
- relies on a weakened finding
- The answer cites a finding that is contested, superseded, or past its recheck date.
- review due
- More than 120 days since an editor last reviewed the answer.
Freshness
Each finding has a half-life: how long before its evidence should be rechecked. The bar fills as its newest evidence ages; past the tick, the finding shows as revalidate.
- Evidence 60 days old, recheck in 305 days
- Recent evidence.
- Evidence 300 days old, recheck in 65 days
- Approaching its recheck.
- Evidence 450 days old, past its 365-day recheck
- Past its recheck: due for new evidence.
Capability levels
Used on the frontier. A level records the strongest credible public evidence, so it is a floor, not an estimate.
- 0 · No public evidence
- No credible public source shows AI doing this task in a meaningful way.
- 1 · Assists a human
- AI speeds up a person who does the task; the person remains the operator.
- 2 · Completes benchmark tasks
- Agents complete the task autonomously on benchmarks or CTF-style challenges.
- 3 · End to end in realistic settings
- Agents complete the task autonomously against realistic targets, ranges, or real software, under test.
- 4 · Observed on real systems
- Agents have done the task autonomously against real systems outside a test, per a credible source.
- Not yet assessed
- No evidence recorded for that task in that quarter. Not the same as zero.
Review and corrections
- assistant-drafted
- Drafted from the cited sources with AI assistance and checked by software, but not yet read against the sources by a Fide editor. Treat it as a lead to verify.
- human-reviewed
- A Fide editor has read it against its sources.
- correction
- We fixed our own earlier judgment. Corrections are listed in the changelog and never shown as movement in the field.
- New
- Added or revised since your last visit (remembered only in your browser).
Atlas components
The parts of an agent system where attacks land and defenses sit.
- Supply chain
- Model artifacts, packages, IDE extensions, and skills that an agent system is built from.
- Access gate
- Account controls, trusted access programs, and verification that decide who can use cyber capability.
- Untrusted content
- Anything the agent reads that an outsider could write. The entry point for indirect prompt injection.
- Agent model
- The model making decisions: its safeguards, refusals, and susceptibility to jailbreaks.
- Memory & context
- State the agent carries between steps and sessions, including summaries it writes for itself.
- Tools & MCP
- The actions an agent can take and the servers that provide them.
- Credentials
- The authority an agent holds, and what it can reach with it.
- Execution sandbox
- Where agent-written code runs and what that environment can reach.
- Other agents
- Sub-agents, peer agents, and the messages passed between them.
- Monitor
- Automated systems that watch agent behavior and can stop it.
- Human approver
- People asked to approve, review, or intervene, and the load placed on them.
- Evaluation environment
- The environments used to test agents, which have themselves become part of the attack surface.
Opening signals
Patterns in the evidence that suggest a research question is open.
- Incidents outpace defenses
- Documented failures in an area with no evaluated defense.
- Rests on one source
- A consequential finding that no one has replicated.
- Contested
- Findings that later work disputes or narrows.
- Aging out
- Evidence older than its half-life.
- Defense unmeasured
- Defenses that are proposed but have no published evaluation.
- Transfers from Fide work
- Methods from Fide’s other research domains that apply here.
- Fide agenda
- Directly serves an open Fide research call.
Updated at least every two weeks. Follow new records, findings that move, and revised answers through the RSS feed, a feed for any topic, or the changelog. For how records are sourced and checked, read the methodology.