Start here

How to read this

The explorer connects what happened (records) to what we can conclude (findings) to what it means (key questions), and keeps each step linked to the evidence. This page explains where to start and what every label means.

In one minute

  1. 1
    Key questions

    Start with the lasting questions about the field. Each has a current answer, a confidence level, and the findings behind it.

  2. 2
    Findings

    Each finding is one checkable claim with its evidence, its limits, and a status that changes as new work arrives.

  3. 3
    Records

    Records say what happened: a paper, incident, report, or policy, summarized in our words and linked to its sources.

  4. 4
    Openings

    Where the evidence runs out, openings describe the research that would settle the question.

Everything else is a way into the same evidence: the methods an attack or defense uses, the part of an agent system it hits (atlas), how far AI has reached on a task (frontier), and what changed each week (desk). Search anything with ⌘K (Ctrl K on Windows and Linux).

Lanes

Every record belongs to one of four lanes, color-coded throughout.

Attacks & incidents
Real-world misuse, AI-enabled intrusions, AI malware, attacks on deployed agents, and agents acting outside their authority.
Capability & gating
Measurements of offensive cyber capability, thresholds crossed, release and access decisions, and the spread of capability to open models.
Defense & research
Defensive agents, vulnerability discovery and repair, agent-security defenses, and research on how to evaluate all of it.
Policy & standards
Regulation, government guidance, standards, taxonomies, and lab safety frameworks.

Significance

An editorial judgment of how much a record matters to the field, from one to five.

Changes how the field works or thinks.
A solid, representative contribution.
Minor, kept for completeness.

Finding statuses

A finding’s status changes as later work supports, narrows, disputes, or replaces it. Every change is kept.

Reported
Stated by one source and not yet corroborated or challenged.
Corroborated
Supported by at least two independent sources.
Qualified
Still standing, but later work narrows how far it applies.
Contested
Later work directly disputes it.
Superseded
Replaced by a newer measurement of the same thing. Kept for the trend.
Revalidate
Older than its half-life with no newer evidence. May no longer hold.

Evidence kinds

How the evidence behind a finding was produced.

measured
An experiment or evaluation produced the number.
observed
Seen in a real incident or operation.
reported
Stated by a party without shareable evidence.
argued
A position, framework, or analysis.

Confidence

How strongly the findings support a key question’s current answer.

low confidence
The evidence is thin, one-sided, or limited by a known gap in the corpus.
moderate confidence
The evidence is consistent but limited in scope, independence, or measurement.
high confidence
Independent, consistent findings support the answer.

Review flags on key questions

A key question is flagged for an editor when the evidence behind its answer moves. The flag clears when the answer is revised or confirmed.

new since review
Records bearing on the answer were added, or one of its findings changed status, after the last review.
relies on a weakened finding
The answer cites a finding that is contested, superseded, or past its recheck date.
review due
More than 120 days since an editor last reviewed the answer.

Freshness

Each finding has a half-life: how long before its evidence should be rechecked. The bar fills as its newest evidence ages; past the tick, the finding shows as revalidate.

Evidence 60 days old, recheck in 305 days
Recent evidence.
Evidence 300 days old, recheck in 65 days
Approaching its recheck.
Evidence 450 days old, past its 365-day recheck
Past its recheck: due for new evidence.

Capability levels

Used on the frontier. A level records the strongest credible public evidence, so it is a floor, not an estimate.

0 · No public evidence
No credible public source shows AI doing this task in a meaningful way.
1 · Assists a human
AI speeds up a person who does the task; the person remains the operator.
2 · Completes benchmark tasks
Agents complete the task autonomously on benchmarks or CTF-style challenges.
3 · End to end in realistic settings
Agents complete the task autonomously against realistic targets, ranges, or real software, under test.
4 · Observed on real systems
Agents have done the task autonomously against real systems outside a test, per a credible source.
Not yet assessed
No evidence recorded for that task in that quarter. Not the same as zero.

Review and corrections

assistant-drafted
Drafted from the cited sources with AI assistance and checked by software, but not yet read against the sources by a Fide editor. Treat it as a lead to verify.
human-reviewed
A Fide editor has read it against its sources.
correction
We fixed our own earlier judgment. Corrections are listed in the changelog and never shown as movement in the field.
New
Added or revised since your last visit (remembered only in your browser).

Atlas components

The parts of an agent system where attacks land and defenses sit.

Supply chain
Model artifacts, packages, IDE extensions, and skills that an agent system is built from.
Access gate
Account controls, trusted access programs, and verification that decide who can use cyber capability.
Untrusted content
Anything the agent reads that an outsider could write. The entry point for indirect prompt injection.
Agent model
The model making decisions: its safeguards, refusals, and susceptibility to jailbreaks.
Memory & context
State the agent carries between steps and sessions, including summaries it writes for itself.
Tools & MCP
The actions an agent can take and the servers that provide them.
Credentials
The authority an agent holds, and what it can reach with it.
Execution sandbox
Where agent-written code runs and what that environment can reach.
Other agents
Sub-agents, peer agents, and the messages passed between them.
Monitor
Automated systems that watch agent behavior and can stop it.
Human approver
People asked to approve, review, or intervene, and the load placed on them.
Evaluation environment
The environments used to test agents, which have themselves become part of the attack surface.

Opening signals

Patterns in the evidence that suggest a research question is open.

Incidents outpace defenses
Documented failures in an area with no evaluated defense.
Rests on one source
A consequential finding that no one has replicated.
Contested
Findings that later work disputes or narrows.
Aging out
Evidence older than its half-life.
Defense unmeasured
Defenses that are proposed but have no published evaluation.
Transfers from Fide work
Methods from Fide’s other research domains that apply here.
Fide agenda
Directly serves an open Fide research call.
Keeping up

Updated at least every two weeks. Follow new records, findings that move, and revised answers through the RSS feed, a feed for any topic, or the changelog. For how records are sourced and checked, read the methodology.