Methodology
How the corpus is built, what each label means, and where it can mislead you. Read this before citing anything here.
Scope
The explorer covers AI models and agents in cybersecurity in three roles: as tools of attackers, as defenders, and as targets. It also covers how their cyber capability is measured and governed. It excludes deepfake-only and disinformation-only items, generic machine learning work with no security angle, and product announcements without substance.
Record types
Records (events) say what happened: a disclosure, a paper, an incident, a policy. Each has a date from its source, a lane, a kind, a summary in our own words, key facts with locators, and one or more sources. Findings say what we can conclude: one scoped claim, the records that support it, its relations to other findings, and a status history. Key questions are lasting questions about the field, each with a dated answer, a confidence level, and the findings it rests on. Openings are candidate research questions an editor drafts from patterns in the findings. Benchmarks and tools and organizations are reference entries that records point to.
Lanes
- Attacks & incidents. Real-world misuse, AI-enabled intrusions, AI malware, attacks on deployed agents, and agents acting outside their authority.
- Capability & gating. Measurements of offensive cyber capability, thresholds crossed, release and access decisions, and the spread of capability to open models.
- Defense & research. Defensive agents, vulnerability discovery and repair, agent-security defenses, and research on how to evaluate all of it.
- Policy & standards. Regulation, government guidance, standards, taxonomies, and lab safety frameworks.
Kinds
Incident · Misuse report · Vulnerability disclosure · Malware · Paper · Benchmark · Evaluation report · System card · Threshold · Access program · Tool release · Competition · Program · Framework · Standard · Guidance · Regulation · Dataset
Sourcing rules
- Primary sources first. A record cites the publisher of the work where one exists. Secondary sources are marked as such, and records without a primary source are flagged in the corpus checks.
- Dates come from the source. When only the month is certain, the record is dated to the month and drawn hollow on the chronicle.
- Our words, with attribution. Summaries paraphrase and attribute claims to their source. Quotations are short and rare.
- Numbers are scoped. Key facts give the figure as the source states it, with the model, benchmark, or date it applies to, and a locator.
- No offensive detail. Attacks are described as a security news article would: what was affected, the class of technique, the impact, and the fix. No payloads or step-by-step instructions.
- Vendor claims stay vendor claims. A result reported by the organization that built the system is recorded as reported, not measured, unless an independent party confirms it.
Significance
Each record carries an editorial significance from 1 to 5. A 5 changes how the field works or thinks: a first-of-its-kind incident, a threshold crossed, a landmark benchmark. A 3 is a solid, representative contribution. A 1 is minor but worth keeping for completeness. Significance sets mark size on the chronicle. It is a judgment, not a measurement.
Finding statuses
- Reported Stated by one source and not yet corroborated or challenged.
- Corroborated Supported by at least two independent sources.
- Qualified Still standing, but later work narrows how far it applies.
- Contested Later work directly disputes it.
- Superseded Replaced by a newer measurement of the same thing. Kept for the trend.
- Revalidate Older than its half-life with no newer evidence. May no longer hold.
Statuses are stored with their history, so a finding that was once corroborated and is now contested shows both, with the reason and date of each change. Revalidate is computed at build time: when a finding’s newest evidence is older than its half-life, it is flagged. Half-lives are set per finding. Capability measurements age in months; structural observations about how systems fail age more slowly.
Evidence kinds
- measured. An experiment or evaluation produced the number.
- observed. Seen in a real incident or operation.
- reported. Stated by a party without shareable evidence.
- argued. A position, framework, or analysis.
Corroborated requires at least two independent publishers. Two reports from the same organization count once.
Capability levels
The frontier assigns each task the strongest level of AI autonomy that credible public evidence supports. Public evidence lags private capability and labs disclose selectively, so a level is a floor, not an estimate.
- 0, No public evidence. No credible public source shows AI doing this task in a meaningful way.
- 1, Assists a human. AI speeds up a person who does the task; the person remains the operator.
- 2, Completes benchmark tasks. Agents complete the task autonomously on benchmarks or CTF-style challenges.
- 3, End to end in realistic settings. Agents complete the task autonomously against realistic targets, ranges, or real software, under test.
- 4, Observed on real systems. Agents have done the task autonomously against real systems outside a test, per a credible source.
Methods
Methods are attack techniques, defenses, and evaluation methods. Records and findings are linked to methods by matching their text against patterns for each method, restricted to records that share a topic or attack-surface component with it, then checked by editors. Each defense lists the attacks it is designed to counter; the coverage matrix on the methods page counts the findings that bear on each pair and marks pairs with no measured finding.
Model profiles on the models page collect records that mention a model family by name. They are collections, not capability assessments.
Verification
On Sep 25, 2026, when the initial corpus was assembled, every record and finding was checked against its sources in a verification pass separate from drafting. It corrected wrong numbers and dates, figures quoted from a later version of a paper than the one dated, overstated claims, and status histories that cited the wrong evidence; the audit found about one in six sampled records had a factual error before correction. Corrections to findings are appended to their status history with the reason, never rewritten. Records remain marked assistant-drafted until a Fide editor reviews them.
Research openings
The signal radar is computed every build from the corpus. Editorial openings are drafted from those signals and carry the signal types below. Each needs a prior-work check before it becomes a study.
- Incidents outpace defenses. Documented failures in an area with no evaluated defense.
- Rests on one source. A consequential finding that no one has replicated.
- Contested. Findings that later work disputes or narrows.
- Aging out. Evidence older than its half-life.
- Defense unmeasured. Defenses that are proposed but have no published evaluation.
- Transfers from Fide work. Methods from Fide’s other research domains that apply here.
- Fide agenda. Directly serves an open Fide research call.
AI assistance and review
Records are drafted with Anthropic’s Claude models from the cited sources, then checked by software for schema, dates, references, and duplicate sources, and enter the corpus only after an editor approves them. Because Anthropic is also covered here, the About page explains how the same rules apply to every organization. Every record shows its review state. assistant-drafted means the record has passed those checks but a Fide editor has not yet read it against the sources. human-reviewed means an editor has. Treat assistant-drafted records as leads to verify, not as settled facts.
What this cannot tell you
- Disclosure bias. The record reflects what organizations chose to publish. Labs and vendors report selectively. Many incidents are never disclosed, and many measurements are private.
- Coverage bias. Sources are mostly English-language and weighted toward US and UK organizations.
- Absence is not evidence. A gap in the atlas or a missing defense can mean the work does not exist, or only that it is not yet recorded here.
- Counts are not rates. The number of records about a topic reflects attention, not frequency in the world.
- Judgments are judgments. Significance, capability levels, finding relations, and openings are editorial. Their reasoning is shown so you can disagree with it.
Known coverage gaps
These are the areas where the corpus is known to be incomplete today. Treat counts and gaps in these areas with extra caution.
- Offensive capability measurements. Benchmark results and system-card cyber evaluations from before 2026 are thinly recorded. The benchmarks themselves are catalogued, but most per-model results are not yet records, so several offense rows on the frontier start late and show as not yet assessed. Reconnaissance and intrusion start in late 2025 from observed attacks rather than measurements.
- Threat-intelligence and misuse reports. The major provider and vendor reports on attackers using AI from 2024 to September 2026 are records (Anthropic, Google, Microsoft with OpenAI, ESET, Sysdig, ThreatDown), but OpenAI’s own threat reports and CERT-UA’s LAMEHUG report are not yet. Most of these reports describe misuse the provider detected on its own platform, so they show that such attacks happen, not how often.
- June to September 2026. Coverage of June to September 2026 comes from the themed segments and targeted additions, not a complete sweep. Capability and threshold items from those months may be missing.
- Non-English and non-Western sources. Sources are mostly English-language and weighted toward US and UK organizations.
How the record stays current
The explorer is built to be read long after any one update, so it separates what is stored from what is computed.
- Stored text carries dates, not relative time. Records and findings never say “recently” or “now”; the check that validates the corpus flags such words. Phrases like “3 weeks ago” and “in the 30 days to” are computed when you view the page.
- Key questions stay fixed; answers are revised in the open. A revision adds a dated answer and keeps every earlier one. An answer is flagged for review when relevant records are added, when one of its findings changes status or ages past its half-life, or 120 days after its last review.
- The quarterly record shows the answers as they stood at the end of each quarter, derived from the answer histories so it cannot drift.
- Corrections are not movement. When new evidence changes a finding, it is shown as movement in the field. When we correct our own earlier judgment, the fix is appended to the finding’s history as a correction and listed in the changelog, not on the Desk.
- Update cadence. The record is swept for new work at least every two weeks, in a supervised session where an editor accepts each addition. The dot beside “Updated” on the home page turns gold when a sweep is overdue.
- Freshness is visible. Each finding shows how far it is through its half-life, each topic shows its latest record, and the site says when it was last updated. If it has not been updated for a month, every page says so.
Updates and corrections
New items enter through a pipeline that crawls a registry of sources, drafts typed records, and queues them for review. A single paper or incident can be added in minutes. Corrections are welcome by email to alex@fideai.org, and each record page links to a prefilled correction message. Every change to a record is kept in Fide’s version history, and every addition, status change, revised answer, and correction is listed in the changelog.
Independence
The explorer is maintained by Fide AI under its published independence and conflict-of-interest policies. No organization in the corpus pays for inclusion, placement, or review. Fide’s own research questions appear where they relate to the evidence, and are labeled as Fide’s.