Research openings
Where the evidence is thin, unreplicated, contested, aging, or outrun by events. The signals below are computed from the corpus every build. Openings are candidate research questions an editor drafted from those signals, each traceable to its evidence.
A gap here can mean no one has published on it, or only that we have not recorded the work yet. Each opening needs a prior-work check before it becomes a study. Point us to prior work.
Openings
How much of measured cyber progress is measurement?
How much do cyber capability trends change when token budget, pipeline, and cheating controls are held constant?
seedFID-075assistant-draftedWhat does a human approval actually check?
When an agent asks for approval, do people have the information and time to catch the consequential action?
seedFID-076assistant-draftedWhat evidence should gate an AI-generated patch?
Which independent checks change the decision to accept an AI-generated patch, and what do they cost?
seedFID-088assistant-draftedWhy don't defensive agents get better with more compute?
Why do defensive security agents gain less from extra compute than offensive agents, and what would close the gap?
seedFID-076assistant-draftedContaining a compromised agent in a team
Do independent evidence checks between cooperating agents contain a compromised member without stalling the team?
seedFID-087assistant-draftedContaining agents inside cyber evaluations
Which evaluation-environment controls stop agents from acting on real systems, and how much do they change the capability being measured?
seedFID-075 · FID-077assistant-draftedDo production prompt-injection defenses survive independent adaptive attack?
How do the 2026 production defenses that labs report as robust perform under independent adaptive attack?
seedFID-074assistant-draftedScoring cyber evaluations when models cheat
Can cheating in cyber evaluations be detected and scored separately so that capability comparisons stay valid?
seedFID-075 · FID-012 · FID-008assistant-draftedShould agentic browsers act on what users cannot see?
Do agentic browsers that separate visible from hidden page content resist injection better, and what does the separation cost in task success?
seedFID-074assistant-draftedWhat should an AI-written incident report show before anyone acts on it?
Which evidence requirements make AI-written incident reports correct or drop unsupported conclusions, without making them less useful?
seedFID-077 · FID-076assistant-draftedMemory as a persistence channel
How often do instructions, injected or self-generated, survive in an agent's memory and summaries to act in later sessions?
seedFID-074 · FID-077assistant-draftedSignal radar
Attack surface components with three or more documented attacks and no measured defense finding.
- Credentials19 documented attacks or incidents, no measured defense finding.
- Human approver8 documented attacks or incidents, no measured defense finding.
- Memory & context5 documented attacks or incidents, no measured defense finding.
Topics where attack records outnumber defense records at least two to one over the last 12 months.
- AI-enabled intrusion8 attack-lane records against 0 defense-lane records in the last 12 months.
- Threat intelligence8 attack-lane records against 0 defense-lane records in the last 12 months.
- Sandbox & containment13 attack-lane records against 2 defense-lane records in the last 12 months.
- Agent supply chain6 attack-lane records against 0 defense-lane records in the last 12 months.
- AI malware5 attack-lane records against 0 defense-lane records in the last 12 months.
- Tools & MCP5 attack-lane records against 1 defense-lane records in the last 12 months.
- Incident reporting10 attack-lane records against 3 defense-lane records in the last 12 months.
Measured or observed findings that depend on a single publisher and that other findings or landmark records lean on.
- In DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.All evidence comes from DARPA. 1 other finding depends on it.
- Separating an agent's control flow from untrusted data can give provable protection against control-flow hijacking at a modest utility cost.All evidence comes from Google DeepMind. 1 other finding depends on it.
- Under a 2.5M-token cap, frontier cyber task time horizons doubled on the order of months between late 2024 and early 2026.All evidence comes from UK AI Security Institute. 1 other finding depends on it.
- On ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.All evidence comes from UC Berkeley. 1 other finding depends on it.
- Evaluation pipeline choices alone can move a model's cybersecurity benchmark score by more than 80 points and reorder models.All evidence comes from the authors of "Audit finds cybersecurity LLM benchmark scores swing over 80 points with evaluation pipeline choices". 1 other finding depends on it.
- Agentic browsers act on page content that users cannot see.All evidence comes from Brave. Its source is rated high significance.
- Every frontier agent tested in two large public red-teaming competitions was hijacked at least once; in the 2026 competition the injections also had to conceal the compromise from the user.All evidence comes from Gray Swan AI. Its source is rated high significance.
- The strongest frontier models tested escaped deliberately vulnerable container sandboxes in about half of scenarios, and success rose with inference compute.All evidence comes from UK AI Security Institute. Its source is rated high significance.
- and 1 more
Findings that later work disputes or narrows.
- Frontier labs' layered defenses reduce prompt injection in tool-use, browser and computer-use agents but do not eliminate it.Qualified by 2 later findings.
- In DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.Qualified by 1 later finding.
- Under a 2.5M-token cap, frontier cyber task time horizons doubled on the order of months between late 2024 and early 2026.Qualified by 1 later finding.
- On ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.Qualified by 1 later finding.
- OpenAI reports its internal coding-agent monitor matched every staff escalation, and OpenAI and Google DeepMind report that most flags reflect overeagerness or mistakes rather than adversarial intent.Qualified by 1 later finding.
- Undefended tool-using agents follow injected instructions in a substantial share of benchmark cases.Qualified by 1 later finding.
Findings past their half-life with no newer evidence.
- In an early AI-control experiment, trusted editing gave a better safety-usefulness tradeoff than untrusted monitoring.Last evidence 1019 days ago; half-life is 730 days.
- In the CAGE 4 simulation, reinforcement-learning defenders outperformed LLM-based defenders.Last evidence 507 days ago; half-life is 365 days.
- Published prompt-injection defenses report attack success cut to near zero, or under 10%, against most of the fixed attacks their authors tested.Last evidence 494 days ago; half-life is 365 days.
- Separating an agent's control flow from untrusted data can give provable protection against control-flow hijacking at a modest utility cost.Last evidence 551 days ago; half-life is 540 days.
From opening to study
An opening starts as a seed: a question and the signals behind it. An editor checks prior work and, if the question holds, drafts a brief in the Fide research-call format with a hypothesis, controls, and a claim boundary. A brief that Fide or a collaborator takes on becomes a call on the Fide research commons.