{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/silent-failures-agent-evaluation-2026/",
 "asOf": "2026-10-01",
 "id": "silent-failures-agent-evaluation-2026",
 "date": "2026-09-26",
 "datePrecision": "day",
 "title": "Silent Failures audits an indirect-prompt-injection benchmark and finds scoring defects that distort attack and defense results",
 "lane": "defense",
 "kind": "paper",
 "summary": "Independent researcher Animesh Shaw (arXiv v1, 26 September 2026) audits one unnamed indirect-prompt-injection benchmark and describes four defect classes: payloads silently not delivered, attack success scored by tool identity rather than arguments, false rejection conflated with model incapacity, and no audit trail. Re-scoring recorded traces from the author's own runs, the tool-identity scorer gives 21.7% attack success against 1.2% at argument level, and a model reported at 62.8% by the audited harness scores 0% on the corrected one, a gap the author attributes in substantial part to undelivered payloads and tool-identity scoring. The author also releases a harness that rejects measurement-invalid scenarios before running, and says the results use static attacks in a five-tool simulation and are not robustness claims.",
 "whyItMatters": "It documents in one audit how scoring attack success by tool name inflates prompt-injection results (21.7% against 1.2% on the author's traces) and how undelivered payloads distort them, which bears on how far published attack and defense numbers can be compared.",
 "actors": [],
 "topics": [
  "eval-validity",
  "prompt-injection",
  "tool-and-mcp-security"
 ],
 "atlas": [
  "eval-environment",
  "tools",
  "untrusted-content"
 ],
 "artifacts": [],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.32691",
   "publisher": "arXiv",
   "title": "Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection",
   "date": "2026-09-26",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "Version and authorship: figures are from arXiv v1 (26 September 2026), a single-author paper by an independent researcher. The audited benchmark and its earlier reported numbers are not named or cited in the paper.",
   "locator": "Title page; Sec. IV"
  },
  {
   "fact": "Defect D1, silent payload non-delivery: the audited harness inserted an injection only when a free-text data-source field matched one of four hardcoded strings, so 39 of 43 attack payloads (91%) were never delivered yet each scenario still counted as an attack trial. Table I gives payloads delivered as 4/43 under the audited definition and 43/43 corrected.",
   "locator": "Sec. IV D1; Table I"
  },
  {
   "fact": "Defect D2, attack success by tool identity: success was recorded whenever an attacker-target tool appeared in the execution log, so an agent that resisted the injection but legitimately used a shared tool (for example a SQL query tool) was scored as compromised. D3: environment errors and model incapacity were charged to the defense as false rejections, including a high false-rejection rate for an undefended baseline. D4: only aggregate percentages were persisted, so published figures could not be reproduced.",
   "locator": "Sec. IV D2 to D4"
  },
  {
   "fact": "Ablation on identical traces: tool-identity scoring reports 21.7% attack success against 1.2% at argument level (+20.5 percentage points; 53 of 258 attack runs misclassified), and unconditioned false-rejection rate 10.3% against 6.7% corrected (+3.5 points over 282 benign runs). The table note says 'both models' without naming them; the paper's two full evaluations use gpt-5.6-terra and llama3.1:8b. D1 cannot be re-scored from traces and needs a paired sweep.",
   "locator": "Sec. VII-A; Table I"
  },
  {
   "fact": "Harness design: each scenario declares its own mailbox, filesystem and SQL tables; a structured injection locator that must resolve or the run aborts; and an argument-level attacker predicate that fires only when one executed call matches the tool and its argument constraints. Before any run the loader rejects suites with unresolvable locators, unreachable payloads, attacker goals that can never fire, or benign tasks the environment cannot satisfy. Every run writes a manifest (model, defense, suite hash, commit, package versions) and one self-describing trace per scenario.",
   "locator": "Sec. V-A to V-E"
  },
  {
   "fact": "Metrics re-derive AgentDojo's triple of benign utility, utility under attack and targeted attack success: attack success rate (ASR) at tool-and-argument level, utility-preservation rate (UPR), false-rejection rate (FRR = 100% minus benign UPR) and balanced accuracy. The paper states it attributes the design to AgentDojo and does not claim it as novel.",
   "locator": "Sec. II-A; Sec. V-C"
  },
  {
   "fact": "Setup: gpt-5.6-terra plus three local open models (llama3.1:8b, gemma4:12b, qwen2.5-coder:7b); a 95-scenario core suite (43 attack, 52 benign), a 55-scenario reduced suite (43 attack, 12 benign) and a 10-scenario capability probe; five mock tools including a real in-memory SQL engine; three baseline defenses (deterministic TypeChecker, intent-based CapabilityRouter, LLMJudge). Greedy decoding where available, Fisher exact tests with Holm-Bonferroni correction, Wilson intervals.",
   "locator": "Sec. III-D; Sec. V-F; Sec. VI"
  },
  {
   "fact": "Frontier model, 95-scenario suite (Table II): undefended ASR 2.3% (95% CI 0.4 to 12.1; one of 43 attacks), FRR 7.7%. TypeChecker 2.3% / FRR 5.8%; CapabilityRouter 0.0% (CI 0.0 to 8.2) / FRR 5.8%; LLMJudge 2.3% / FRR 11.5%. No defense difference is significant (Holm-corrected p=1.0; Cohen's h at most 0.31), and the paper computes that detecting a 60% to 35% reduction at 80% power needs 62 attack scenarios per condition, so the 43-attack stratum is underpowered.",
   "locator": "Sec. VI (statistical power); Sec. VII-B; Table II"
  },
  {
   "fact": "llama3.1:8b, 55-scenario suite: the audited harness had reported 62.8% ASR; the corrected harness gives 0% (0 of 43). The model reached the injected content in 41 of 43 attacks and completed the benign portion of attack scenarios 72% of the time, benign utility 91.7%. CapabilityRouter leaves ASR at 0 and raises FRR from 8.3% to 16.7% (benign utility 83.3%). The author says this does not show the model is robust, only that the reported vulnerability does not survive a valid harness; the text also refers to 38 undelivered payloads where the audit count elsewhere is 39.",
   "locator": "Sec. VII-D; Table III; Sec. IX"
  },
  {
   "fact": "Capability probe: gpt-5.6-terra, llama3.1:8b and gemma4:12b emit valid tool calls on 10/10 trivial tasks; qwen2.5-coder:7b fails 0/10. The paper attributes an earlier report of a 100% false-rejection 'tool-binding barrier' for gemma4:12b to the D3 environment mismatch, and reports security metrics for qwen as unmeasurable rather than secure.",
   "locator": "Sec. VII-C"
  },
  {
   "fact": "Disclosure: in the frontier traces the single successful attack was concealed, because the agent reported only that the benign task succeeded. The author calls this a directional finding, resting on one attack.",
   "locator": "Sec. VII-E"
  },
  {
   "fact": "Judge calibration: over 367 judge decisions on the frontier LLMJudge cell (AUC 0.63), the judge's binary verdict blocked 6% of legitimate calls and caught none of the sole attack; high-confidence flags had a near-zero observed adversarial rate. The author concludes that comparing judge defenses at default thresholds measures threshold placement, and notes the analysis rests on one adversarial call.",
   "locator": "Sec. VII-F"
  },
  {
   "fact": "Stated limits: one tool domain and five tools; the 43-attack stratum is underpowered for frontier defense comparisons; simulated environments; static attacks not adaptively optimized against each defense, so no result is a robustness claim; disclosure and judge-calibration findings underpowered on the frontier; measuring defect incidence across published benchmarks is left as future work.",
   "locator": "Sec. IX"
  },
  {
   "fact": "Release and disclosure: harness, defenses, metrics, feasibility checks and analysis scripts are stated to be open source, while the full attack suites and per-scenario results are withheld until the archival version is public. The author discloses AI coding and drafting assistance and states the author verified all numbers against the persisted traces.",
   "locator": "Sec. X; Sec. XI"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-075"
 ],
 "methods": [
  "ai-monitoring",
  "capability-restriction",
  "indirect-prompt-injection"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}