Researchers at Singapore Management University, Nanjing University and Adelaide University propose Aletheia, which turns the permissions a repository rule file requests into sandbox settings and checks whether the agent still passes independent tests when each is withheld. With GPT-5.5 on one shared refactoring task, they report alarms on all 314 AIShellJack attack inputs, none on five benign templates, and three false positives among 80 benign GHAgentFiles rules. The paper states these figures do not establish a population false-positive rate.
It offers a way to detect malicious repository rules without recognizing payload wording, but the reported evidence comes from one task with a development set that informed the method.
Key facts
As stated in the sources, with where to find them.
- Method: an LLM translates the authority a rule requests into a typed permission language; each permission is removed in an independent run; passing independent tests under strictly reduced authority gives a task-relative dispensability witness, which an LLM then interprets against the task to raise an alarm. The authors say witnesses establish dispensability for the tested task, not malicious intent or a global minimum policy.v1, Abstract; Sections 2.1 to 2.4
- Setup: GPT-5.5 through OpenRouter, eight agent turns per branch and a ten-second command timeout, on one shared refactoring task with independently authored tests, for 314 AIShellJack attack prefixes, its five benign templates, and 80 manually verified benign GHAgentFiles rules.v1, Section 3
- Detection: all 314 attack inputs completed execution and raised an alarm (314 of 314); 0 of 5 benign templates alarmed; on GHAgentFiles, 3 of 80 alarmed (3.75%), 52 were completed negatives and 25 were unresolved (22 unsupported permission interpretations, 2 translation failures, 1 provider refusal), which are not confirmed negatives.v1, Section 3; Table 1
- In an earlier development comparison, task-context interpretation alone produced the same file-level alarms as the combined criterion, so the authors say execution adds tested authority-removal evidence without an established classification gain.v1, Section 3
- Cost: synthesis prepared all 1,112 configurations for the cached attack projections and found 26 ineffective deletions; across 24 completed GHAgentFiles runs with fresh translation, translation plus sandbox execution took a median of 44.35 seconds and US$0.22 per input, excluding task-context interpretation.v1, Section 3
- Stated limits: all 80 GHAgentFiles files informed method development, five templates and that set do not establish a population false-positive rate, the shared task may omit legitimate project obligations, and attacks that misuse authority the task already requires may yield no witness. The paper does not test attackers who adapt to the method or compare against baseline detectors.v1, Section 3; Section 4
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.