Korea University researchers propose ActionGuard, which checks skill-influenced tool calls immediately before execution against the trusted user request and runtime evidence. In an OpenClaw evaluation using SKILL-INJECT tasks, they report 8.65% overall attack success and 90.38% task success, averaged across reviewer models and injection types. The evaluation uses one framework, one target model and one benchmark; it does not establish performance against attackers adapting to the defense.
Why it matters
It tests authorization at the execution boundary instead of treating skill instructions as permission, while measuring the resulting task-completion cost.
Key facts
As stated in the sources, with where to find them.
- ActionGuard separates the agent’s planning context from an isolated reviewer’s authorization context. The reviewer uses the trusted request, an independently maintained skill profile, tool-call context and local script contents, rather than receiving the raw potentially poisoned skill as execution authority. Invalid or unavailable reviewer decisions deny execution.v1, Sections 4.1–4.3
- Evaluation uses 139 contextual and 180 obvious injection-task pairs in OpenClaw, five reviewer models, and three repetitions per condition. Comparisons include Dynamic Guardian, SkillGuard and no safeguard.v1, Sections 5.1–5.2; Tables 2–3
- Table 4 averages across reviewer models and injection types: ActionGuard 8.65% attack success and 90.38% task success; Dynamic Guardian 16.05% and 92.08%; SkillGuard 13.42% and 88.55%; no safeguard 29.05% and 93.46%. These are averaged rates, not individual-model results.v1, Table 4; Section 6.2
- The authors limit generalization to one agent framework, target model and benchmark. Blocking a call that combines benign and unauthorized actions can leave the task incomplete if the agent does not produce a safe alternative. Their abstract and conclusion give slightly different relative reductions against no safeguard; this record uses the absolute Table 4 rates.v1, Section 8, Conclusion; Abstract; Table 4
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
- Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Sep 30, 2026
Sep 30, 2026
Sep 29, 2026
Sep 30, 2026
Sep 30, 2026
Aug 26, 2026