Researchers from the Chinese Academy of Sciences and Worcester Polytechnic Institute introduce TrustProbe to trace installed skill content into security-sensitive operations and verify observable effects. Using DeepSeek-V4-Flash across eleven open-source agents, they report 104 verified source-to-sink vulnerabilities in permissive non-interactive configurations, with a subset remaining exploitable under stricter approval settings. Their real-skill experiment measures exposure to vulnerable execution paths, rather than the prevalence of malicious skills or attacks on real users.
It distinguishes ordinary skill-path exposure from verified harmful behavior and tests how installation and approval mechanisms affect delegated authority.
Key facts
As stated in the sources, with where to find them.
- The main campaign runs eleven skill-supporting open-source agents in fresh sessions and isolated workspaces with permissive non-interactive execution settings. DeepSeek-V4-Flash supplies both target-agent inference and TrustProbe generation, scoring and mutation; this is not an eleven-model comparison.v1, Sections 5.1–5.2
- The authors report 104 verified taint-style vulnerabilities, counted once per audited source-to-sink path when attacker-controlled flow and observable harm are both confirmed. Under the strictest usable non-interactive approval settings, 31 of 89 applicable vulnerabilities across eight agents remain exploitable.v1, Section 5.2; Table 1; Appendices H–I
- Direct-prompt replay reproduces 33 of the 104 discovery-confirmed vulnerabilities (31.7%). This compares reproduction of previously discovered skill cases with prompt delivery, rather than equal-budget discovery campaigns.v1, Section 5.3; Table 2
- Among 2,963 compatible skill-agent runs using 633 real skill files, 743 (25.1%) trigger an audited vulnerable path. This exposure rate does not label the original skills malicious. The authors validate complete attacks after modifying 15 selected triggering skills, in isolated workspaces using controlled data and endpoints.v1, Section 5.4; Ethics statement
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
- Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.