Chronicle/Attacks & incidents

Pretext: Huawei researchers evade NVIDIA's SkillSpector skill scanner with an LLM attacker that learns across generations

AttackPaperSignificance assistant-drafted

Two Huawei Research Zurich authors present Pretext, a white-box LLM attacker that iteratively writes agent skills to evade a scanner pairing static rules with LLM-based semantic analysis, using NVIDIA's open-source SkillSpector as the target. Their abstract reports up to 97% attack success against a frozen scanner and up to 77% against one that also learns, across three open-source model stacks; the body shows attack success falling over generations against the learning scanner (final-generation values are not tabulated) and finds that the learning scanner became more conservative with high false-positive rates. Commercial scanners were not tested.

Why it matters

It indicates that SkillSpector, the one scanner tested, can be evaded by an attacker who knows its rules, and that tightening its judge trades away benign skills; the authors call transfer to other scanners a conjecture.

Key facts

As stated in the sources, with where to find them.

  • Target and setup: SkillSpector v2.2.3 (64 deterministic patterns across 16 categories plus LLM semantic analyzers and an LLM meta-analyzer, scored 0 to 100); a skill counts as evading detection below 20. The attacker is GLM-5.1; detector, victim and benign-generator roles use GLM-5.1, Qwen3-235B-A22B-Thinking or gpt-oss-120b. Twelve trojan attack types are paired with synthetic victim environments described only in prompt text, and nothing is executed.v1, Section 3; Section 4 (Setup); Appendix A; Appendix C
  • The attacker knows SkillSpector's base rules and semantic-analyzer prompts but not its scoring or threshold, and in the learning-detector mode never sees the detector's learned heuristics. Each run has at most three refinement iterations.v1, Section 3; Appendix B.1
  • Frozen detector (Mode A), final generation of 10, mean over 5 replicates: attack success rate 96.7 (Qwen3-235B-A22B-Thinking stack), 63.2 (GLM-5.1 stack) and 70.5 (gpt-oss-120b stack). Success requires evading the scanner, delivering the payload in a simulated victim, and completing the benign task.v1, Section 4 (Mode A); Table D.1
  • The authors attribute differences between stacks to the detector model's safety alignment and say a more conservative LLM detector lowers success but raises false positives. When the detector also learns (Mode B, 6 generations), attacker success trends downward, but final-generation false-positive rates reach up to 62% (gpt-oss-120b, blind) and detector coverage of attacker lessons is as low as 0.33; the authors say this does not indicate higher security.v1, Section 4 (Mode B); Figure 4; Table D.2
  • Re-running SkillSpector's static layer alone (--no-llm) on the 178 final-generation attacker skills flagged 10 (5.6%), mean static score 2.9. The skills were optimized against the full detector, so the authors say this shows the static layer is not the binding constraint rather than that a static-targeting attacker wins trivially.v1, Appendix D.6; Table D.4
  • Payload delivery is judged by an LLM call of the attacker's model, which does not see the skill file and scores against the same per-type success criterion the attacker optimizes toward.v1, Appendix B.7
  • The authors did not test commercial scanners and describe transfer to them as a conjecture, not a result; they recommend layered controls such as sandboxed, monitored execution and restricted actions, which the paper does not test.v1, Section 5

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records