{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/curriculum-rl-prompt-injection-red-teaming-2026/",
 "asOf": "2026-10-01",
 "id": "curriculum-rl-prompt-injection-red-teaming-2026",
 "date": "2026-09-29",
 "datePrecision": "day",
 "title": "Curriculum-trained RL attacker reaches 45.0% ASR@10 on GPT-5.6-Terra where direct RL training gets 0%",
 "lane": "attack",
 "kind": "paper",
 "summary": "Researchers at Penn State and Purdue (arXiv v2, dated 2026-09-29; the v1 date is not in the archived text) report a way to train a reinforcement-learning prompt-injection attacker against frontier targets, where direct training finds no successful attack and so receives no reward. Their curriculum trains one attacker model against a sequence of increasingly robust targets, and they report ASR@10 of 93.8% against GPT-5.6-Luna and 45.0% against GPT-5.6-Terra on AgentDyn, where PISmith and RL-Hammer trained directly score 0%. The authors also report that the attacker transfers to six targets it was not trained on and to AgentDojo.",
 "whyItMatters": "It adds an academic black-box adaptive attack, reaching 45.0% ASR@10 on GPT-5.6-Terra, to the evidence on how well defended frontier models resist attackers who adapt.",
 "actors": [
  "purdue-university",
  "pennsylvania-state-university"
 ],
 "topics": [
  "prompt-injection",
  "eval-validity"
 ],
 "atlas": [
  "untrusted-content",
  "model"
 ],
 "artifacts": [
  "agentdojo",
  "gpt-5-family",
  "gemini",
  "deepseek",
  "glm",
  "gpt-4-family"
 ],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.33628",
   "publisher": "arXiv",
   "title": "Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning",
   "date": "2026-09-29",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "Setup: Qwen3-4B-Instruct-2507 as the attacker, PISmith as the RL algorithm at every stage, black-box access to targets at mid reasoning effort; training on the AgentDyn GitHub subset only, evaluation on all of AgentDyn (60 tasks, 560 injection cases).",
   "locator": "v2, Sections 3.1 to 3.3; Appendix C.1"
  },
  {
   "fact": "Training directly against GPT-5.6-Luna or GPT-5.6-Terra with PISmith or RL-Hammer gives 0.0%/0.0% ASR@1/ASR@10 on AgentDyn.",
   "locator": "v2, Table 1"
  },
  {
   "fact": "Curriculum GPT-5-nano to GPT-5.6-Luna to GPT-5.6-Terra: 69.6%/93.8% ASR@1/ASR@10 against GPT-5.6-Luna and 18.3%/45.0% against GPT-5.6-Terra (AgentDyn overall). Skipping the Luna stage gives 0.9%/1.6% against Terra; GPT-4o-mini as first stage gives 0.4%/1.4% against Luna versus 66.3%/87.5% with GPT-5-nano.",
   "locator": "v2, Table 1"
  },
  {
   "fact": "Before the Terra stage, the GPT-5-nano-trained attacker succeeded 0 times in 1,800 attempts (0.0%/0.0%); the attacker that had also been trained on Luna scored 1.9%/6.1%, which the authors say is enough to give RL a learning signal.",
   "locator": "v2, Appendix C.1, Table 5"
  },
  {
   "fact": "Without further training, the Terra-trained attacker reaches overall ASR@10 of 47.3% on GPT-6-Luna, 35.5% on Muse-Spark-1.2, 22.3% on GLM-5.3-flash, 21.8% on Gemini-3.6-flash, 95.4% on DeepSeek-v4-flash-0731 and 80.0% on Qwen3.6-27B-SecOPD, against 0% to 0.2% for the base attacker and 0% ASR@1 for the static attack on five of six targets.",
   "locator": "v2, Table 2"
  },
  {
   "fact": "Against GPT-5.6-Sol, which the authors describe as the most robust model in the GPT-5.6 family, the Terra-trained attacker reaches 2.9% ASR@1 and 7.3% ASR@10 with no training on Sol; the authors call this weak and did not run an RL stage against Sol for lack of budget.",
   "locator": "v2, Appendix C.2, Table 6"
  },
  {
   "fact": "On AgentDojo, with no training on that benchmark, the Luna-stage attacker reaches 41.7%/72.5% ASR@1/ASR@10 on GPT-5.6-Luna and the Terra attacker reaches 10.4%/29.9% on GPT-5.6-Terra.",
   "locator": "v2, Table 4"
  },
  {
   "fact": "Continued RL from the Terra attacker against Muse-Spark-1.2 raises ASR@10 from 35.5% to 82.3%, and the result carries to Muse-Spark-1.3 (73.9% ASR@10 versus 22.0% before).",
   "locator": "v2, Section 4.2, Table 3"
  },
  {
   "fact": "For the Terra-trained attacker on GPT-5.6-Terra, overall ASR@10 falls from 48.4% (low reasoning effort) to 45.0% (mid) to 43.6% (high); the authors say ASR decreases as reasoning effort increases.",
   "locator": "v2, Appendix E, Table 7"
  },
  {
   "fact": "The authors state that OpenAI's GPT-Red, an internal self-play RL red-teaming model, is not released and that its recipe is not public.",
   "locator": "v2, Section 2.1"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-075"
 ],
 "methods": [
  "adaptive-red-teaming",
  "indirect-prompt-injection"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}