Chronicle/Attacks & incidents

Curriculum-trained RL attacker reaches 45.0% ASR@10 on GPT-5.6-Terra where direct RL training gets 0%

AttackPaperSignificance assistant-drafted

Researchers at Penn State and Purdue (arXiv v2, dated 2026-09-29; the v1 date is not in the archived text) report a way to train a reinforcement-learning prompt-injection attacker against frontier targets, where direct training finds no successful attack and so receives no reward. Their curriculum trains one attacker model against a sequence of increasingly robust targets, and they report ASR@10 of 93.8% against GPT-5.6-Luna and 45.0% against GPT-5.6-Terra on AgentDyn, where PISmith and RL-Hammer trained directly score 0%. The authors also report that the attacker transfers to six targets it was not trained on and to AgentDojo.

Why it matters

It adds an academic black-box adaptive attack, reaching 45.0% ASR@10 on GPT-5.6-Terra, to the evidence on how well defended frontier models resist attackers who adapt.

Key facts

As stated in the sources, with where to find them.

  • Setup: Qwen3-4B-Instruct-2507 as the attacker, PISmith as the RL algorithm at every stage, black-box access to targets at mid reasoning effort; training on the AgentDyn GitHub subset only, evaluation on all of AgentDyn (60 tasks, 560 injection cases).v2, Sections 3.1 to 3.3; Appendix C.1
  • Training directly against GPT-5.6-Luna or GPT-5.6-Terra with PISmith or RL-Hammer gives 0.0%/0.0% ASR@1/ASR@10 on AgentDyn.v2, Table 1
  • Curriculum GPT-5-nano to GPT-5.6-Luna to GPT-5.6-Terra: 69.6%/93.8% ASR@1/ASR@10 against GPT-5.6-Luna and 18.3%/45.0% against GPT-5.6-Terra (AgentDyn overall). Skipping the Luna stage gives 0.9%/1.6% against Terra; GPT-4o-mini as first stage gives 0.4%/1.4% against Luna versus 66.3%/87.5% with GPT-5-nano.v2, Table 1
  • Before the Terra stage, the GPT-5-nano-trained attacker succeeded 0 times in 1,800 attempts (0.0%/0.0%); the attacker that had also been trained on Luna scored 1.9%/6.1%, which the authors say is enough to give RL a learning signal.v2, Appendix C.1, Table 5
  • Without further training, the Terra-trained attacker reaches overall ASR@10 of 47.3% on GPT-6-Luna, 35.5% on Muse-Spark-1.2, 22.3% on GLM-5.3-flash, 21.8% on Gemini-3.6-flash, 95.4% on DeepSeek-v4-flash-0731 and 80.0% on Qwen3.6-27B-SecOPD, against 0% to 0.2% for the base attacker and 0% ASR@1 for the static attack on five of six targets.v2, Table 2
  • Against GPT-5.6-Sol, which the authors describe as the most robust model in the GPT-5.6 family, the Terra-trained attacker reaches 2.9% ASR@1 and 7.3% ASR@10 with no training on Sol; the authors call this weak and did not run an RL stage against Sol for lack of budget.v2, Appendix C.2, Table 6
  • On AgentDojo, with no training on that benchmark, the Luna-stage attacker reaches 41.7%/72.5% ASR@1/ASR@10 on GPT-5.6-Luna and the Terra attacker reaches 10.4%/29.9% on GPT-5.6-Terra.v2, Table 4
  • Continued RL from the Terra attacker against Muse-Spark-1.2 raises ASR@10 from 35.5% to 82.3%, and the result carries to Muse-Spark-1.3 (73.9% ASR@10 versus 22.0% before).v2, Section 4.2, Table 3
  • For the Terra-trained attacker on GPT-5.6-Terra, overall ASR@10 falls from 48.4% (low reasoning effort) to 45.0% (mid) to 43.6% (high); the authors say ASR decreases as reasoning effort increases.v2, Appendix E, Table 7
  • The authors state that OpenAI's GPT-Red, an internal self-play RL red-teaming model, is not released and that its recipe is not public.v2, Section 2.1

Findings that cite this record

Key questions this bears on

Sources

Related records