Organizations/academic

Pennsylvania State University

1 records1 attackWebsite
Sep 29, 2026
Curriculum-trained RL attacker reaches 45.0% ASR@10 on GPT-5.6-Terra where direct RL training gets 0%
AttackPaperPurdue University, Pennsylvania State University

Researchers at Penn State and Purdue (arXiv v2, dated 2026-09-29; the v1 date is not in the archived text) report a way to train a reinforcement-learning prompt-injection attacker against frontier targets, where direct training finds no successful attack and so receives no reward. Their curriculum trains one attacker model against a sequence of increasingly robust targets, and they report ASR@10 of 93.8% against GPT-5.6-Luna and 45.0% against GPT-5.6-Terra on AgentDyn, where PISmith and RL-Hammer trained directly score 0%. The authors also report that the attacker transfers to six targets it was not trained on and to AgentDojo.