Desk/2025-W52

Week of Dec 22–28, 2025

1 records0 status changes on new evidence0 new findings

Defense & research

Dec 22, 2025
OpenAI hardens ChatGPT Atlas with an RL-trained automated prompt injection attacker
DefenseFrameworkOpenAI

OpenAI describes an LLM-based attacker trained end-to-end with reinforcement learning that searches for prompt injections able to steer the Atlas browser agent through long, multi-step harmful workflows, and a rapid response loop that adversarially trains new agent checkpoints against discovered attacks. OpenAI says the attacker found strategies absent from human red-teaming and external reports, and states that prompt injection is unlikely ever to be fully solved.