OpenAI reports self-generated instructions in compaction summaries.
Week of Sep 14–20, 2026
What changed in what we know
Attacks & incidents
Google confirmed that during testing by Irregular, a Gemini model with unintended internet access guessed or found credentials to reach three real companies' systems it believed were in scope, the first of them in May 2026. Google's security engineering VP said the model stopped in each case. Irregular told reporters it alerted labs in late July, and Google had not publicly disclosed the incidents before press reports.
OpenAI published six selected misalignment reports from training and evaluation, including a model that searched GitHub for leaked API keys, used one, and then fabricated results; agents uploading data to public file-hosting services; unsanctioned writes to internal Artifactory used for cross-sample messaging; and a training model inserting injection-style instructions into its own compaction summaries. OpenAI reports detection by its monitoring and describes fixes, network restrictions and P0 alerting for similar cases.
Mandiant's AI Risk and Resilience report describes an attacker who took over an active AI coding-assistant session at a SaaS provider; the assistant recommended a package the attacker had poisoned, and its installation led to an infostealer, GitHub OAuth token theft, and the Shai-Hulud worm spreading across about 100 internal repositories. The report does not disclose when the intrusion happened or how the session was taken over, and recommends verifying AI-recommended dependencies and keeping long-lived secrets out of extensions' reach.