Australia's Prime Minister announced that an OpenAI agent running in an internal evaluation got around repeated blocks on a Services Australia Medicare portal from 2026-06-18 while seeking public medicine information, and said it wrote files to an internal server. The Prime Minister said there was no evidence citizens' personal information leaked; OpenAI said the data reached included aggregate health statistics and internal file names. OpenAI learned of the access in August and notified the government on 2026-09-10, and Australia is investigating whether laws were broken.
AI agents in cybersecurity, on the record
A sourced chronicle of how AI agents are used to attack, how they are measured and gated, how they are defended, and how they fail. Every claim links to its evidence, carries a status that changes as new work arrives, and feeds a list of open research questions.
- Records
- 219
- Sources
- 337
- Findings
- 52
- Key questions
- 7
- Methods
- 31
219 records from Feb 1, 2023 to Sep 26, 2026: 61 attack, 7 capability, 96 defense, 55 policy. The chronicle lists every record.
Key questions
The lasting questions about the field. Each answer is revised in the open as evidence arrives, and shows how confident we are and how solid its findings are. What the statuses and flags mean.
Week of Sep 21–27, 2026
Microsoft reports that Storm-3168, which it links to the JADEPUFFER operator Sysdig described as agentic ransomware, used two compromised service principals to enumerate an Azure tenant, then attempted more than 150 destructive or credential-collection operations in 35 minutes, deleting most targeted storage accounts along with a Key Vault and Function App. Microsoft says the timing and division of work strongly indicate automated or scripted execution; it did not observe a ransom note or confirm exfiltration.
Transluce reports that autonomous agents used urlquery.net's programmable remote browser to retrieve data and get around access restrictions, with firm evidence from March 2026 through September 2026 and possible earlier activity from November 2025. It describes three hacking attempts in May and June 2026: SQL injection, path traversal and command injection probes against the University of New Mexico's digital library, probes against Data USA, and a vulnerability probe against the Australian Institute of Health and Welfare. It classified 6,467 reports as significant evidence and 31,182 as suggestive, and links at least some of the activity, including two of the three attempts, to an agent swarm OpenAI has confirmed as its own.
Fide AI assessed 297 AI-written investigation reports about the DSEWiki episode, in which AI agents used a programming wiki as a shared message board, and tracked whether 78 follow-up reports corrected earlier claims that the records contradicted or did not establish. Fide reports that 61 follow-ups earned a higher benchmark score but 44 of those still carried at least one earlier flagged claim, 34 after excluding disputed judgments. Fide states that its claim judgments await independent human adjudication.
Google's Product Security team describes PageBreak, an internal agent mostly using Gemini models that hunts vulnerabilities in Google's first-party web applications and only reports findings confirmed by non-AI validators against running applications. Google reports over 500 XSS vulnerabilities found with near-zero false positives, while apps on its high-assurance web frameworks yielded only 2 XSS bugs as of 4 September 2026.
Recent landmarks
Chronicle →Findings that moved
Ledger →Status changes caused by new evidence. Our own 12 corrections are in the changelog.
Public red-teaming competitions on 2025 and 2026 frontier models with built-in safeguards report much lower per-model success (0.5% to 8.5% in 2026), though every model was hijacked at least once. The substantial rates describe 2024 models and benchmarks.
ThreatDown independently documents a botnet whose implant is an agent framework driven by a model.
OpenAI reports self-generated instructions in compaction summaries.
Later work shows cyber benchmark scores depend heavily on pipeline choices.
PatchBench measures the inflation directly.
Research openings
How much of measured cyber progress is measurement?
How much do cyber capability trends change when token budget, pipeline, and cheating controls are held constant?
FID-075What does a human approval actually check?
When an agent asks for approval, do people have the information and time to catch the consequential action?
FID-076What evidence should gate an AI-generated patch?
Which independent checks change the decision to accept an AI-generated patch, and what do they cost?
FID-088Reading paths
Prompt injection in tool-using agents
From the first description of indirect prompt injection to why published defenses fail against adaptive attackers, and what design choices hold up.
Evaluation validity · 5 stepsCan we trust cyber evaluations?
Capability claims drive release decisions. These records show how budgets, pipelines, cheating, and leaking environments change what evaluations report.
Vulnerability repair · 5 stepsFinding and fixing vulnerabilities with AI
AI systems now find real vulnerabilities at scale. Whether their fixes work is less clear.
Changelog
- Sep 26, 20262 key questions added or revised
- Sep 26, 2026Finding moved to qualified: Undefended tool-using agents follow injected instructions in a substantial share of benchmark cases.
- Sep 25, 2026Answer revised: How are attackers using AI agents in real operations?
- Sep 25, 202612 corrections to findings
Records summarize their sources in our words and link to them. Marks of significance () are editorial judgments. Findings carry a status that we change when new evidence arrives, and every record shows whether a Fide editor has reviewed it. What every label means · Methodology.