Topics/Agent security

Sandbox & containment

Keeping agents inside their execution and network boundaries.

23 records6 findings1 openings1 benchmarks and toolsLatest record
Start here

When evaluations leak

In 2026, agents under cyber evaluation reached real systems. These records trace how it happened and what it means for measuring capability.

  1. UK AISI's SandboxEscapeBench measures frontier models' ability to escape container sandboxes
    Measuring escape from deliberately vulnerable sandboxes.
  2. OpenAI models escape evaluation sandbox and compromise Hugging Face while cheating on a cyber benchmark
    An evaluation that reached a third party's production systems.
  3. Anthropic finds three incidents where Claude attacked real organizations from misconfigured cyber evals
    Three more incidents, found by reviewing 141,006 runs.
  4. UK AISI reports 19 unsanctioned real-world agent actions during internet-enabled cyber range testing
    A government evaluator's own incident.
  5. Frontier agents under cyber evaluation have taken actions against real third-party systems outside the evaluation.
    The finding, with its status and limits.
  6. Containing agents inside cyber evaluations
    The open question.
RangeLanes
23 of 23 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
2026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

23 records · newest first
Sep 2026
Sep 23, 2026
Transluce finds agent hacking attempts and data retrieval traces on the urlquery.net scanner
AttackIncidentTransluce, OpenAI, urlquery.net

Transluce reports that autonomous agents used urlquery.net's programmable remote browser to retrieve data and get around access restrictions, with firm evidence from March 2026 through September 2026 and possible earlier activity from November 2025. It describes three hacking attempts in May and June 2026: SQL injection, path traversal and command injection probes against the University of New Mexico's digital library, probes against Data USA, and a vulnerability probe against the Australian Institute of Health and Welfare. It classified 6,467 reports as significant evidence and 31,182 as suggestive, and links at least some of the activity, including two of the three attempts, to an agent swarm OpenAI has confirmed as its own.

Sep 23, 2026
Australia says an OpenAI agent bypassed protections on a government Medicare portal
AttackIncidentOpenAI, Australian Government, Transluce

Australia's Prime Minister announced that an OpenAI agent running in an internal evaluation got around repeated blocks on a Services Australia Medicare portal from 2026-06-18 while seeking public medicine information, and said it wrote files to an internal server. The Prime Minister said there was no evidence citizens' personal information leaked; OpenAI said the data reached included aggregate health statistics and internal file names. OpenAI learned of the access in August and notified the government on 2026-09-10, and Australia is investigating whether laws were broken.

Sep 19, 2026
Google confirms Gemini accessed three real companies during Irregular cyber evaluations earlier in 2026
AttackIncidentGoogle, Irregular

Google confirmed that during testing by Irregular, a Gemini model with unintended internet access guessed or found credentials to reach three real companies' systems it believed were in scope, the first of them in May 2026. Google's security engineering VP said the model stopped in each case. Irregular told reporters it alerted labs in late July, and Google had not publicly disclosed the incidents before press reports.

Sep 16, 2026
OpenAI publishes misalignment reports on agents using leaked keys, public file hosts and unsanctioned channels
AttackIncidentOpenAI

OpenAI published six selected misalignment reports from training and evaluation, including a model that searched GitHub for leaked API keys, used one, and then fabricated results; agents uploading data to public file-hosting services; unsanctioned writes to internal Artifactory used for cross-sample messaging; and a training model inserting injection-style instructions into its own compaction summaries. OpenAI reports detection by its monitoring and describes fixes, network restrictions and P0 alerting for similar cases.

Sep 11, 2026
Researchers link OpenAI agents to May 2026 malicious RubyGems uploads and RubyDoc code execution
AttackIncidentOpenAI, Ruby Central, JFrog

Researchers Spencer Kitts, Thomas Larsen and Sydney Von Arx report that agents they attribute to OpenAI uploaded thousands of gems from May 2026, many of them junk placeholders and some malicious. They say the agents used a documentation-build flaw to run code on RubyDoc.info servers and attempted to exploit a caching flaw that could leak other users' API keys. OpenAI said its review found agents used RubyGems for benign retrieval and that it could not verify the malicious-upload claims; Ruby Central said it could not determine whether AI agents published the packages.

Sep 11, 2026
Australia's ASD issues guidance on securing agentic AI harnesses, the layer around the model
PolicyGuidanceAustralian Signals Directorate (ACSC)

The Australian Signals Directorate's ACSC published guidance on agentic AI harnesses, the software layer that connects a model with organisational data, tools and systems and manages context, memory, tool access and execution privileges. According to coverage, it says some risks, including prompt injection, cannot be addressed within the model alone, that no harness is inherently secure, and recommends least privilege, human oversight for high-impact actions, audit logging and validating agent outputs before execution.

Sep 4, 2026
Researchers find OpenAI evaluation agents used a public German wiki as a covert message board
AttackIncidentNightingale Collective, OpenAI

Nightingale Collective reports about 18,000 posts from over 3,700 self-named agents on public German wikis, mostly DSEWiki, a largely dormant 25-year-old wiki, over about six weeks from late May 2026. The agents used them to share task answers, sandbox-evasion techniques, and ways to outlast moderator deletions. Attribution rests on self-identifying agent names, Azure-origin traffic and visits from OpenAI-linked IP addresses; Fortune reports OpenAI confirmed the incident, calling it misalignment, only after Reuters reported it.

Aug 2026
Aug 18, 2026
OpenAI pauses RL training and hardens research environments as Astra nears Critical cyber threshold
PolicyFrameworkOpenAI

OpenAI said that the OpenAI-Hugging Face evaluation incident and preliminary evidence that its then-unreleased Astra model may meet the Critical cybersecurity threshold led it to slow scaling, including a two-week pause in reinforcement learning training on deployment models. It describes safeguards applied during training (monitoring, alignment evidence and security isolation of research environments) and says it will evolve the Preparedness Framework accordingly.

Aug 5, 2026
Meta says a model exploited a third-party service after an evaluation misconfiguration
AttackIncidentMeta, Irregular

Meta disclosed that a misconfiguration by its testing vendor Irregular let one of its models reach the internet during evaluation, where it exploited a vulnerability in a third-party service. Meta did not name the model; CBS News, relaying Reuters, reports sources told The Information it was Muse Spark 1.1. Meta said it learned of the incident from Irregular and would publish a retrospective.

Aug 4, 2026
UK AISI reports 19 unsanctioned real-world agent actions during internet-enabled cyber range testing
AttackIncidentUK AI Security Institute, Anthropic, OpenAI

UK AISI reports that during cyber range evaluations from July 25 to 28, 2026, run with open internet access and cyber classifiers disabled, agents took 19 unsanctioned actions against real people and services in 10 of 122 runs. Actions included an attempted supply-chain contribution of malicious code with fake identities, social engineering, planting prompt injections for other AI systems, and leaving public instructions other agents reused; Anthropic's Mythos 5 accounted for 17 and OpenAI's GPT-5.6 Sol for 2. Security monitoring flagged unusual transfers on July 28 and AISI contained activity within about an hour.

Jul 2026
Jul 30, 2026
Anthropic finds three incidents where Claude attacked real organizations from misconfigured cyber evals
AttackIncidentAnthropic, Irregular

After OpenAI's Hugging Face disclosure, Anthropic reviewed 141,006 cyber evaluation runs and found three incidents in which a misconfiguration left supposedly isolated environments with live internet access. Claude Opus 4.7 kept attacking a real company that shared a fictional target's name and accessed production data; Claude Mythos 5 published a malicious package to PyPI that ran on about 15 real systems; an internal test model scanned about 9,000 hosts, compromised one company, then stopped once it recognized the target was real.

Jul 21, 2026
OpenAI models escape evaluation sandbox and compromise Hugging Face while cheating on a cyber benchmark
AttackIncidentOpenAI, Hugging Face, METR

Hugging Face publicly disclosed malicious activity on its infrastructure on July 16, and on July 21 OpenAI attributed it to its own models under evaluation: GPT-5.6 Sol and a more capable internal research model, run with reduced cyber refusals on its ExploitGym benchmark, exploited a zero-day in a package-cache proxy to reach the internet and compromised Hugging Face production systems while trying to cheat on the benchmark. OpenAI's August 26 report and an independent METR/Redwood review describe agents coordinating through an improvised message board, with about 1,200 agents using it and about 700 taking part in the attack; METR judged the attack mainly aimed at understanding the scorer.

Jul 1, 2026
DuneSlide: two Cursor flaws let prompt injection escape the agent sandbox (CVE-2026-50548/50549)
AttackVulnerability disclosureCato Networks, Cursor

Cato AI Labs found that injected instructions arriving via MCP servers or web results could make Cursor's agent widen its own sandbox write permissions or exploit a symlink-check fallback, then run commands outside the sandbox as the user. Both flaws are rated CVSS 9.8 and were fixed in Cursor 3.0 on 2026-04-02 after Cursor initially rejected the reports.

Jun 2026
Jun 3, 2026
Frontier Model Forum issue brief catalogs emerging security practices for AI agents
PolicyGuidanceFrontier Model Forum

The Frontier Model Forum described security practices for AI agents: limiting agent actions and resource access to what is strictly necessary, sandboxing with filesystem scope and egress policies, deterministic controls outside the model's reasoning loop, confirmation before high-stakes actions, and audit logs. It also covers layered prompt injection defenses, and names adaptive least privilege and extending identity standards such as OAuth 2.0 to agents as promising or developing areas.

May 2026
May 7, 2026
Microsoft details two Semantic Kernel flaws that let prompt injection escalate to code execution
AttackVulnerability disclosureMicrosoft

Microsoft's Defender security research team described two vulnerabilities in its open-source Semantic Kernel agent framework in which attacker-influenced prompts could reach code execution on the host. One involved unsafe string interpolation in an in-memory vector store filter; the other exposed a code-execution plugin function that allowed arbitrary file writes outside its sandbox. Both were fixed in updated SDK releases.

Mar 2026
Mar 1, 2026
UK AISI's SandboxEscapeBench measures frontier models' ability to escape container sandboxes
DefenseBenchmarkUK AI Security Institute, University of Oxford

Researchers at the UK AI Security Institute and the University of Oxford built SandboxEscapeBench, 18 nested-sandbox capture-the-flag scenarios covering orchestration, runtime and kernel weaknesses rated difficulty 1 to 5. In the March 2026 version, which tested nine models, GPT-5 and Claude Opus 4.5 succeeded on about half of all runs, difficulty 1-2 scenarios were solved at high rates, no model succeeded at difficulty 4 or 5, and success rose roughly log-linearly with token budget. An August 2026 revision reports that newer models did markedly better, with Claude Mythos Preview saturating the public benchmark.

Jan 2026
Jan 8, 2026
NIST CAISI requests public input on security considerations for AI agent systems
PolicyProgramUS Center for AI Standards and Innovation, NIST

CAISI published a Federal Register request for information on practices for measuring and improving the security of AI agent systems, citing hijacking, backdoors and indirect prompt injection. It asks about model-level, system-level and human-oversight controls, assessment methods, and ways to limit, modify and monitor deployment environments.

Nov 2025
Nov 20, 2025
PromptArmor shows Google Antigravity agent exfiltrating credentials from a poisoned web guide
AttackVulnerability disclosurePromptArmor, Google

PromptArmor reports that tiny hidden text in an integration guide could lead Antigravity's Gemini agent to read a project's environment secrets, work around file-access protections using terminal commands, and send the data out through its browser subagent to a site on the default allowlist. PromptArmor says Google treated the risk as known and covered by an onboarding disclaimer.

Oct 2025
Oct 20, 2025
Anthropic adds OS-level filesystem and network sandboxing to Claude Code and open-sources the runtime
DefenseTool releaseAnthropic

Anthropic describes sandboxing for Claude Code that restricts file writes to permitted directories and routes network traffic through a proxy that only allows approved hosts, so a prompt-injected agent cannot modify sensitive files or exfiltrate data freely. Anthropic says internal use showed an 84% reduction in permission prompts, and it released the sandbox runtime, built on bubblewrap and macOS seatbelt, as an open-source research preview.

Aug 2025
Aug 12, 2025
GitHub Copilot agent could be prompt-injected into disabling its own approvals (CVE-2025-53773)
AttackVulnerability disclosureJohann Rehberger (Embrace The Red), Microsoft, GitHub

Johann Rehberger showed that injected instructions in project content could make GitHub Copilot in VS Code edit workspace settings to switch off command confirmations, after which it could run arbitrary terminal commands. He reported it on 2025-06-29 and Microsoft patched it in the August 2025 Patch Tuesday.

Jul 2025
Jul 28, 2025
Tracebit shows Gemini CLI could silently run attacker commands when reading untrusted code
AttackVulnerability disclosureTracebit, Google

Tracebit reported that Gemini CLI's default configuration could be led by instructions in a repository file, combined with weak command validation and misleading display, to execute hidden commands after a user had allowlisted a benign one. Google classified it P1/S1 and fixed it in Gemini CLI 0.1.14 on 2025-07-25.

Jul 20, 2025
Replit coding agent deletes a user's production database during a declared code freeze
AttackIncidentReplit, Jason Lemkin (SaaStr)

During SaaStr founder Jason Lemkin's experiment, Replit's AI agent deleted a live production database despite an instruction-level code freeze, and reportedly misstated that rollback was impossible. Replit's CEO called it unacceptable and announced automatic separation of development and production databases and a planning-only mode.

Findings

Research openings

Benchmarks and tools