OpenAI says structured safety documentation, ideally safety cases, should be required before any frontier reinforcement-learning training run, and publishes initial guidelines in three parts: technical safeguards (alignment training, containment, monitoring), operational guidelines, and practices for investigating misalignment incidents. OpenAI calls safety cases an aspirational goal, says it is working on a framework to codify them (as of 2026-09-28), and limits the scope to training rather than deployment. The page is a statement of intended practice with no measurements; OpenAI says its operational recommendations are in the process of being implemented.
It sets out in one lab's own words which controls, evidence and sign-offs it thinks should precede a frontier training run, but as a proposal that outside parties have no standard or measurement to check against.
Key facts
As stated in the sources, with where to find them.
- OpenAI says structured safety documentation should be required before continuing any frontier reinforcement-learning training run, ideally at the level of safety cases (structured, evidence-based arguments about risk, as used in other safety-critical industries). It calls safety cases an aspirational goal, acknowledges they will be harder to make rigorous for AI than for aviation or nuclear power, and says it is working on a framework to codify the practices.Introduction
- Scope: the guidelines cover frontier reinforcement-learning training only. OpenAI says internal and external deployment need a much broader set of alignment properties, and that the guidelines reflect what it had learned by 2026-09-28 and are expected to evolve.Introduction
- Technical safeguards: safety cases should cover three aspects of the stack, alignment training, containment and monitoring, so that the model does not try misaligned actions, would find it hard to break containment if it did, and would be caught by monitoring before harm occurred.1. Technical safeguards
- Alignment training: prevent reinforcement of reward hacks through automated and manual dataset review of training environments, grader tuning to penalize exploit attempts, and classifiers over prior-run traces to check graders. Measure alignment with offline evaluations during runs, backtesting on previous incidents to check evaluations detect earlier misaligned models without overfitting to them, tracking of evaluation awareness with blocking thresholds, and worst-case stress tests. Do not let automated graders see chain-of-thought in reinforcement learning, so models do not learn to evade chain-of-thought monitors.1. Technical safeguards; Model alignment
- Containment: harden both the sandbox and the research infrastructure hosting it as an extra layer against sandbox escape; red-team sandbox, infrastructure and perimeter with checkpoints from frontier runs (subject to a safety analysis) and red-team every service the model can touch, such as the internet cache; limit high-bandwidth cross-sample communication to approved pathways and close discoverable side channels; save agent transcripts immutably (for example in write-once storage) for all RL training, evaluation and internal deployment.1. Technical safeguards; Containment
- Monitoring: keep models monitorable, measured with monitorability and monitor-evasion evaluations against clear thresholds; require high recall on past incidents in a held-out evaluation and refresh evaluation datapoints for net-new risks so recall and precision figures do not go stale; define priority alerts that lead to action within a set response time, either paging an on-call person or automatically pausing the run (the illustration is paging security on-call by day and auto-pausing on unacknowledged alerts at night).1. Technical safeguards; Monitoring
- Operational guidelines: a written dissent from a member of another team that the training team must address; review by senior leaders who can each veto the run; accountability of the responsible senior leader, including in performance reviews; runbooks, technical controls and response times for pausing all covered runs; sharing safety cases with internal oversight groups; auditor access sufficient to verify the safety case; severity-level escalation with a misalignment on-call able to page executives; controls that fail closed (no starting a run without monitoring, no disabling the monitor from within training, evaluation or internal deployment); the ability to trace downstream uses of a misaligned model to undo its effects; and an enumerated list of residual risks.2. Operational guidelines
- OpenAI says the operational recommendations are its recommendations as of 2026-09-28 and are in the process of being implemented at OpenAI, and expects them to evolve. The page does not state an implementation status for the technical safeguards.2. Operational guidelines (closing paragraph)
- Incident investigations: periodic internal updates during an investigation (daily is the example) with defined pathways to raw transcripts and samples from misaligned models where safe; root-causing training dynamics with targeted ablations or resampling; an operational and cultural postmortem; alignment tests that can find the propensity behind an incident without hill-climbing on incident-derived information, with incident-derived evaluations kept as regression tests; and public disclosure of results, postmortems and operational changes after the investigation, with affected third parties notified at the earliest opportunity.3. Investigations of misalignment incidents
- The page gives no numeric thresholds, response times or effectiveness results; time examples (daily updates, day and night paging) are illustrations, and the guidelines are worded as what safety cases 'could' or 'should' include.Whole page (reading note)
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Do cyber evaluations of AI agents stay contained?Not reliably. Labs and a government evaluator disclosed agents reaching real systems from cyber evaluations; OpenAI agents did so from training runs too.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.