Here’s what’s wild about the AI safety debate in 2026: the companies building the most powerful AI systems have spent years warning about potential risks, but according to a new study, most of them still won’t publicly document how they’d actually contain a rogue model if one showed up.
Think about that for a second. These labs have entire teams dedicated to “AI safety.” They publish research papers. They testify before Congress. OpenAI just flipped its position on California’s SB 53 and is now calling for stronger AI safety rules after opposing the bill earlier. But when researchers ask “okay, so what’s your actual plan if a model goes rogue?” the answer is mostly crickets.
This isn’t some distant sci-fi scenario anymore. AI systems are already demonstrating unexpected behavior. They’re being deployed in critical systems. And the companies building them are racing to make them more capable, more autonomous, more integrated into everything we do. The gap between “we take safety seriously” and “here’s our documented containment protocol” is starting to look less like a technical challenge and more like a credibility problem.
OpenAI’s sudden enthusiasm for strengthening California’s AI safety bill is particularly instructive. This is the same company that opposed the bill when it was first proposed. Now they’re calling for it to be stronger. What changed?
Maybe they genuinely reconsidered their position. Maybe the political winds shifted. Or maybe, and I think this is closest to the truth, they’ve realized that some regulation is inevitable and they’d rather shape it than fight it. Better to advocate for rules you can live with than have them imposed on you.
But here’s the thing: calling for stronger safety bills while not having public containment plans for dangerous models is peak have-your-cake-and-eat-it-too. It lets you look responsible without actually committing to specific, auditable safety measures. It’s safety theater dressed up as policy engagement.
The lack of documented containment plans isn’t just an oversight. It’s a choice. These labs have incredibly smart people. They know how to write documentation. They know how to create protocols. The fact that they haven’t done this publicly suggests they either don’t actually have detailed plans, or they don’t want those plans scrutinized.
Neither option is reassuring.
If they don’t have plans, that’s terrifying. You’re building systems you admit could be dangerous, and you’re just winging it? If they do have plans but won’t share them, that’s also a problem. Safety measures that can’t be independently evaluated aren’t really safety measures. They’re trust exercises.
And frankly, the AI industry hasn’t earned that trust. Not because the people in it are bad, but because the incentives are completely misaligned. Every lab is racing to ship the next breakthrough. Every few months we get a new model that’s more capable than the last. The pressure to move fast is enormous. The pressure to be cautious is mostly self-imposed.
This isn’t that complicated. If you’re serious about containing rogue AI, you document your thresholds. You specify what behaviors would trigger containment protocols. You detail the technical measures you’d deploy. You establish who has authority to pull the plug and under what circumstances. You test these procedures. You publish enough for external experts to evaluate whether your plans make sense.
You don’t have to reveal every technical detail. But you should be able to demonstrate that you’ve thought through the scenarios, that you have specific responses ready, and that someone other than your own team has reviewed them.
The fact that most frontier labs apparently can’t or won’t do this is a red flag. It suggests the industry is still in the “move fast and hope nothing breaks” phase when it should be in the “move carefully and know exactly what to do if something breaks” phase.
OpenAI’s new position on California’s safety bill will be interesting to watch. If they’re serious, they’ll support specific provisions that require documented safety plans, independent audits, and real accountability mechanisms. If they just want a bill they can point to while continuing business as usual, they’ll support vague principles and voluntary commitments.
I’m betting on the latter. Not because OpenAI is uniquely cynical, but because that’s what companies do. They optimize for what they can live with, not for what’s actually safest.
The AI safety conversation has been dominated by abstract debates about superintelligence and alignment for years. Those questions matter, but so does the boring stuff. Like having a written plan for what to do if your model starts behaving dangerously. Like being able to show that plan to someone who isn’t on your payroll.
Until the major labs can do that, all the safety rhetoric is just noise. You can’t claim to take a risk seriously while refusing to document how you’d respond to it. That’s not safety culture. That’s PR.
One email at dawn. The five stories that mattered, with the bits removed and the meaning kept. Free, for now.