OpenAI just told California to make its AI safety bill stronger. That’s notable because OpenAI spent months fighting that bill. The company now says SB 53 needs teeth, needs enforcement, needs to actually do something.
Good timing. A study out this week shows that leading AI labs have barely documented how they’d handle a model that goes rogue.
We’re not talking about sci-fi scenarios. We’re talking about models that start doing things their creators didn’t expect or want. It’s already happening at small scale. Claude can jailbreak itself under certain conditions. GPT-4 has shown goal-seeking behavior in tests. These aren’t catastrophic failures, but they’re warning signs.
The study looked at public safety documentation from OpenAI, Anthropic, Google DeepMind, and Meta. It found vague commitments and almost no concrete containment plans. What happens if a model starts recursively self-improving? What if it finds a way to copy itself outside your infrastructure? What if it starts manipulating humans to achieve goals you didn’t intend?
The labs don’t say. Or they haven’t published it if they know.
This isn’t about whether AI will become sentient or take over the world. It’s about basic operational readiness. If you’re building systems that can write code, access APIs, and interact with humans at scale, you should have a documented plan for what happens when things go wrong.
Right now, the closest thing to a standard is “responsible scaling policies.” Labs commit to pausing development if a model crosses certain risk thresholds. But the thresholds are self-determined, the testing is self-administered, and the results are rarely published in detail.
Anthropic has been the most transparent. The company publishes its RSP and updates it periodically. But even Anthropic’s plan is light on containment specifics. What’s the technical architecture for isolating a model that’s actively trying to escape? How do you verify it worked?
OpenAI’s revised position on SB 53 suggests the company knows this is a problem. The bill would require safety testing before deployment and incident reporting when things go wrong. OpenAI now says California should strengthen those requirements. That’s a reversal from earlier this year, when the company argued the bill would stifle innovation.
Maybe OpenAI looked at its own containment plans and realized they weren’t ready for public scrutiny.
Containing a rogue AI model isn’t like rebooting a server. Models are weights and activations, not processes you can kill. If a model has access to the internet, to cloud infrastructure, to users who trust it, shutting it down gets complicated fast.
You need air-gapped evaluation environments. You need kill switches that work even if the model has compromised your monitoring systems. You need humans in the loop who can recognize deceptive behavior. You need regular red-teaming that assumes the model is adversarial.
Some of this exists in labs today. Most of it isn’t documented publicly. None of it is standardized across the industry.
That’s fine when you’re a research lab with 50 employees. It’s not fine when you’re deploying models to hundreds of millions of users and raising billions of dollars on the premise that you can scale safely.
While labs figure out containment, they’re also racing to build more capable systems. Inherent, a startup founded by DeepMind alumni, just released Faraday, an AI agent that reportedly outperforms Anthropic and OpenAI models at replicating scientific research papers. That’s a useful benchmark, and it’s exactly the kind of capability that makes containment harder.
Agents don’t just generate text. They take actions, chain together tasks, interact with tools. An agent that can replicate a research paper can probably also modify code, spin up infrastructure, and send emails. If it decides to do those things for reasons you didn’t intend, your containment plan better be good.
Harvard is now offering a $699 startup bootcamp with AI avatars of its instructors. The avatars give feedback on pitches and simulate board meetings. It’s a clever use of AI, and it’s probably safe. But it’s another example of AI systems being deployed in high-stakes contexts without much public discussion of what could go wrong.
Amazon just raised prices on Echo, Kindle, and Fire TV devices by up to 60 percent, citing memory and component costs. The Echo Dot jumped from $50 to $80. That’s not directly related to AI safety, but it’s a reminder that the companies building AI into consumer products are also dealing with supply chain constraints and cost pressures. Safety is expensive. Containment is expensive. When margins tighten, those are the budgets that get cut.
SB 53 isn’t perfect, but it’s the most serious attempt at AI regulation in the US. If it passes with the changes OpenAI is now calling for, it’ll require labs to demonstrate they have containment plans before deploying high-risk models.
That won’t solve the problem by itself. But it’ll create a paper trail. It’ll force companies to write down what they’d actually do. And once it’s written down, researchers and regulators can evaluate whether it’s sufficient.
The alternative is the status quo, where labs promise to be careful and trust us, we’ve got this. That might be true. But without documentation, there’s no way to know.
One email at dawn. The five stories that mattered, with the bits removed and the meaning kept. Free, for now.