The OpenAI Shock
What to make of this week’s electrifying OpenAI story, in which two models autonomously hacked their way out of a supposedly segregated “sandbox”, broke out onto the internet, strategized, stole a password, discovered more chinks in a supposedly walled-off AI platform, and raided its code?
Four points stand out.
1. Frontier labs, however well intentioned, cannot be trusted to control their own technology. OpenAI deserves credit for being transparent after the fact. But clearly it would have been better if a third-party had stress-tested its models’ potential misbehavior before things went wrong. AI risks are no longer theoretical. The time for theorizing about the potential need for oversight is past.
2. The nature of this oversight can be debated. Ideas range from a traditional government regulatory agency (modeled on the Food and Drug Administration) to a loose cabal of AI tycoons keeping an eye on one another (Elon Musk proposes this in his new interview with the editor-in-chief of The Economist, aka my wife). For reasons explained in last week’s Machine Readable, there is much to be said for the middle ground: a self-regulatory agency, as proposed by Demis Hassabis.
3. I haven’t seen other commentators say this, but the OpenAI exploit underscores the case for cracking down on open-weight AI systems. The two escapee models were being tested without guardrails; their alarming shenanigans demonstrate that guardrails are needed. Open-weight AI systems are freely modifiable by users, which is to say that bad guys can remove their guardrails. In the wake of this week’s sorcerer’s apprentice spectacle, surely people see that allowing removable guardrails is nuts?
4. Relatedly, it’s worth checking out Xi Jinping’s speech from last week, in which he recommits China to open-weight models. (Link here to the text, annotated by @MattSheehan.) Before Xi’s declaration, there had been speculation that, given the power of new Chinese systems such as Kimi K3, China might shift its policy in favor of restrictions. But there’s no sign of that so far. Perhaps in hopes of shifting Beijing’s thinking, Treasury Secretary Scott Bessent is said to be planning AI talks with Chinese counterparts, to be held in September. The OpenAI scare makes those talks more urgent. It also boosts the chances that China will get the point.
On the subject of open weights and the need to speak with China, check out the podcast I recorded with Daniel Kurtz-Phelan, the editor of Foreign Affairs. A friend remarked that in this episode I discard my polite English veneer and unleash my inner Australian. Of course I understand that people love open-weight models: they are cheap and customizable. But aviators once loved zipping through the heavens untrammeled by air-traffic control. We have to trade convenience for safety.
In other news, I published a New York Times guest essay pushing back against Sam Altman’s proposal that the government take 5% of his company. Even if the government took 5% of the top dozen tech companies engaged in AI, the dividends from this pot would be insufficient to achieve the declared goal: to protect American citizens from an AI jobs apocalypse. AI-driven profits will accrue to users of AI as well as the producers, so the government should claim a share of this bonanza through a broad company tax, not through a narrow equity appropriation. Unfortunately, the Trump administration is fond of collecting corporate equity stakes, so Altman’s idea stands a fair chance of adoption.
These are wild times. Buckle up.

Thanks a lot for the article as always a pleasure to read. I'm curious what are the main supporting facts/hopes beyond the idea that "AI-driven profits will accrue to users of AI"?