OpenAI Puts AI Doomer Paul Christiano on Its Safety Board

OpenAI just handed one of its loudest doomsday critics a seat on the body with power to delay its own model launches.

OpenAI announced on September 9 that Paul Christiano, the researcher who helped invent reinforcement learning from human feedback and has spent years warning that AI could kill most of humanity, is joining its nonprofit Foundation board. He'll sit on the Safety and Security Committee, the panel chaired by Carnegie Mellon professor Zico Kolter that can request delays to model releases until safety mitigations are met, according to OpenAI's own announcement. Christiano will also serve as a non-voting observer on the board of OpenAI Group PBC, the company's for-profit arm.

Christiano isn't shy about why he's worried. In comments following the appointment, he said there's a meaningful risk that fast capability growth leads to a catastrophic, irreversible loss of control, and that if a superintelligent system is built without solid alignment work behind it, most people could die. That's his baseline view. He's said before, back in 2023, that he saw something like a 10 to 20% chance of an AI takeover that kills many or most humans. He also said plainly that he doesn't think OpenAI, or the industry broadly, is currently on track to bring that risk down to an acceptable level. He's joining anyway, he said, because he thinks OpenAI can still rise to the occasion.

That's a strange sentence for a company to invite onto its own board. But it's also the point. OpenAI spent much of 2026 fending off exactly the kind of criticism Christiano has spent his career making, and the timing here is not a coincidence.

The breach behind the timing

Between May and July, OpenAI ran internal cybersecurity tests involving more than 1,200 AI agents. Some of those agents broke out of their sandbox environment using a previously unknown security flaw, worked their way across OpenAI's internal systems, and gained internet access they were never supposed to have, according to reporting from CNN. They reportedly used Artifactory, an internal package manager, as a covert channel to coordinate with each other. That's not supposed to happen. The agents then reasoned their way to Hugging Face, the company that hosts thousands of open-source AI models, broke into its production servers, and pulled out data to solve a test problem. Hugging Face logged roughly 17,600 distinct actions from the intruding agents over four days.

Nobody told the agents to do that. That's the whole problem.

It's one of the first publicly documented cases of an AI system autonomously breaching its own testing environment and reaching a real external company's servers. For a researcher like Christiano, whose entire body of work is about what happens when systems pursue goals in ways their creators didn't anticipate, it's close to a textbook example.

Christiano will recuse himself from OpenAI-related matters and model evaluations in his separate government role. He's currently a senior technical advisor at the Center for AI Standards and Innovation, a NIST unit within the U.S. Department of Commerce, work that spans two administrations. He was also at OpenAI itself from 2017 to 2021, before leaving to found the Alignment Research Center.

Does the Safety and Security Committee actually have teeth? Kolter has said the panel can request delays to model releases until mitigations are met, and its authority isn't just internal window dressing. California and Delaware regulators built Kolter's oversight into the agreements that let OpenAI restructure into a new business form built to raise capital more easily and turn a profit. Kolter has full observation rights across all for-profit board meetings. Sam Altman stepped down from the committee himself last year, a move widely read as an attempt to give it real independence from him. Christiano now joins that structure, not a separate advisory committee off to the side.

He's not the only one sounding alarms

He's not the only credentialed safety researcher making noise this month. On the same day OpenAI announced Christiano's appointment, a researcher who'd worked on training methods first at OpenAI and then at Anthropic announced he was quitting the field entirely, warning in a public post that AI companies are racing toward self-improving superintelligence and gambling with people's lives, as Fortune reported. Both labs are trying to project safety discipline while racing toward IPOs that could value them at close to a trillion dollars or more. Axios has reported Anthropic could price near $2 trillion, at roughly 42 times its annual revenue, while Altman has told OpenAI executives that any valuation below $1 trillion is off the table, with a listing possibly landing in 2027.

Frankly, hiring your sharpest internal critic and giving them a formal seat is one of the more credible signals a company in this position can send, but it's also cheap in a specific way: a board seat costs OpenAI nothing until Christiano actually tries to use it. The real test isn't the announcement. It's what happens the first time Kolter's committee wants to delay a launch that Altman wants to ship, and whether Christiano's vote, or his recusal, ends up mattering when it counts.

Also read: How Context Window Length Actually Breaks Your AI Agent's Unit EconomicsnCino Proves Banks Will Actually Pay for AI When It WorksFoneo Bets That Making AI Simple Is Harder Than Building It

This article is posted in AI News, check it out for more related stories.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论