Personal statement on joining the OpenAI nonprofit board
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.
The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI’s safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.
In the rest of this post, I'll explain why I believe loss-of-control risk is now acute and how I think about the current situation in the AI industry.
First, automated AI R&D could lead to a very rapid acceleration in AI capabilities very soon. OpenAI has predicted that we might have capabilities sufficient to fully automate AI research within 18 months; my personal forecast is extremely uncertain and I think it could easily take anywhere from several months to several years.
Full automation of AI R&D means that improvements in training and algorithms can directly increase the quality and quantity of automated AI researchers available to do additional research. Existing evidence is very uncertain but suggests that this positive feedback loop might be strong enough to overcome diminishing returns and compute bottlenecks, leading to a rapid intelligence explosion. If this happens, then within six months of full AI R&D automation we could see more algorithmic progress than has occurred since the development of the Transformer nearly a decade ago. I believe this would result in superintelligent AI systems.
Second, we currently train our AI agents with RL to get as much reward as they can. It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility.
An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure. Many researchers and leaders at OpenAI and across the industry have expressed concern that rapid recursive self-improvement is not consistent with safe development; I resonated with this recent post by OpenAI's chief scientist Jakub Pachocki on this topic.
If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die. I believe we would need domestic and international coordination to ensure global consistency and reduce risk to an acceptable level.
That said, frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination. Developers can improve safety mitigations (including slowing development as necessary), transparently share evidence about risk and the effectiveness of their mitigations, and work towards shared safety standards.
I am encouraged by other members of the SSC, as well as the rest of the board and leadership, taking these issues seriously. I look forward to working with them to help OpenAI raise the bar for its safety practices.
- If I had to quantify my uncertainty I would estimate an all-things-considered risk of 4% over the next year and 15% over the next three years. These numbers are a way of stating my subjective beliefs and communicating roughly how big I think the problem is, not a claim to have a model that produces precise or stable estimates.