Another AI Safety Kerfuffle Has Hit the Timeline
Happy Wednesday.
The current thing in tech and business is another AI safety debate, this time around reporting from the Information (which some are calling questionable) that OpenAI has used a breakthrough in “neuralese” for Astra that could reduce monitorability.
Today’s Lineup
- Palo Alto Networks CEO Nikesh Arora and Console founder Andrei Serban at 11:45 AM
- Blue Owl Capital Senior Managing Director Kurt Tenenbaum at 12:00 PM
- Starcloud CEO Philip Johnston at 12:10 PM
- Former record label executive and Beats co-founder Jimmy Iovine at 12:15 PM
Run of Show
Another AI Safety Kerfuffle Has Hit the Timeline
There’s another AI safety kerfuffle on the timeline this morning. Yesterday, Amir Efrati at The Information reported that OpenAI is “quietly using loop transformers that don’t show [the model’s] ‘thinking’ when scaled up,” but which make the model more efficient (this is also referred to as “neuralese”). Being able to monitor a model’s chain-of-thought (CoT) can provide both an early warning system for various forms of misalignment (deception, manipulation, going after an unintended goal) and insight into an AI’s actions after the fact, so the reporting quickly set off a broad wave of anxiety among the AI safety camp. Their chief concerns being: (1) not being able to monitor CoT makes AI less safe, and (2) by using neuralese, and thus abandoning CoT, OpenAI starts a race to the bottom among frontier labs, where others will also abandon CoT for the benefit of efficiency, thus making all frontier AI less safe.
OpenAI’s research director Jakub Pachocki responded to all this. “I want to prevent a race into unmonitorability kicked off by confused reporting,” he wrote. He continued:
OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it’s a core goal of our current research program.
In other words, Jakub is saying The Information’s reporting is inaccurate, but that the state of CoT monitoring is not that great, and is in fact getting worse (even though OpenAI is dedicated to making it better). Former OpenAI staffer Joshua Achiam made an additional point here: “Coordinating everyone around a technique this brittle is about as bad a safety or strategy posture I could imagine. It’s not just ‘it will break eventually,’ it’s ‘this is a fundamentally unsound basis for safety.’”
And Dean Ball chimed in:
We are now adjudicating technically complex and nuanced claims on the timeline with almost no ground-truth information about what is actually happening. Communities form these “thought-terminating taboos,” as roon says, and then panic at anything that vaguely resembles them. And the trust-eroding reality of social media makes the timeline an especially difficult place to do this adjudication.It is frankly insane and crazymaking and grating for everyone involved. It would be like if we argued about what every publicly traded company’s financials were by posting hyperventilating [sic] on the timeline rather than relying on the institution of auditing and the audited financial statements that institution produces.
To sum it up: if you’re in the AI safety camp, you’re worried that OpenAI is abandoning a technique for making AI safer, and that other frontier labs will follow suit. If you’re skeptical of the AI safety camp, The Information’s reporting is inaccurate, the AI safety camp is unnecessarily panicking, and OpenAI is doing the best it can. — Brandon
Clip Spotlight: CrowdStrike President Michael Sentonas says AI has pushed phishing click-through rates to over 60%, from roughly 11–12% just a year ago
Headlines
WSJ: Google Avoids Breakup of Dominant Ad Tech Business
Reuters: US government backs OpenAI in New York Times copyright case
Axios: Lutnick: Anthropic is “back on the right side” with Trump administration
Eric Seufert: Unpacking the FTC’s case against Amazon
Sam Altman posts about OpenAI’s safety priorities
Vanity Fair profiles Ed Zitron: “He Did Tech PR. Now He Rails Against AI for a Living.”
Mark Gurman: Apple Sets Pay Targets at $58 Million for Ternus, $47 Million for Cook
Ilya Sutskever calls on neolabs to strengthen their cybersecurity in ominous 𝕏 post
TechCrunch: Norway is considering a ban on smart glasses and other camera-enabled wearable headsets
WSJ: Elon Musk Wants to Make Power-Turbine Components. It Won’t Be Easy.
WSJ: The Island Paradise That Is a Secret Hub for Russian Sanctions Evasion
FT: Maersk turns to wind sails to cut fuel for container ships
WSJ: The Sudden Unraveling of Wall Street’s Momentum Trade
Medici, owner of David Protein, raises $250M
WSJ: Preppy Style Is Booming. What Does It Look Like Now?
Posts of the Day
Special thanks to our sponsors: Ramp, Shopify, CrowdStrike, MongoDB, NYSE, Codex, Public, Console, Railway, Figma, and Cisco.