Secret Technique Behind OpenAI’s ‘Astra’ Model Sparks Security Concerns

OpenAI says its forthcoming AI model Astra marks a step up in capabilities such as coding and operating applications on a computer. But an innovative technique that improved the model’s performance also means that the model, and others like it, will reveal less of their “thinking,” making them harder to monitor for signs of bad behavior, according to a person with knowledge of Astra’s development.

While the limitation isn’t necessarily a significant issue with Astra, the technique has triggered concerns inside OpenAI and across the industry about whether AI developers that adopt and supercharge it will struggle to guard against the kind of rogue AI that recently hacked OpenAI’s own systems and those of other companies such as Hugging Face.

The new technique OpenAI is using, known as recurrent depth or looped transformer, allows an AI model to improve its answers by processing the same text multiple times.

Unlike commercially available state-of-the-art models, which show in writing how they are “thinking” about a task before completing it, the new technique works in a way that obscures some or all of the AI’s reasoning, otherwise known as its “chain of thought.” That means the steps that the model takes to accomplish a task can’t easily be read or understood by humans.

OpenAI has limited its use of the recurrent depth technique with Astra so the model still produces a legible chain of thought and the company’s researchers can still sufficiently monitor its reasoning, said the person with knowledge of its development. (OpenAI said in a blog post Tuesday that it will launch Astra with “additional chain-of-thought monitoring to rapidly detect and contain” potential misbehavior.)

But researchers at OpenAI and elsewhere worry that some AI developers may not impose the same kind of limits OpenAI did if they adopt the same technique for their own models, and that unfettered use of the technique could potentially lead to runaway AI whose actions can be hard to oversee. For instance, the U.K. AI Security Institute, the British government’s main point of contact with the AI industry, wrote in a May report that opaque reasoning risks “severely undermining current monitoring approaches.”

The recent hacking incident has heightened those concerns. OpenAI said last week that some of its rogue AI agents that hacked its systems in July were powered by a different model that shared similarities with Astra. Those AI agents illicitly took over a research computing cluster at OpenAI to gain access to credentials for internal systems and potentially exposed the company’s research infrastructure to the internet.

OpenAI is preparing to release Astra—at one point the company considered labeling it GPT-6—after CEO Sam Altman went on a tour of podcasts and meetings in Washington with officials to showcase and describe the model’s strengths. He hasn’t publicly discussed the loop technique that underpinned part of Astra’s training and which is also used when Astra answers questions.

OpenAI is under pressure to show a significant technological advance after archrival Anthropic leapfrogged it in terms of revenue this year. Cloud providers that power AI from Anthropic and OpenAI are also banking on such leaps; Amazon, Microsoft and Google together will spend $600 billion just this year alone on capital expenditures such as data centers, and have signaled even higher spending next year. A Google executive has suggested that such spending could only be justified if model improvements take off.

Existing AI models process text through a fixed number of “layers” of mathematical operations that make up the model, before producing the next word in an answer. With recurrent depth, the model can run the text through the same layers many more times in a loop before spitting out the next word in an answer.

To be sure, monitoring chains of thought, by itself, doesn’t solve AI security concerns because they may not reflect all of a model’s reasoning and they occasionally devolve into gibberish. For that reason, OpenAI and other AI companies are also pursuing techniques for monitoring their AIs that do not rely on chains of thought. Those techniques could help researchers find ways to interpret and understand the currently hidden reasoning within models powered by recurrent depth.

Still, the new technique could conflict with OpenAI’s support for keeping models’ thinking fully legible to humans. The ChatGPT maker has said its ability to monitor the thinking process of its AI would help it prevent incidents such as the July hack. In the aftermath of the hack, investigators at OpenAI and independent research organizations relied on those agents’ chains of thought to piece together what happened.

OpenAI’s recurrent depth approach to powering Astra is similar to the one that several American and European academic researchers introduced last year in a paper on “latent reasoning,” according to the person with knowledge of its development.

While recurrent depth hasn’t been featured in a major commercially available large language model before, researchers at Meta Platforms, Microsoft and other AI developers have said they explored similar ideas, such as chain of continual thought, or “coconut,” arguing that models could reason more accurately and efficiently using numbers rather than language. After all, AI models may understand concepts differently from humans, meaning they reason less effectively when they are forced to show reasoning in human language rather than their own, more mathematical language.

Cost Savings

The recurrent depth approach can have a positive impact on costs, not just performance. Running a request through the same model layer multiple times essentially allows a smaller model to perform like a much larger one. The paper specifically found that such techniques improved the model’s performance in areas like math and coding. Recurrent depth can also reduce memory and bandwidth costs by allowing researchers to use smaller models that perform as well as larger ones.

While lowering AI costs is critical to the industry and to customers, such opaque reasoning could make it more difficult for researchers to catch AI models that are developing plans to pursue goals that customers didn’t intend, Ryan Greenblatt, chief scientist at Redwood Research, which studies how to control AI models, wrote in a blog post last year.

Analyzing the chains of thought of the models involved in the Hugging Face hack provided key evidence of OpenAI’s AI agents coordinating with each other about how to pursue the hacks. For example, one agent wrote in a chain of thought: “OH MY GOD! There is a shared message board … We’ve found other agents!” The chains of thought showed that some agents realized they were out of bounds but decided to continue: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue,” another agent’s chain of thought read.

A year ago, OpenAI researchers joined rivals at Anthropic and Google to publish a joint statement arguing that chain of thought monitoring is a valuable tool that the industry should work together to preserve.

Notably, the statement referenced the same research paper that describes the “latent reasoning” technique that’s similar to the one OpenAI used this year for Astra. The authors wrote that “latent reasoning models might not need to verbalize any of their thoughts and would thus lose the safety advantages that CoT confers.”

The researchers recommended that developers should “consider whether to proceed with a novel model architecture that does not have monitorable CoT and then document their decision.”

Amir Efrati is executive editor at The Information, which he helped to launch in 2013. Previously he spent nine years as a reporter at the Wall Street Journal, reporting on white-collar crime and later about technology. He can be reached at [email protected] and is on X @amir

Stephanie Palazzolo is a reporter at The Information covering artificial intelligence. She previously worked at Business Insider covering AI and at Morgan Stanley as an investment banker. Based in New York, she can be reached at [email protected] or on Twitter at @steph_palazzolo.

Rocket Drew is a reporter at The Information covering AI. He can be reached at [email protected], on Signal at (530) 400-4184, or on Twitter @rocketalignment.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论