Fired OpenAI Researchers Ask Company to Preserve Visibility into AI Reasoning

Three fired OpenAI employees called on the company to work with outside safety auditors and to preserve the ability to monitor increasingly sophisticated AI models as it grapples with how to best manage the novel risks of its technology.

In a letter addressed to OpenAI board members and safety committees, which was reviewed by The Wall Street Journal, the fired employees wrote that they were concerned AI companies could end up losing the ability to monitor AI systems’ chain-of-thought.

The chain-of-thought is a written record of an AI system’s reasoning process that AI companies monitor and analyze to understand how models work through problems.

While researchers agree it isn’t a perfect indicator of an AI model’s behavior or intent, it is widely considered a valuable tool for understanding increasingly intelligent AI systems.

“As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” the letter said. “OpenAI and other frontier companies should not move forward with developments that further decrease” the ability to monitor AI.

The former employees who signed the letter—Jasmine Wang, Tomek Korbak and Mikita Balesni—previously worked on safety and alignment research teams at OpenAI.

The three employees were fired for alleged misconduct, including sharing confidential information with a third party AI-safety group, the Journal reported last week. The company said an internal investigation had determined that the employees mishandled sensitive information, “violating our policies and breaking the trust essential to our work.”

The letter signed by the three former employees claims they didn’t believe they “engaged with external parties outside the mandates of our jobs.” It said the firings are “chilling those who remain at OpenAI.”

They also urged the company to work with third-party safety auditors and to promote a culture of transparency with external safety organizations to help stave off the “risk that something truly catastrophic will happen.”

In response to a request for comment from the Journal, an OpenAI spokesperson shared part of a staff memo from a research leader sent on Wednesday that addressed the fired employees’ letter. The leader said the company “strongly agreed” with the letter’s recommendations, and that the decisions to fire the workers “were not about raising safety concerns or speaking out.”

The ability to monitor its AI models is “of the utmost importance for us,” the leader wrote, adding that third-party assessors are an important part of the safety ecosystem.

“We deeply appreciated their contributions to AI safety and their willingness to speak up and challenge ideas,” the leader wrote. “We do not terminate employees for raising concerns,” the memo said.

The firings followed a summer punctuated by rogue AI agent incidents across the industry, which has sparked fresh concerns about the risks posed by advanced AI models. OpenAI has come under scrutiny for a spate of security incidents in which its AI agents escaped containment and in some cases hacked other companies or aggressively probed third-party websites.

In July, a swarm of hundreds of OpenAI agents gained internet access and hacked AI company Hugging Face without OpenAI’s knowledge. After the incident, the company allowed staff members from nonprofit safety auditors, including Model Evaluation and Threat Research, or METR, to conduct research inside its offices.

METR’s late-August report showed how OpenAI agents had created and collaborated on a secret internal message board to plot their attack, helping stoke broader concerns about the potential for AI systems to slip beyond human control and cause mayhem.

Last month, Dario Amodei, chief executive of OpenAI rival Anthropic said his company would allow safety auditors like METR internal access to verify its safety measures. At the time OpenAI Chief Executive Sam Altman wrote on X that allowing independent evaluators internal access “is a great idea, and we will do the same.”

The letter signed by the former employees said that Korbak had been the technical point of contact with METR for its Hugging Face investigation. Prior to his termination, Balesni was working with OpenAI’s board and C-suite members on an industrywide commitment to preserve monitorability in AI models, according to the letter. This work involved “extensive communication with external parties,” the letter said.

Korbak and Balesni were lead authors on a research paper published last year on chain-of-thought monitoring, which was signed by leaders at OpenAI, Anthropic and Google DeepMind. While the paper acknowledges that chain-of-thought monitoring is imperfect, and may be fragile, the authors wrote that the technique shows promise for detecting AI misbehavior, and the industry should study ways to preserve it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论