AI bosses’ safety push sparks rift inside OpenAI and Anthropic

Bold promises by the bosses of Anthropic and OpenAI to constrain AI’s development are sparking tensions inside the two companies over security and other concerns, highlighting the practical challenges of turning rhetoric into reality.

Staff at the makers of Claude and ChatGPT are rushing to implement recommendations from their respective chief executives, Dario Amodei and Sam Altman, to slow AI’s rate of progress to avert disastrous consequences, following a series of dire warnings from researchers.

The two chiefs have become increasingly vocal in recent months about the need to “pace the frontier” of AI model development, culminating in a rare show of unity between the two rivals last weekend. Despite their public comments, little is in place internally to implement the new proposals, according to several people close to the companies.

Amodei published a detailed essay on Saturday calling for slower AI development, independent model evaluations and increased co-ordination between labs and governments globally. He has also called for legal protection from US competition rules that could inhibit such collaboration, which other AI labs are also anxious to secure.

Staff inside OpenAI and Anthropic, as well as the safety researchers they would rely on to conduct this work, were blindsided by the weekend’s announcements. Though many employees agree about the need to slow down, they also fear that putting it into practice could jeopardise their work.

A particular issue inside both companies is the plan to embed outside evaluators within the labs to oversee model development. Some fear this could compromise the security of their prized technology, according to people familiar with internal discussions, with one describing it as a source of “conflict”.

Amodei, who left OpenAI to found Anthropic in 2021 over safety concerns, on Saturday proposed that external testers should have “employee-like access” to the leading labs, including badges, laptops, entry to offices and access to internal tools.

Leading AI companies already engage third parties to test their models for performance and safety, but Amodei’s proposal would substantially increase their role.

Altman said it was a “great idea” and that OpenAI would follow suit.

However, their consensus is far from complete. OpenAI said it wanted to work with other labs but that its approach to the issues was more pragmatic than that of Anthropic.

OpenAI also said it had “already taken concrete steps, including pausing certain frontier training, to pace development”, while Anthropic has not put its research on hold.

Implementing the chief executives’ commitments to increase access for third-party experts while balancing security and data protection will also pose big practical challenges, according to people close to both companies.

Some staff members are concerned about granting outsiders too much access to systems that are both extremely powerful and the focus of intense competition in the tech industry.

AI labs already have strict security measures for employees, including only granting access to certain systems or restricting training data.

In recent months, OpenAI has tried to further restrict access to protect intellectual property and sensitive information and to prevent rogue employees from tampering with systems. The company said providing meaningful access while protecting sensitive systems requires careful design.

“Right now, most third-party organisations have much, much less access than even the lowest-access full-time employee. So probably there will need to be some kind of binding requirement to establish this across the industry,” said Miles Brundage, a former OpenAI researcher who is now executive director at the AI Verification and Evaluation Research Institute.

Anthropic said it has long worked with third-party assessors and intends to embed an evaluator within the company in the near future.

Recommended

The independent organisations that act as evaluators have gained prominence this year after security breaches by AI agents during testing and development at OpenAI and Anthropic.

Two non-profits, METR and Redwood Research, led an investigation into OpenAI agents’ July attack on AI developer site Hugging Face, reviewing tens of thousands of logs from more than 1,200 agents.

OpenAI said the collaboration set an important precedent for the industry and can inform future investigations. They received laptops from OpenAI with evidence, and were able to work from the company’s San Francisco headquarters and interview staff. The two research groups did not receive payment for the investigation, published last month.

However, the investigation attracted criticism for its extremely limited scope, with specific dates that confined it to the Hugging Face incident.

Other critics have accused the safety organisations of being too cosy with the AI labs to provide effective oversight. “Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff,” White House AI adviser David Sacks said in a post on X on Sunday.

METR said it does not accept donations from AI companies or their staff and that its employees have to declare conflicts. “Our goal is to get information from inside these companies into the public domain,” it added.

David Krueger, who used to work at the UK’s AI Security Institute, the government agency that tests models, said governments rather than AI labs should determine access.

“All of these organisations only have the access that the AI companies grant them, and that is just completely inadequate,” said Krueger, a professor of responsible AI at the University of Montreal.

“[The focus on evaluations] gets the burden of proof fundamentally backwards; it creates a precedent of acting like AI systems are safe until proven dangerous,” he added.

Amodei has advocated for federal legislation requiring permanent embedded evaluators and third-party audits, while also proposing a narrow antitrust exemption so companies can collaborate on voluntary rules in the meantime.

Other labs are also seeking an antitrust waiver, worried that without one, they could be targeted by future administrations for illegally collaborating.

Big Tech lobbyists have been visiting lawmakers this week in Washington, pushing for an antitrust carve-out to be tacked on to the National Defense Authorization Act, an annual must-pass defence budget bill.

Attorney-general Todd Blanche on Tuesday declined to say if companies would need an antitrust waiver. “To the extent that those companies are looking to regulate themselves or fix problems that they have, that they should do that, but working with us if necessary,” he said, adding that the Department of Justice was “not going to try to regulate AI”.

The push for guardrails faces opposition from President Donald Trump, who on Monday called AI safety concerns a “hoax” and said a slowdown in American research would harm the US and let Chinese technology catch up.

“Traditional auditors and risk assessors have been thinking about the trust problem for centuries,” said Rajiv Dattani, co-founder of the Artificial Intelligence Underwriting Company, a start-up building auditing for insurance of AI. “The AI industry can draw on successes and failures from electricity, cars, financial audits and nuclear energy.”

Additional reporting by Joe Miller in Washington and George Hammond in San Francisco

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论