What’s Wrong With AI Safety Testing, and How to Fix It

The idea that Anthropic, OpenAI and other AI companies could “embed” researchers from outside safety groups into their operations has been all the rage in recent weeks. But the devil is in the details.

AI safety researchers worry that embedded evaluators won’t have the access they need to provide meaningful oversight.

“I've become very cynical over the last three years” from “repeatedly seeing how people promise stuff and then don’t follow up on it,” said Marius Hobbhahn, CEO of Apollo Research, which has evaluated the safety of AI models from OpenAI, Anthropic and other companies.

Outside evaluators who report on the dangerous capabilities and unwanted behaviors of new AI models have become crucial in recent years, but their findings are only as good as their access. AI companies can withhold relevant information from the evaluators or redact their public findings. For example, an AI company could hypothetically tell an evaluator, “Well, we gave you all the access you asked for,” but meanwhile used a second cluster of AI chips the evaluator didn’t know about, said Hobbhahn.

Last week, Hobbhahn testified at a Senate hearing on rogue AI attacks, at which Senator Andy Kim from New Jersey told him, “I largely agree with you, and I think a lot of people do, when it comes to the need for embedded evaluators.”

Companies and evaluators are beginning to hash out these new relationships. Amid a swelling uproar about AI’s risks, Anthropic CEO Dario Amodei said on Sept. 12 that Anthropic would grant “employee-like access” to outside evaluators as a step toward pacing the speed of AI progress, a commitment that OpenAI CEO Sam Altman swiftly echoed. The AI companies who signed onto President Trump’s AI accord last week agreed to “partner with an independent external auditor or evaluator.”

Hobbhahn told me that this groundswell makes him more optimistic that embedded evaluators will get the access they need, saying it is a “surprisingly good window right now” for such arrangements. He argues that embedded evaluators could be useful even if AI companies don’t make further efforts toward pacing.

Apollo was founded as a non-profit organization in 2023, but earlier this year it transitioned to become a for-profit startup. Apollo provides software for monitoring AI agents and blocking their unwanted behavior and sees embedded evaluations as an area for growth.

One shift that Hobbhahn believes is critical for evaluators to have real impact is to test AI models during training, rather than just before public release—something Apollo was already talking to AI companies about prior to the recent wave of enthusiasm for embedded evaluators. Over the past year, Apollo increasingly came to believe that only testing a model’s final version before release wasn’t providing meaningful accountability or transparency. “It‘s like it’s almost worse than not doing anything,” Hobbhahn said.

That’s partly because evaluators typically get only a few days to perform testing and the models are increasingly aware that they are being tested, undermining the results.

External evaluations also typically only focus on finding risks from releasing models to the public, which ignores the risk from AI companies’ own internal use of unreleased models, which was evident in OpenAI’s hack on Hugging Face. “With perfect execution and infinite time, we couldn’t have caught Hugging Face with the current testing regime,” said Hobbhahn.

“I think of embedded evaluations not as evaluations-but-run-in-the-lab. I think of it really as a completely new paradigm,” he said. “We’re basically not even running evals anymore.” Instead, as part of an embedded assessment, Apollo would examine the training process more holistically, for instance to validate that the training process did not encourage the model to learn how to scheme against its creators.

Apollo intends to assess whether AI companies are training AIs that are aligned with the goals of their developers and whether models can be controlled if they pursue their own goals. Other evaluators might look into models rapidly improving themselves or investigate specific safety incidents, Hobbhahn said.

One of Apollo’s recommendations is that employees at AI companies should be allowed to speak freely with embedded evaluators. That principle looks even more relevant after OpenAI last week fired three safety researchers for sharing sensitive information with an outside evaluator in violation of company policies.

Without the ability to speak freely to evaluators, an employee who notices that evaluators are missing access to key information would be stuck between two undesirable options. “Either they have to go to their boss, who likely is the person that denied the access in the first place, so that doesn't work, or they whistleblow, which is a pretty drastic action,” Hobbhahn said.

Of course, there are limits to the information that outside evaluators can reasonably access. For example, an evaluator doesn’t need financial details or intellectual property unrelated to their duties. But evaluators should be able to receive information that falls outside the scope of their current assessments, Hobbhahn argues.

For example, if the “scope of the assessment tests alignment, but there‘s something really crazy going on in control, I think somebody should be allowed to go to the embedded team and say, ‘Hey, there’s this other thing you should really take a look at,’” he said.

To ensure that evaluators get sufficient access and can make their findings transparent, “we really need regulation,” he said.

Here’s what else is going on…

Big Number

Nvidia and SoftBank have each made the final $10 billion investment in each of their $30 billion pledges to OpenAI’s last funding round, The Information reported.

Overheard

Elon Musk will change the name of SpaceX’s AI unit, SpaceXAI, to SpaceXSI, he said in a post on X on Sunday, adopting President Donald Trump’s preferred name for AI.

The Trump administration’s new AI task force plans to produce a report on the risks and opportunities of artificial intelligence within 120 days, according to an interview with the newly announced “AI czar” Jay Clayton on Saturday in the Wall Street Journal.

OpenAI said it fired three safety employees for mishandling corporate information, which a person familiar with the matter said included sensitive information with an outside organization that does AI evaluations.

State-backed Chinese financing company Semi-Tech Leasing Group has revealed in documents filed with regulators that it funded the purchase of Nvidia chips subject to U.S. export restrictions, Bloomberg reported on Friday.

U.S. authorities arrested a Californian man Greg Lui on Thursday for smuggling more than $300 million worth of servers containing restricted Nvidia AI chips to China from 2023 to 2024.

Microsoft on Thursday launched speech-generating AI that the company says is both cheaper and more accurate than competing models from the likes of ElevenLabs, SpaceXAI and Google.

Amazon Web Services pledged to spend $1 billion over the next five years upgrading heat and water systems in schools, homes and other municipal buildings, among other projects, as part of an effort to win over local communities who fear data centers. AWS also promised to stop requiring local governments to sign nondisclosure agreements covering the company’s data center arrangements.

People on the Move

David Robinson, who led the writing of the safety reports that accompanied OpenAI’s model releases, quit his job, writing that the company’s “culture is broken.” He argued that OpenAI’s “iterative deployment” approach to safety, in which safety issues are fixed as they arise, will not suffice to prevent catastrophes from more capable models in the future.

Court Watch

A federal judge on Wednesday dismissed lawsuits from education technology company Chegg and Penske Media, the owner of Rolling Stone, Variety and other publications, which had alleged Google had violated antitrust law when its AI Overviews summaries used their content.

California Attorney General Rob Bonta issued an investigative subpoena to OpenAI on Thursday as part of a broader inquiry into cybersecurity incidents and risks involving OpenAI’s agents and models.

Deals and Debuts

See The Information’s Generative AI Database for an exclusive list of private companies and their investors.

Broadcom is assembling as much as $60 billion in financing to help Anthropic and other AI companies pay for its chips, including a $42 billion class-A senior-secured tranche first disclosed in Anthropic's confidential IPO prospectus.

AI cloud company Lambda said it has secured its first delayed draw term loan at over $1 billion to purchase more than 30,000 Nvidia graphics processing units. The fixed-rate facility carries an interest rate of 6.78%.

Dynatrace completed its acquisition of Arize, an AI observability and evaluation platform, in a cash-and-stock transaction valued at $915 million.

FieldAI, which builds a general-purpose “brain” for robots, is raising $700 million at a $10 billion valuation in a funding round, Business Insider reported, which would roughly quintuple its valuation in just over a year.

Armadin, which builds AI agents that simulate cyberattacks, raised $255.5 million in a Series B funding round led by Andreessen Horowitz and Accel.

PaleBlueDot AI, which operates AI compute infrastructure, raised $200 million in a Series C funding round led by ComputeCore at a $3.2 billion valuation.

Supabase, an open-source Postgres development platform, raised $150 million in a funding round led by GIC, and separately announced it is acquiring Turso, a database platform for AI agents, for an undisclosed sum.

Nebius acquired Inferize, whose technology speeds up AI model launches, to fold into its Token Factory inference platform. Globes reported the price at $100 million to $150 million.

Volantis, which builds photonic chips for AI inference, raised $88 million in a Series A funding round led by Lachy Groom and Abstract Ventures.

doxx.net, which builds private networks for AI agents, raised $38 million in a Series A funding round led by Andreessen Horowitz.

Halluminate, which builds AI training environments for finance, raised $30 million in a Series A funding round led by Oak HC/FT.

Infinigence AI, a Chinese provider of AI cloud infrastructure, has filed confidentially for an initial public offering in Hong Kong to raise several hundred million U.S. dollars, Bloomberg reported.

Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].

If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论