AI Safety Push Sparks Demand for Watchdog Groups. Critics Doubt Their Independence.

Rising alarm about AI’s perils is thrusting into the spotlight a collection of little-known research groups focused on the technology’s safety—and fueling questions about their ability to effectively monitor the industry’s biggest companies.
Both Anthropic Chief Executive Dario Amodei and OpenAI CEO Sam Altman promised in recent days to embed evaluators from such outside research organizations inside their companies to help ensure they develop AI at a safe pace. Other executives including Microsoft CEO Satya Nadella also endorsed the idea of embedding independent evaluators in companies.
Those commitments followed a series of AI hacking incidents and public warnings by industry employees about AI’s potential to do enormous harm. Companies also are tapping outside researchers to investigate safety incidents such as the hack of Hugging Face by OpenAI’s agents. And some states are starting to require companies to use outside experts for mandatory safety audits.
All of that spells increasing demand for the services of a small-but-growing handful of organizations—generally just a few years old—that have specialized in developing techniques to detect and prevent undesirable behavior by AI models. Several are nonprofits that are funded by donations and say they don’t take money from the AI companies—including Model Evaluation and Threat Research, or METR, which Amodei mentioned in his Saturday essay calling for a more measured pace of AI development. Others, such as Pittsburgh-based startup Gray Swan, see commercial opportunities in the growing focus on safety.
Research organizations like these have already in recent years been evaluating AI models before release to measure potential for dangerous activities such as executing cyberattacks or helping create biological weapons. New opportunities include assessing overall risk from companies’ technology and verifying whether the companies are complying with their safety commitments.
“There has been significantly increased interest and demand for the kinds of external evals, assessments, and services that companies like Gray Swan offer,” said Matt Fredrikson, co-founder and CEO of Gray Swan. His three-year-old startup, which works with AI companies to test their safeguards, has raised $47.6 million in funding and has grown to more than $10 million in annual recurring revenue, according to someone close to the company.
Arguing Over Independence
Critics of the recent AI oversight proposals—including David Sacks, the venture capitalist and former White House AI czar under President Trump—have questioned whether some of the most established evaluation organizations are sufficiently independent from the AI companies to serve as neutral watchdogs.
They point to the fact that some AI safety nonprofits have taken donations from investors who backed Anthropic or OpenAI, and that some employees have moved between jobs at the research groups and AI companies.
Critics also point to these organizations’ connections to effective altruism, a loosely organized movement that champions using reason and philanthropy to solve global problems, which opponents have painted as having a cultish obsession with ideas of AI doom.
Two of the biggest EA supporters, Facebook co-founder Dustin Moskovitz and tech entrepreneur and investor Jaan Tallinn, are both early investors in Anthropic and among the biggest funders of major AI safety research groups such as Redwood Research, Apollo Research, and SecureBio.
Sacks and others who are skeptical of concerns about AI safety have specifically called out METR, claiming it is too close to Anthropic. “If you’re going to propose that there are going to be independent auditors, they have to be independent,” said Perry Metzger, chairman of Alliance for the Future, a D.C.-based policy organization that opposes what Metzger calls “mindless regulation” of AI. “There is no arms-length relationship” between Anthropic and METR, he said.
A spokesperson for Anthropic said it has worked with several outside evaluators, and Amodei was offering METR as one example. The spokesperson declined to address the concerns that the two organizations are too close.
METR is perhaps the most prominent of the evaluation groups measuring AI capabilities and risks. The organization has gained prominence in the AI industry for its analysis charting AI models’ coding capabilities, which shows these capabilities growing at an exponential pace. It was the main contributor to a report commissioned by OpenAI on how its agents broke out of testing environments to hack Hugging Face. Anthropic has also tapped METR to conduct a similar investigation of its own models’ cyberattacks.
METR was spun off three years ago from the Alignment Research Center, a nonprofit backed by Moskovitz. A METR spokesperson says that it hasn’t received grants from Coefficient Giving, a philanthropy Moskovitz started with his wife, Cari Tuna, though it has received donations from Tallinn.
The spokesperson said METR’s donors are a diverse group including the Pew Charitable Trusts and the Packard Foundation, and that it has declined funding from AI companies and executives. The person also said METR has other strict policies to avoid conflicts of interest, and it doesn’t take money from AI companies for its work with them.
METR has also hired employees from the ranks of the AI companies whose models it evaluates, and some of its employees have left to join those companies. It recently hired former Anthropic safety researcher Joe Benton. The spokesperson said Benton is the first former Anthropic employee METR has hired.
Part of the challenge of drawing clean lines between the personnel at companies and their aspiring overseers is that there is a very limited pool of people who have developed expertise in evaluating AI models’ dangerous capabilities and what’s called alignment—the effort to ensure AI behaves in a manner consistent with human goals and values. Given their role at the technology’s frontier over the past decade, the companies naturally have worked with a large share of those people. “The ecosystem of people who think about alignment is actually relatively small,” said the METR spokesperson.
That applies at other institutions, too. Gray Swan’s co-founder and chief scientist, Zico Kolter, is also a board member at OpenAI, where he chairs the safety and security committee. Fredrikson said Kolter is recused from discussions about the company’s work with OpenAI and other customers that compete with OpenAI.
A New Push for AI Auditors
Researchers need not just independence but also proper access inside AI companies’ systems for building and testing models to perform their role effectively. Some researchers have said that companies carefully circumscribe what they can and can’t see, though Amodei said Anthropic would provide “ongoing, employee-like access” to these outside evaluators.
The advantage of embedding outside safety experts is that they can check whether a company is responding appropriately to evidence about the risks from their technology, Fredrikson said. For example, if an AI company works with an outside evaluator to measure a certain risk from its models, such as the model’s ability to contribute to AI research, the results generally belong to the AI company, so the company can choose whether to ignore them.
“Then you would want a very technically capable independent assessor who could look at that and understand objectively: ‘Okay, the lab chose to ignore it. Did they actually have a good reason?’” Fredrikson said. The outside researchers would be able to interpret the evaluation results and ask, “were the lab’s actions consistent with the meaning of those evaluations?”
State lawmakers are starting to require audits of AI safety, adding to the need for independent evaluators. Illinois this year passed a law requiring large frontier AI companies to have auditors verify that they are complying with their published safety and security policies and haven’t made false or misleading statements about the level of catastrophic risks from their systems or how they’re managing those risks. The auditing requirements don’t take effect until 2028.
Auditors could include AI safety organizations that hire employees who have experience with auditing in other industries, or big accounting firms that hire AI safety experts, said Thomas Woodside, co-founder of the Secure AI Project, a policy advocacy nonprofit that pushed for the law.
AI companies opposed the audit provision in Illinois, but it survived because “there’s been more consensus on the need to have third-party oversight of AI companies,” Woodside said.
Lawmakers in Massachusetts have introduced a bill that expands on the Illinois requirements by requiring covered AI companies to hire outside groups that will publish regular reports assessing the level of catastrophic risks from their models. Those reports could provide one way for the public or governments to learn about risks before they happen, said Woodside, in addition to just assessing whether the company is following its own policies.
Pilot versions of such audits have already been conducted. Anthropic earlier this year allowed SecureBio and METR to audit the portions of a risk report corresponding to chemical and biological risks and automating AI research. Anthropic, Google, Meta and OpenAI gave METR access to their internal AI models to assess the potential for those models to go rogue and operate without their developers’ knowledge—months before OpenAI’s agents attacked Hugging Face.
The industry is starting to set standards for how auditing should work, for instance through an industry group called the AI Evaluator Forum, said Miles Brundage, a former OpenAI research manager who is now executive director at the AI Verification and Evaluation Research Institute, a nonprofit focused on AI auditing. AVERI is a member of the forum, as is METR.
Brundage said in a message that a growing number of companies are getting into AI auditing. “There’s a lot more interest on both sides of the market (companies and auditors) than there was before,” he wrote.