AI Evaluators Raise Ideological Bias Concern for Trump, Industry

Groups that perform the niche work of assessing artificial intelligence models have suddenly found themselves at the center of a global debate over safety that could shape the pace of AI innovation and the direction of financial markets.

Evaluators such as METR are drawing increased interest from Anthropic PBC, Meta Platforms Inc. and other AI giants as the preferred way to guard against harmful models, with Microsoft Corp. emerging as the latest to offer its endorsement for the idea.

“We should have independent evaluators,” Microsoft President Brad Smith said Wednesday in a Bloomberg Television interview. “Ultimately the world has to have confidence in these evaluators. They better be chosen properly, and they need to have genuine independence.”

But as momentum for standard setting grows, so are concerns that the leading groups equipped to vet new models are saddled with potentially troubling conflicts of interest. Several have close ties to the doomsaying effective altruism movement that regards AI advances as a threat, leaving some government officials and company executives questioning these evaluators’ independence and objectivity.

In Washington, US officials are seeking to learn whether the groups could help create safety standards that give companies across the economy confidence to deploy AI tools. The Commerce Department’s Center for AI Standards and Innovation, or CAISI, has been meeting evaluators, including METR, which investigated a breach involving OpenAI models, and the AI Evaluator Forum, a consortium formed last year to help set best practices for assessments, according to people familiar with the matter.

Those discussions remain in the early stages, and the goal is to explore how these groups might help set standards and evaluate AI systems, said the people, who spoke on condition of anonymity to describe private conversations.

Read More: Trump Rejects Push for International Pact to Manage Risks of AI

Yet ties between many safety reviewers and the effective altruism movement threaten to turn third-party evaluators into a nonstarter with President Donald Trump. His former AI czar, David Sacks, has dismissed outside assessments as “pseudoscience,” and the president has belittled existential AI fears as “a hoax” and warned against reining in the technology.

While Sacks, a venture capitalist who now leads the President’s Council of Science and Technology Advisors, has said that wider transparency and audits of AI companies could be a good idea, he has also expressed skepticism of independent assessors. In a post on X last week, he called it a “plan to embed unfireable EA minders inside every AI company as the new trust and safety layer.”

Sacks was referencing his distaste for effective altruism, a popular philosophical movement in Silicon Valley whose adherents strive to ground every decision in mathematical probabilities they believe would maximize benefit for other people. Its supporters generally favor slowing AI development because some working on the technology believe it could lead to human extinction.

Spokespeople for the Commerce Department and CAISI didn’t respond to requests for comment. A METR spokesperson confirmed that the group has met with CAISI to share its evaluation results and at times discussed AI testing standards. A spokesperson for the AI Evaluator Forum declined to comment.

Tensions over the role independent groups should play in assessing AI safety reflect the broader battle under way among policymakers and the technology industry over how to address concerns that AI could slip beyond human control. For Trump, tapping the brakes on AI development would undermine his economic agenda and hinder a technology credited for historic stock market gains, one of his benchmarks of presidential performance.

In response to a spate of high-profile security breaches, Anthropic Chief Executive Officer Dario Amodei has promised to slow the pace of development of cutting-edge systems and bring in outside evaluators with employee-level access to vet new models for safety. Some top AI users, including Target Corp. and the Mount Sinai Health System in New York City, have voiced support for third-party assessments to boost public trust in the technology.

Calls for third-party assessors gained traction following the collapse of efforts over the summer to bring the government and various industries together on AI safety standards, which included Treasury Secretary Scott Bessent’s proposal for a FINRA-like model. Since then, tech companies have mobilized to pursue standards without the administration’s help, according to a person familiar with the matter.

Read More: OpenAI’s Sam Altman ‘Confident’ Industry Can Police AI Safety

OpenAI, Anthropic and Alphabet Inc.’s Google DeepMind have already begun discussions about standards. Last week, Anthropic said it would embed evaluators from Accenture to test the safety of its models, and OpenAI said Tuesday that it also plans to let outside groups vet its models for safety risks in earlier phases of the development cycle.

“We’ve hit a starting gun moment,” said Andrew Freedman, the co-founder and chief executive officer of Fathom, a nonprofit that focuses on AI policy and has been advocating for a regulatory framework of third-party evaluations. “There are dozens of groups lined up to be independent verification organizations, but there needs to be a competition for how good you are.”

Beyond those steps by industry, Freedman said he sees a role for the government in “verifying the verifiers.” That’s a task already performed by the Commerce Department’s National Institute of Standards and Technology for other industries.

Read More: Trump’s AI Defense Defies Voter Unease Heading Into Midterms

But the idea of vetting outside assessors has fueled questions about the reviewers’ objectivity. Many of the groups fit to conduct those tests — such as members of the AI Evaluator Forum, which include METR — are closely connected with effective altruism. That’s the same community that supports reining in AI development for safety reasons.

One person familiar with the AI industry said that while the concept of third-party auditors makes sense, the current ecosystem is small and lacks diversity in thought leadership because most reviewers are tied to effective altruism. US officials have been sensitive to those relationships, the person added.

When it tapped Accenture to conduct audits, Anthropic said that Faculty — a company Accenture acquired earlier this year — would lead the partnership. Prior to the acquisition by Accenture, Faculty had raised $54.1 million, including funding from Jaan Tallinn, a top AI safety advocate who co-founded the effective altruist Future of Life Institute.

Read More: Altman, Amodei Call on UN, World Leaders to Boost AI Safety

In response to a request for comment, Anthropic spokespeople pointed to their policy page, where the company lays out its support for third-party evaluations and model reviews by governments. Anthropic also says there that it submits its models to CAISI and the UK AI Security Institute for pre-deployment testing. Spokespeople for Accenture declined to comment.

METR, which stands for Model Evaluation and Threat Research, has worked with AI companies including OpenAI and Anthropic to assess risks. Led by Beth Barnes, a former OpenAI employee, METR got its start inside another nonprofit called the Alignment Research Center before spinning off as its own project and has received funding from the nonprofit, known as ARC. According to a 2024 tax filing with the Internal Revenue Service, METR reported it received a $4.5 million grant from ARC that year.

In response to Sacks’s public comments, METR said that it doesn’t believe any small group should have authority over what happens with AI and that its goal is to bring information from inside the companies to light. METR also strives to hire people with competing views and its published results do not “cleanly map” on to a doom-centered or AI acceleration-centered worldview, the spokesperson added.

ARC is run by Paul Christiano, a former OpenAI employee who recently joined OpenAI’s nonprofit board. Christiano is also a senior technical adviser for CAISI, the organization within NIST that helps test and develop guidelines for AI models.

Read More: World Labs Founder Fei-Fei Li Urges Independent Oversight of AI

Christiano’s nonprofit ARC has received more than $1.5 million from Coefficient Giving, previously known as Open Philanthropy, one of the best-known grant providers based in the effective altruist movement. In 2022, ARC received one grant of $265,000 and another for $1.25 million, according to Coefficient Giving’s website.

ARC’s vice president of research, Jacob Hilton, said that when METR later spun out of ARC, ARC went through a process to divvy up its assets and determine how much to grant METR. The Coefficient Giving grants were put toward ARC’s ongoing work, rather than to METR, he said.

Beyond METR, other listed members of the AI Evaluator Forum, which helps companies access what they need for evaluations and seeks to set shared best practices for participating model evaluators, also have connections to effective altruism, particularly via funding.

In a post responding to Sacks on X, the group wrote that “trust is essential and it’s earned through transparency, independence, and the freedom to speak truth to power.”

“AEF is asking AI developers to provide these basic guarantees to ensure AI evaluations can be done credibly and to the highest scientific standards,” the group wrote. “This shouldn’t be controversial.”

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论