Is METR A Meaningful Check On Anthropic?

Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.

METR is not capable of being a meaningful check on Anthropic. METR is not meaningfully independent, is not sufficiently staffed, and has no authority over Anthropic that cannot be revoked at Anthropic’s discretion. Suggesting that embedding METR into Anthropic would be a meaningful check on Anthropic is so suspicious that it looks like an attempt to evade oversight and to sabotage attempts at oversight in general.

If Dario does not really mean to suggest that METR could be expected to meaningfully check Anthropic, or he actually believes that METR could check Anthropic in the same way a government regulator can check a bank, something has gone badly wrong in formulating and communicating this policy.

Why METR Cannot Check Anthropic

You don’t have to take my word that they’re closely tied to Anthropic, you can take METR’s:

Also note that some METR staff have strong social ties to employees of AI companies, and METR currently works out of a shared research center (Constellation) which hosts some AI lab staff.

I appreciate METR’s professionalism and candor, both in general and in this report, because those are not always easy to find. I do think that this somewhat understates the problem: If you required METR employees who dealt with Anthropic to be “independent” the way an accountant doing an audit is legally required to be, METR would probably have to recuse itself from consideration for conflicts of interest. An accountant, for example, would generally be considered conflicted if they had lived with or dated the people they were auditing. It would also be important that the management of the accounting firm, controlling the accountant’s career, did not have that sort of history with management in the firm being audited.

This proposed arrangement, given the choice of METR, is nothing like embedding a regulatory supervisor at a bank. Bank regulators work for the government and will send bank employees to prison if they do the wrong thing. METR employees would mostly be hanging out with their friends or their boss’ friends at work while drawing non-profit salaries.

If METR wants to establish that it isn’t hopelessly conflicted here, they should hire an accounting firm to publicly review them for conflicts of interest, or sue me for saying they have them. I think Dario suggesting them as a third-party evaluator analogous to a banking regulator amounts to a malicious lie intended to sabotage regulations and audits, and METR’s failure to correct the record makes them complicit in the lie.

It seems strange to even call this regulatory capture. They are not at all regulators, and they are not being captured; they were never meaningfully independent, and were never regulators. Dario is simply asserting that they are like regulators. This is so far from even being plausible that it is baffling that anyone would say it. If this arrangement came up in court it would be called incestuous, and if this is the basis for anything like a regulation or a legally binding agreement between companies it will come up in court.

This is my first and most serious objection to this suggestion, but there are more.

Staffing is also a problem here. METR does not have sufficient staff. Point blank, as a matter of numbers, there are not enough people at METR. METR tells the press that it has about 35 employees, and I don’t think more than a tiny fraction of those employees are directly working on model evaluations. I am not entirely sure I could start a poker game with METR’s relevant staff. Team sports are right out. Anthropic employs thousands of people and has millions of users. Due to the relevant numbers alone, this is not a serious regulatory proposal. This is a fig leaf, and maybe the plan is to fake it til you make it, to try to hire a ton of people and make this reasonable in the future even if it isn’t now, but that is not good enough given the seriousness of the situation. Anthropic is a trillion-dollar business and one of the most important single endeavors on Earth. Even if they were totally independent, METR could not meaningfully check or even record what Anthropic is doing.

Next, METR’s finances are not meaningfully independent of Anthropic. This applies two ways: first, METR refuses cash compensation from AI labs for its services, but it accepts tokens from those labs for research purposes. As it happens, tokens are worth cash; in fact, they are the main product the labs sell. METR is quite professional, so I am positive they would never re-sell tokens they were given for research purposes, but if they did they would be able to make many millions of dollars doing so. Refusing cash but accepting tokens worth millions of dollars may be a reasonable ethical position, but it is not being “independent”.

The other issue with METR’s financial independence is tricky, complicated, and annoying to explain, but we’re going to try anyway. Coefficient Giving is the main entity disbursing money for non-profit activity concerning AI, its primary funders are investors in Anthropic, and Coefficient’s former CEO is named Holden Karnofsky. Karnofsky is directly thanked in some of METR’s early work, currently works at Anthropic, and is married to Anthropic’s President, Daniela Amodei. Coefficient Giving, which was then called Open Philanthropy, funded the Alignment Research Center, from which METR spun out. METR appears to have, mostly, cut that tie, but its most recent funding includes ten million dollars from a RAND Corporation program that was, in turn, funded by Coefficient Giving. I am going to hazard that the terms of that grant were such that there were only a few plausible grantees here, and METR may well have been the only one. This feels like a shell game: I can’t prove that this money was given to RAND on the tacit understanding that it would end up at METR, but it probably was. (Also, several prominent current METR employees previously held prominent roles at Coefficient Giving.)

Last, and, honestly, least? This is not real regulation. Real regulation is imposed by the government, and you cannot simply opt out of legally required oversight. METR is full of smart people, and they will be well aware that if they don’t play nice they can simply be kicked out of Anthropic. Anthropic could at least have entered into some legal agreement with a real auditor which gave that auditor enforceable oversight powers, beyond publication rights, but instead they chose this.

Dario is calling here for the AI industry to self-regulate. That is not, normally, a thing. The cases where societies allow industries to self-regulate are broadly those where regulation raises free speech questions, centered around art. The most famous examples of these are the age-related rating systems for movies, TV, music, and video games. These are industries that provide entertainment, not fundamental technology. Most importantly, none of these industries claim on a regular basis that their technologies threaten national security or human survival. Dario does.

How Did We Get Here

Why would anyone think that this should fly? If you claimed, in a contract, that you had an independent auditor, and when I looked into it I found out that they got paid with the same money you did, lived in the same house you did, and one of their executives married your sister, you would go to prison. You would go to prison because if you don’t send people to prison for fudging the independence of their auditors and business connections, they tend to end up like Bernie Madoff and Sam Bankman-Fried. AI is more important than the average business, and accountability should be more, and not less, important.

We live in a relatively high-trust society, where you can safely assume that a signed contract will be honored and your bank really does have your money, because we value impartiality and fair process, and we set high standards for what those are. Impartial application of the law is the absolute bedrock of modern civilization, and even the lowest criminal is entitled to a judge and a jury that does not know him. These rules and how we enforce them are what keeps our society honest and fair, and following them, as much as we do, has made us prosperous and safe.

Anthropic’s company culture is apparently not a part of this tradition and does not believe in it. Anthropic doesn’t think it should be regulated the way that banks, insurance companies or law firms would be, or that it should be required to seek impartial auditors to do security reviews of current practices or of past breaches. Close associations do not matter, because impartiality does not matter, and in fact, the more closely connected you are to Effective Altruism, Anthropic and Coefficient Giving, the more trust you deserve. This seems to be why Anthropic thinks that METR is, of course, very trustworthy, and why they have a recent history of giving other important security contracts to very incompetent, but very Effective Altruist, companies.

One of the aims of democratic society is to prevent you from needing to know the proclivities of particular individuals in the Bay Area. I would love it if we had a normal liberal-democratic process here, instead of one where I have to care about who is sharing funding, housing, and fluids in San Francisco and Berkeley.

METR, “Frontier Risk Report (February to March 2026)”, May 19, 2026. See the discussion of evaluator independence in Appendix A.

Business Insider, reporting on METR’s staffing and the AI safety research talent shortage, August 2026.

METR, “Funding update”, August 14, 2026. METR describes its funding policy and the free tokens supplied by frontier AI companies.

METR, “Update on Security at METR”, August 31, 2026, “Incident 1.” METR reports that an attacker used a stolen API key over three weeks to consume credits worth approximately $600,000, supplied to METR for free by the model developer. METR also reports that the key had no spending limit at the time.

Coefficient Giving, “Open Philanthropy Is Now Coefficient Giving”, describes its founding funding partnership with Good Ventures, the foundation of Cari Tuna and Dustin Moskovitz. Anthropic’s Series A announcement names Moskovitz as an investor; TIME’s profile of Tuna, September 5, 2024, also describes the couple’s 2021 investment in Anthropic.

The Alignment Research Center grant records list Open Philanthropy general-support grants of $265,000 in March 2022 and $1.25 million in November 2022, with links to the original grant announcements. METR describes its origins in “ARC Evals is spinning out from ARC”, September 19, 2023.

Grantmaking.ai records a $10 million Coefficient Giving grant to RAND for “AI Evaluation and Testing”, dated September 20, 2025, and links to the Audacious Project’s Canary collaboration between METR and RAND. METR’s “New Support Through The Audacious Project”, October 9, 2024, reports approximately $38 million for Canary, with approximately $17 million supporting METR’s work.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论