Could internal model transparency tame the AI race?

This is a cross-post from my Substack.

Motivation

As AI labs accelerate the automation of AI research itself (recursive self-improvement, or RSI), the speed of AI progress is becoming worrying even to frontier labs themselves. 1400 employees at frontier labs have signed the Pacing the Frontier letter, in support of the claim that:

To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.

Competitive incentives are a large motivation for labs to develop AI much faster and more recklessly than they would prefer. The challenge is that the first firm to automate AI research can run away with it, concentrating enormous power in the hands of one firm, and forcing their competitors to race as well. Thus, we need policy ideas that disincentivize firms from pursuing RSI, by muting their competitive incentives.

This essay proposes one such idea: internal model transparency.

Internal model transparency

A central reason for labs to pursue RSI, despite its risks, is that it gives them better research technology than their competitors. Imagine DeepMind focuses on scientific research while Anthropic focuses on coding. So Gemini gets an A on scientific research but a B- at coding. Anthropic can beat DeepMind by making Claude A- at coding, and then using that Claude to make another Claude that’s A+ at scientific research. As a result, RSI provides labs with a competitive advantage in every domain – meaning that labs that don’t pursue RSI will necessarily lose to ones that do. Thus, despite the risks, RSI is unavoidably attractive.

The advantage is real enough that labs reach for each other’s tools. In August 2025, Anthropic revoked OpenAI’s API access because OpenAI staff were using Claude Code ahead of the GPT-5 launch. This crackdown makes sense; giving competitors access to your best models defeats the purpose of having the best models.

This is the motivation for internal model transparency (IMT), which is the following proposal: once a lab deploys a model for internal use, they must also serve that model to researchers at other labs. In other words, firms cannot use their internal models as a source of competitive advantage in the AI race.

What would IMT do?

To see the effect that IMT would have on labs, it’s helpful to first imagine that labs don’t change their behavior at all, and continue to focus on automating AI research. What would happen then?

In the example above, RSI helped Anthropic beat DeepMind on scientific research, because they were accelerated by Claude being a better tool than Gemini. But with IMT, DeepMind would also have access to the same Claude, so Anthropic can’t use their superior coding capabilities to make a better model than DeepMind. In fact, you could go further; DeepMind could specialize in collecting complements to a frontier coding tool, like better task data on scientific R&D than Anthropic has access to. With this additional data, DeepMind could use Claude to produce a better AI science model than Anthropic can make with Claude. In other words, DeepMind can free-ride off of Anthropic’s coding advantage, but Anthropic can’t exploit DeepMind’s data advantage. Thus, DeepMind beats Anthropic in making a science-optimized AI.

The broader takeaway is that under IMT, having a model that’s really good at AI R&D is no longer a source of competitive advantage. This reduces the competitive incentive to automate AI R&D, with all the risks that it entails.

In fact, the example shows there is at least some incentive to free-ride off other labs: let them develop coding capabilities that will be made available to you, while you spend your money and compute on other parts of the AI training stack. The equilibrium is that every lab invests a lot less in frontier coding capabilities, and more in other sources of advantage. As a result, we can slow down frontier AI development without ever requiring it. Under internal model transparency, slower progress emerges from labs’ own incentives.

I see the disincentivization of RSI as the major reason to have internal model transparency. But there are two other substantial benefits to the proposal:

  1. It makes it much less likely that we end up with one company controlling the most powerful model. If all the intermediate models they had to build to get there were also accessible to competitors, it’s very difficult for one lab to stay ahead of all the others. This prevents a concentration of power within a single lab, in worlds where RSI occurs.
  2. It enables safety testing of internal models by third parties. The HuggingFace incident was led by an internally-deployed, “research-only” OpenAI model. If that model had been accessible to safety teams at Anthropic or DeepMind, then they could have uncovered issues that OpenAI did not. These risks would be even more detectable if IMT was extended to include trusted third-party investigators, like METR or Redwood Research, who led the investigation into the OpenAI-HuggingFace incident. The HuggingFace incident shows how much risk could be posed by internal models; IMT could offer a way to bring those risks to light.

Implementing IMT

A virtue of IMT is that it does not have to be an international agreement. A regulation covering only US labs already delivers the pacing benefit, because the racing dynamic it targets is primarily a race between US frontier labs. The main cost of staying domestic is leakage: Chinese labs could continue to automate AI R&D without constraint, while the US gives up some lead time.

If that cost is intolerable, IMT extends naturally to a US-China deal, where Chinese labs are also party to the regulation. China would likely agree, since IMT benefits follower labs. Of course, the challenge is verifying reciprocal cooperation – ensuring that Chinese labs are not withholding their own advances towards RSI.

But unlike other kinds of international cooperation, the verification surface is tiny. The only thing that the US has to verify is that a model internally available to researchers at a Chinese lab is also served to US labs, which is substantially easier than verifying other kinds of international cooperation (e.g. compute verification).

There are a few implementation details that are quite important to figure out. What is the list of firms included in this transparency list? How is it decided? How do you ensure that labs aren’t serving nerfed versions of their internal models to external users? How do you prevent labs from charging absurdly high rates when they serve internal models to competitors, to effectively restrict usage? These problems seem solvable, but they do actually need to be solved.

How does IMT compare to other AI regulations?

Safety regulations

One straightforward way to tame RSI would be direct rules on internal deployment: codifying the labs’ existing safety policies into law, requiring safety evaluations before a model can be used internally. This formula – proposing rules on AI development/deployment that labs must follow – is the predominant approach to AI governance.

But while these safety regulations regulate behavior, IMT regulates the incentives faced by labs. A safety regulation has to specify what counts as a dangerous capability, keep that definition current as the technology moves, and catch labs that cross the line. In contrast, IMT doesn’t require a regulator to judge whether any model is safe. It changes the payoff structure so that racing toward RSI becomes less profitable.

None of this makes IMT a substitute for direct safety regulation. But it does mean IMT asks much less of regulators than standard approaches to AI governance, while attacking the racing incentive that is the root cause of reckless AI progress.

Total research transparency

AI 2040, the governance plan proposed by the AI Futures Project, features an international coordination plan built on “total research transparency”: every algorithmic innovation ever made is immediately disclosed to all members of the public. The benefit of this approach is that it helps everyone get on the same page about how to align AI, and disincentivizes racing.

But total research transparency is a burdensome requirement, with a huge surface area to be monitored. It dovetails with other features of AI 2040 – in particular, the premise of an international deal to govern data centers and AI training. But total research transparency is not implementable without a sweeping international agreement.

Internal model transparency is a minimum viable form of research transparency: instead of sharing every innovation, labs share only the finished artifact – the model itself. And while AI 2040 argues for this transparency to be total, to all members of the public, I suggest the more modest version where this transparency is extended mainly to other labs, and possibly to trusted third-party evaluators.

Of course, an international form of IMT still requires an agreement, and it still requires some amount of verification. But the surface area of that verification is simply an API for models that is widely internally available in a lab. There is no need to monitor data centers, or researchers, or any other source of advantage for labs. Thus, IMT partially captures the benefits of total research transparency, while being strictly easier to implement.

Caveats

This is a rough idea with clear limitations:

  1. In the short term, IMT likely speeds up aggregate AI progress, by allowing lagging firms back into the race. This could be harmful; it is much easier to coordinate a slowdown across labs when there are fewer of them.
  2. Regulating internal deployment is really hard in general. Verifying that labs are complying with any rules on their internally deployed models requires much more technical capacity than the US government currently has.
  3. Better research technology is not the only motivation to pursue advanced coding models. More capable coding models are a huge source of revenue for labs, as shown by Anthropic’s revenue growth. So labs will still want to make advanced coding models; IMT only mutes one motivation for RSI.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论