Evidence suggests RSI may decelerate

Why recursive self-improvement can be slow

The most familiar way to envision recursive self improvement is as causing accelerating improvement in model capability. As the machine gets smarter, the self-improvement gets faster and the result is exponential growth, or something even beyond exponential. One can even imagine the singularity popping out of the other side of RSI, super-exponential growth converging to infinity in finite time. This is not, however, the only possible form that RSI can take.

As much as possible, imaginative exercises should be disciplined with quantitative reasoning. It does not take much formal investigation to show that recursive self-improvement can lead to many different outcomes, ranging from singularity in finite time to a rate of improvement that exponentially decays. This is because the natural tendency of improved model capability to accelerate the effective optimization power being used for RSI must compete against the tendency of the optimization problem to grow more difficult as the model capability improves. If both the recalcitrance of the model optimization problem and the optimization capability of the model grow as the model improves, then the central tendency for RSI to accelerate or decelerate over time depends on which grows faster. This tendency is of course affected by other factors such as available compute, but it is itself of central interest for understanding the prospects for RSI in different relevant scenarios.

Not too long ago there was no empirical evidence to show how to model recursive self improvement. Now we do have data. In this discussion I will both consider relevant data, and raise some considerations concerning algorithmic optimization. My conclusion is that gains from RSI should be expected to decelerate over time. If this is accurate, the more dramatic scenarios of existential risk become less plausible.

Let’s begin with a simple model of what RSI might look like. Nick Bostrom offers a simple model in Superintelligence. There he models the increase in intelligence of a recursively self-improving model as:

Here the left side of the equation represents the change in intelligence over time. On the right side, represents the effective computational effort directed at (ptimizing) (ntelligence of the AI system). In Bostrom’s formulation is proportional to (that is, the intelligence of the AI engaging in recursive self improvement). The idea is that as improves, will grow proportionately. It follows that if is constant, will increase exponentially.

is also influenced by the amount of compute available to the RSI system. While in practice we of course are interested in accelerating model capability via increasing the amount of compute available, this is a tangential consideration for the central issue around RSI. The question that concerns us here is whether increasingly capable models can make their own improvement accelerate when research and training compute are held fixed, due to the tendency of their own increasing intelligence to accelerate their research process.

stands for recalcitrance. which we will use later, gives the total effective computational effort required to create a model with intelligence . is its derivative, the rate of change in the recalcitrance. Recalcitrance is a very important idea. It turns out that how RSI goes depends really on the relationship between those two quantities, and . The obvious one is that it goes faster as the system puts more effective research power into optimizing itself: this is . depends on how much compute the system has to work with, but, more importantly, it depends on the optimizing capability of the system itself, and this goes up as the system gets more intelligent. This is the thing that people notice which makes them think RSI will accelerate.

But the other factor is how hard it is to optimize intelligence. And this almost certainly gets harder as the system gets more intelligent: we will discuss this a bunch more later, but suffice it to say that optimization problems like this almost always gets harder the more the system is already optimized, and the empirical evidence suggests that in this case the optimization gets harder quite fast, which is exactly what we expect. If this is correct, then is a function of . So even though the optimization power goes up as the AI gets smarter, the difficulty of the intelligence-optimization problem also goes up as the AI gets smarter. So whether RSI accelerates or decelerates over time depends on whether or grows faster as increases.

Recent work by Davidson and collaborators models essentially the same contest in a different parameterization. Their ‘software’ models include a parameter for improvements becoming harder to find, and summarize the balance between this rising difficulty and increasing AI research power with a returns parameter ; progress accelerates only when , although the way changes as the frontier advances is largely assumed rather than derived. The analysis below can be understood as an attempt to model the recalcitrance side of this balance more explicitly.

We can also see that the idea of a singularity, of RSI leading to infinite intelligence in finite time, is implausible if nis an increasing function of . For RSI to accelerate so fast it hits infinity in finite time requires very, very rapid acceleration. It needs some combination of either recalcitrance decreasing as the system gets smarter or the system’s ability to optimize its own intelligence growing much faster than its intelligence does. Far from being the inevitable outcome of RSI, an unbounded intelligence explosion requires highly unrealistic conditions to be possible.

Just because the AI is getting smarter over time does NOT guarantee that its intelligence-optimization research will go faster. The rate at which its intelligence grows may well decline over time. It entirely depends on whether making these systems incrementally smarter gets harder as their intelligence goes up, and if so how much harder it gets as their capabilities improve. So let’s look at what the evidence says.

The empirical evidence

Since there are no AI systems yet that can do recursive self improvement, we do not have any good measure of or a way of directly estimating how the difficulty of improving changes with time. However, we have evidence about the effects of optimization effort on the capabilities of models more generally. This can show us the pattern about how improving the capabilities of an AI gets harder as it gets more capable.. Let be a valid measure of general model capability. If we have data about how much optimization effort was used to produce various models and a measure for these different models, then we can calculate the marginal increase in as increases. Since is a measure of general model capability, the ability of a model to optimize recursive self-improvement (that is, ) is plausibly a function of . Consequently we can determine how grows relative to by looking at how grows relative to and estimating the relation between and . This will then predict whether RSI intelligence growth will accelerate or decelerate over time.

Epoch AI has a measure of general model capability, ECI. This scale is based on a complex methodology that attempts to create a valid linear interval scale for model capability: each point of increase anywhere on the scale should correspond to an equal amount of increase in model capability. It also has data about the training compute that was used to create the many different models it has ECI scores for. I will therefore use ECI as an empirical measure of for the rest of our discussion: other empirical measures of are of course possible. Given Epoch AI’s data, we can calculate , and then discuss how ECI can be used to estimate and what the overall implications for RSI are. Similar analysis has been performed in the recent paper The Economics of Recursive Self Improvement, which I will discuss after I explain and develop the general approach.

The efficiency of our training algorithms is increasing over time. A given amount of compute used to train a model back in 2024 produced worse models with lower ECI than the same amount used in 2026. So we can represent the relationship between ECI and total effective optimization power used to produce the model like so: , where is the compute used to train the model and is an acceleration parameter that increases with time and is empirically estimated from the data. Each model that we have an ECI score for, a release date, and a training compute budget allows gives us a data point for function f. The recalcitrance of ECI improvements is the reciprocal of the slope of .

Here’s a graph showing the Epoch data, the regression line, and the takeaways.

Historical model capability versus effective training input

The key insight this gives us, which is not sensitive to the precision of the regression, is that the recalcitrance of ECI grows geometrically with linear ECI improvements. According to the regression, the difficulty of increasing ECI grows by about 14.6% for every point ECI increases. This is very, very rapid growth. For example, the amount of effort required to increase ECI from 200 to 201 is about 800,000 times the amount of effort required to increase it from 100 to 101. So for automated RSI to be self-accelerating, a system with ECI of 200 would have to be MORE than 800,000 times as effective at using the same amount of training compute to improve LLMs as one with 100.

This estimate is quite similar to the result found by Cunningham et. al., whose estimate of recalcitrance was slightly higher than mine. They calculated that in order to achieve accelerating growth, the optimizing power of AI would have to increase by about 17% per point of ECI. In either case, the key takeaway is that for self-accelerating RSI to even be possible, optimizing power of the AI would have to grow as a geometric function of ECI. If it just grows proportionately to ECI, automated self-accelerating RSI can’t happen.

The Cunningham et. al. paper also tried to estimate the actual relationship between optimizing power and ECI, to see, effectively, whether the evidence suggested it was growing faster than the recalcitrance. To do this, they relied on a reported 4x increase in productivity reported by Anthropic employees switching from Claude Sonnet 3.7 to Opus 4.8 over the course of a year. This is not data that really suffices for an accurate estimate (and to be clear, Cunningham et. al. do not consider it to be). A single estimated data point that treats all-factor productivity increases in the work of human computer scientists on a variety of tasks as entirely caused by an increase in model capability is already on somewhat shaky ground as telling us anything valuable. But using that number to estimate the optimizing power of non-existent models engaging in autonomous research (which neither Sonnet 3.7 or Opus 4.8 can do even a little bit) has very little validity. But if we accept their estimate, it corresponds to about a 9% growth in optimizing power per point of ECI, which is lower than it needs to be whether we take my or Cunningham’s estimate of recalcitrance.

Ramez Naam argues, rightly, that the Cunningham estimate of the relationship between ECI and optimization power is unsatisfactory, and constructs his own estimate based on better data. His estimate is that a point of ECI increase gives about a 3% increase in optimization power, one fifth of the value it would need to be in order to make accelerating RSI possible. However, like Cunningham et. al. Naam’s estimate is based on evidence about the effects of improvements of model ECI on the productivity of human workers using the AI for research tasks. It is therefore a very imperfect proxy for the extent to which the optimizing power of an autonomous AI system engaged in RSI will scale with its general capability level as measured by ECI. However, the best estimates along these lines, whether we accept Naam or Cunningham, suggest that recalcitrance grows faster than optimizing power as a function of model capability, and therefore that automated RSI would decelerate over time.

If anything, I think that the estimates we get from Cunningham and Naam are overestimating the prospects for accelerating autonomous RSI. We will discuss why in the next section.

A second bottleneck: improving the training algorithm

LLM’s are created through a multi-step process, but the central element is ‘pretraining’, a process where large amounts of data and computing power are brought to bear in order to create the parameter values that define the model. These parameters are generated via a training algorithm, which recursively adjusts them throughout the training. A better training algorithm will create a better model from the same data and computing effort, and improving our training algorithms is one of the main research goals of AI labs. Insofar as RSI is being carried out by LLMs, we can represent their ability to enhance the research process by their ability to optimize the training algorithm. Optimizing the training algorithm is a separate problem from improving the capability of the LLM but is a necessary part of the RSI process. Modelling the interaction between training algorithm optimization and model improvement suggests that thus far our considerations around RSI may have been overly optimistic, and self-accelerating RSI may require an implausibly rapid, double-exponential growth of optimization power as model capability increases.

Suppose that, as is currently the case, new models are developed by allocating some amount of training compute to a training algorithm. The optimization effectiveness of the training , where is a representation of the efficiency of the training algorithm. Just as in modern LLM development, the same amount of training compute will have more optimization effectiveness and therefore produce a stronger model if is higher. The strength of the model, measured as its ECI, is a function of and , with given, if we wish, by our empirical estimate derived above. For the purpose of this model, we assume that both research compute and training compute remain constant, since the question at issue is whether model improvement in and of itself is sufficient to generate accelerating RSI capabilities.

We model RSI by assuming that the best model available at a time uses some separate research compute budget to optimize the training algorithm so that the next training run to generate a new version of the model will benefit from the improved algorithm. After each round of research, the new and improved model uses the research compute to develop the next-generation training algorithm, and the new training algorithm is deployed with the training compute to develop the next generation research model.

Given what we already know, we can predict that if the ability of the model to optimize increases by more than 14.6% per point of ECI, we should expect accelerating model improvement, and gradually slowing model improvement if it less. But the key theoretical consideration is this: based on what we know about algorithmic optimization, we should expect that the difficulty of optimizing should itself increase as a function of ; or put in the terminology from earlier, has its own recalcitrance that is almost certainly an increasing function of . This is because algorithm optimization is a search through a possibility space of algorithms, and the fraction of the possibility space that represents an improvement on the algorithm you currently have necessarily declines as your algorithm improves (we will develop this point further shortly). So we have a recalcitrance value, , which is itself an increasing function of : .

At a given time step (t), the current model has capability and research effectiveness . With a fixed research-compute budget , effective research effort is .

The efficiency of the best training algorithm it can discover, , is given by:

where is the amount of effective research effort required to discover a training algorithm of efficiency .

Given fixed training compute , effective training optimization . is the function that gives effective training optimization needed for a model with ECI. So . It follows that

Rearranged:

This gives us the ECI increase from one generation to the next. The three key input functions are the two recalcitrance functions and the efficiency function on . The efficiency function on relates ECI to optimization power: Cunningham et. al. as we saw gave a rough estimate which Naam attempted to improve. Expressed as functions, their estimates are:

Cunningham and I both have very similar estimates of the recalcitrance function on :

But the main takeaway of this modelling exercise is to highlight the recalcitrance of . Cunningham et. al.’s model effectively incorporates the recalcitrance of via estimating the research effort used on algorithmic improvement and the rate of algorithmic improvement. In the terminology used here, Cunningham et. al.’s model corresponds to a power-law recalcitrance function:

The historical data they use places beta approximately equal to 1, which corresponds to linear growth of the recalcitrance function. If grows linearly with as Cunningham et. al.’s historical estimate suggests, the recalcitrance of effectively acts as a constant multiplier to recalcitrance function on . In that case, if and are both geometric increasing functions what matters is which has the larger exponent. and increases the necessary exponent for to suffice for accelerating improvement by some constant multiplier. But if grows geometrically, then would have to grow super-exponentially in order to be sufficient for accelerating self-improvement.

The historical data Cunningham uses are not really sensitive enough to let us estimate accurately whether grows linearly or exponentially. It would appear, therefore, that the discussion we find in Cunningham and Naam represents a kind of best-case scenario for accelerating RSI, since they assume linear growth in and both already estimate that recalcitrance is too high for self-accelerating automated RSI to be possible. If, as I will shortly argue, it is most plausible that grows at least geometrically, then whether Cunningham or Naam’s estimate of ECI’s optimizing power is more accurate will be irrelevant: neither estimate is super-exponential so RSI will decelerate.

Searching the possibility space of training algorithms

Considerations from algorithmic optimization theory can give us some guidance on how to evaluate the likely form of . We can consider the algorithmic improvement process undertaken by the model as a search through the possibility space of training algorithms. The effective ability to search the space is given by , and then the recalcitrance function for algorithmic improvement should represent the intrinsic amount of effective search required to reach an algorithm of quality . A useful way to understand this is in terms of the geometry of the search space: as the current frontier improves, the set of candidate algorithms that outperform it occupies an increasingly extreme and, normally, increasingly sparse region of the space: the ‘tail’ of the distribution of efficiency within the space. So algorithms that are better than the frontier become more sparse as the frontier improves, and the amount of effective search required to improve the algorithm increases as the improvements grow more sparse and thus harder to find. Because algorithm quality is usually produced by many interacting constraints and trade-offs, extremely good solutions are expected to be exceptional combinations that rapidly grow rarer in the space as performance increases. In a ‘light’ tailed distribution of this kind, the probability of finding an algorithm beyond a frontier falls at least exponentially with . This corresponds to geometrically increasing recalcitrance, or worse.

If search through a light-tailed possibility space of training algorithms is a good model of the process of algorithmic improvement, therefore, accelerating growth in model capability is only possible under RSI if the efficiency gains in the model’s ability to search the possibility space of training algorithms is a super-exponential function of ECI. This seems unlikely.

A bit of context can be given to these considerations by looking at Chen et. al..’s Lion experiment. Chen et. al. implemented an automated, evolutionary search for better training algorithms starting from an existing optimizer AdamW. They describe the search space their evolutionary search algorithm explored, which is the search space of possible training algorithms, as effectively infinite and very sparse. Their search algorithm mutates successful optimizer programs and tests the resulting candidates on proxy training tasks. Progress is initially fairly rapid but then exhibits strong diminishing returns, eventually plateauing. I estimated the best fit for their improvement curves as a saturating exponential curve, suggesting a very sparse search space indeed, but the general inference that the search space is very sparse is what matters.

When their search saturated, the researchers respond by restarting the evolutionary search around the best optimizer found so far, which opens up another period of improvement before returns again begin to diminish. This process ultimately produced LION, an optimizer that performs competitively with AdamW while substantially reducing training compute on some tasks. Overall, all that research compute seems to have produced only a modest improvement in the overall efficiency of the training algorithm: the released Lion data does not suffice to estimate it.

The evidence from Lion suggests a sparse search space. Search repeatedly exhausts relatively accessible improvements and needs to be restarted using the best found candidate, only to saturate again. Estimating the overall recalcitrance of the space with any accuracy is not possible, because we don’t have good evidence about how the curves will change with repeated restarts of the research process. But insofar as we can draw any conclusions about the space of possible training algorithms from the Lion experiment, a geometric increase in recalcitrance seems if anything optimistic.

Summing up, the evidence suggests a very different picture than what one might intuitively imagine automated RSI to look like. While the capacities of the model improve as RSI progresses, the difficulty of improving the model itself and the difficulty of improving the training algorithm both also increase. Whether RSI overall leads to accelerating or decelerating increase in model intelligence depends on the interaction of these factors. On what seem to be the most plausible understanding of the empirical evidence and the nature of algorithmic optimization processes, the rate of improvement will decline with time, possibly quite rapidly if the search spaces of both models and training algorithms are sparse. The discussion by Cunningham et.al. and Naam is probably too optimistic, since Cunningham’s model effectively uses a very optimistic value for the recalcitrance of training-algorithm optimization. And even with that very optimistic assumption, their empirical estimates still support decelerating RSI.

Just because RSI will tend to decelerate, doesn’t mean that progress at improving models will slow in the near term. We have been very successful at increasing research and training effort, so the rate of improvement in ECI has been more or less constant despite the rapidly increasing recalcitrance. However, long term, the viability of accelerating model improvement would seem to depend on the ability of RSI to function given more or less stable compute. And certainly doomsday worries about hostile superintelligence developing massively beyond human levels in short time periods without being noticed or prevented tend to rely on the possibility of rapidly accelerating RSI. Those scenarios seem quite implausible if the considerations here are accurate. The hostile superintelligence’s ability to improve itself recursively will slow over time, perhaps quite rapidly.

Our ability to estimate these key factors will no doubt improve over time. Data about the recalcitrance of optimizing training algorithms would be especially useful: it could let us estimate the actual change curve for RSI. Until those better estimates are available, our expectation should be that recursive self improvement will tend to decelerate.

  1. Bostrom, Nick. 2014. Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press, pp. 65, 75–77. The expression of the formula is modified slightly from Bostrom to fit better with the clearest formulations of subsequent models.
  2. Davidson, Tom, and Tom Houlden. 2025. “How quick and big would a software intelligence explosion be?” Forethought Research, August 4. See “Model dynamics” and “Limitations.”
  3. This is very similar to the value calculated above (14.6%), which is unsurprising since both estimates are based on the same data set. Cunningham takes Epoch’s capability-versus-training-compute relationship and estimates the slope of ECI against log training compute. Then they assume algorithmic efficiency and training compute are interchangeable through so that the compute slope can stand in for the algorithmic-efficiency slope. The regression developed for this paper instead constructed effective optimization power for the 72-model dataset that had reported training compute values by adjusting raw training compute for algorithmic efficiency, and then directly regressed ECI against that effective-compute quantity.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论