Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion

Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000x faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028-2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040-2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off.

Fast Reasoning, Slow Learning

The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for that model, followed by automated next-model building. This teaches LLMs deep skills, but operates at the speed of next-model building (weeks to months for one iteration of advancing the deep skill frontier) rather than at the speed of next-token generation (100-1000x the human speed). Humans are not the key bottleneck to the speed of next-model building loops, there's still a lot of waiting for the compute to do its thing in training, so achieving prosaic RSI by teaching the near-future LLMs all the skills necessary to perform it automatically doesn't make it go too fast. Using smaller LLMs to make everything faster doesn't work because the current frontier LLMs are probably borderline insufficient for learning the prosaic RSI skills that automate the next-model building loop. The LLMs of 2028-2031 (that are very likely sufficient) will be even bigger, though the cost of training or running them is not as bad as "quadrillion total params" sounds. This cost can't be circumvented by using different hardware that makes LLM inference much faster, because different hardware doesn't reduce the necessary number of FLOPs, which are not terribly wasted even in RL training and inference that involve bandwidth bound decode. Fundamentally, cost is the amount of compute, and the only thing that overcomes it is the scale of the global buildout.

Compute Slowdown, Industrial Explosion

Since LLMs become ready to close the next-model building loop (that enables AGI) just as the human industry runs out of various kinds of fuel for quickly increasing the scale of the compute buildout, there is no opportunity for another near-term 1000x increase in the available compute (and thus the speed of next-model building loops), the way compute was increasing in 2022-2027, and the way it'll keep increasing (a bit slower) in 2028-2032 until the pace of decommissioning old compute somewhat catches up to the 2028+ pace of producing new compute, set by factors like availability of EUV machines and skilled human labor. Thus the big compute slowdown of 2032+, stronger than the end of the current exponential scaling of compute by 2028+.

Without paradigm-breaking algorithmic innovations, the slow-learning AGIs can't quickly invent such innovations, and so the more predictable component of the pace of progress is set by the pace of the compute buildout. But also, the AGIs (likely available since 2028-2032) make the industrial explosion of robot-building robots a predictable medium-term development. It probably doesn't start right away, since the AGIs are not much faster than humanity at taking care of all the novel engineering challenges (requiring many next-model building loops to get good), and before it goes into full swing humanity still needs to handle the industrial side of things (at the human pace).

The automotive industry and the compute buildout acceleration of 2022-2027 seem like good anchors for how this might unfold. The process starts once AGIs unlock an outsized demand for robots (by making them very useful for everything), and the industry starts reshaping itself to increase the supply as fast as it can, similarly to the consequences of the ChatGPT moment. Robot production exhausts the industrial capacity of the supply chains within 3-5 years (similarly to how it took 5-6 years to reshape compute production). At that point, the scale of the robot supply (anchored to the current automotive industry) approaches a significant portion of human labor, so the process continues right past the limits of human industry without another big slowdown. If the doubling time of the robot-building industry (autonomously operated by AGIs using the existing robots) is around 1 year, then 10 years of this process increase the industrial capacity about 1000x. The countdown should probably start from the end of the 3-5 year period of industrial conversion, when the robot industry first matches a sufficient fraction of the human industry to also start producing as much compute (together with all the other precursors), and the slow-learning AGIs probably need the time for the next-model building cycles to figure out how to automate everything.

Prosaic Timeline to Takeoff

The timeline starts with prosaic RSI in 2028-2032. The resulting slow-learning AGIs first make robots very useful generally within 2-3 years, in 2030-2034, setting off the industrial conversion of human industry towards robot production. This lasts another 3-5 years, and by 2034-2039 the automated robot-building industry operated by robots and AGIs first matches the human industry's capacity in terms of the compute buildout it can support. If this industry can quickly reach a doubling time of 1 year, it can 1000x the compute buildout by 2044-2050. At that point, the next-model building cycles take 1000x less time, and so the unpredictable paradigm-breaking algorithmic innovations necessary to set off a software-only singularity happen on the scale of months instead of centuries, and would've happened at some point earlier than that, probably by 2040-2045, but certainly by 2050. This is the upper bound on the timing of feasibility of superintelligence, which mostly assumes just the current paradigm (extremely hobbled in its efficacy at superhuman invention) rather than any particular future breakthroughs.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论