An "Anthropic Principle" for Formulations of AI Alignment
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I explore one path toward formulating AI alignment without needing to formalize humans or our values, looking instead at the well-integrated computational power characteristic of agents capable of confronting the alignment problem.
One natural formulation of AI alignment, the study of how to be sure our highly capable software acts according to our interests, starts from unambiguous descriptions of human values (whether an AI starts with hardcoded values or learns them). This concept is mostly speculative, as we are far from knowing how to formalize our values precisely enough. My last article considered two reasons why we might want to hesitate to build in a model of humans. In my mind, the more important reason is that, from an engineering perspective, it seems wickedly difficult to formalize human values precisely enough, let alone formalize values of some extrapolation of “better” humans. There are also other objections that are more philosophical, e.g.: if our future AI systems meet Star Trek-style aliens who are like us but blue, is it really proper to ignore the aliens’ interests? Is it safe to assume that some extrapolation to “better” humans necessarily respects blue aliens, given humanity’s dismal record in other first contacts?
For the rest of this article, let’s step back to a technical problem: how do we formalize identification of which agents alignment should focus on? We have one conventional “test case” that is almost too obvious to state: in our present circumstances, this rule should identify people as the agents to focus on. There should also be good reason to believe the rule makes proper decisions in other circumstances. We also want it to be plausible that the same rule is substantially easier to implement in software than any approach based on reverse-engineering human values.
I’ll start with a big-picture consideration of how intelligence has increased in the universe over time, but I promise I’ll return to the original goal.
Who Could Think About Alignment?
The anthropic principle explains the apparently improbable fact that we inhabit a universe capable of developing intelligent life. An awful lot of random chance needed to line up properly to produce us. However, the question “how did we get here?” can only be asked by life intelligent enough to think abstractly and formulate the question. Yes, there may be many possible universes compatible with a broad formulation of physics, and likely only a small minority of them are compatible with intelligent life having evolved. Hence, contingent on someone being able to ask the question of how we wound up here, we must be in one of the universes conducive to intelligent life evolving.
Let me now develop this kind of thinking toward a principle that any agent at roughly the computational scale to support thinking about AI alignment should be defined as a proper focus for alignment. Instead of making a somewhat arbitrary choice up-front about which agents qualify, let’s try to define alignment once and for all to capture all beneficiaries. (We’ll come shortly to a twist that keeps literal application of this idea from taking us where we expect, motivating a fix.)
In this explanation, I’m going to embrace the computational theory of mind shamelessly, because it has always felt natural to me. This theory says that brains are understood as special kinds of computers, so thinking about computation broadly subsumes thinking about (biologically embodied) intelligence. For now, we’ll discuss at an abstract level, but the next section will concretize to societies of humans (not atomized individuals) as critical computational systems.
Some computational architectures can host the kind of abstract thought that leads to asking big questions about the universe. Some can’t. Let’s simplify for the moment and think in terms of one linear dimension of computational power. Consider how much such power is needed to be able to design recursively self-improving AI. Presumably there is some relatively well-defined minimum level. If self-improvement is kicked off thoroughly enough to cause an intelligence explosion, then quickly all the rules change, and the familiar alignment problem no longer applies, or at least its context becomes so different that the ideal rules are probably different. For example, AI designing AI is likely to have better explicit, formal/algorithmic understanding of its own values than we do, if it is successful creating its own successors.
These observations combine to suggest that there is a narrow window of computational power that naturally evolved intelligence can have when it is working on formulating alignment: it must have passed the minimum power threshold but not yet have begun an intelligence explosion. The window is narrow if we assume that solving alignment requires fairly similar tools to designing recursively self-improving systems. And let’s remember that we’re comparing slow biological evolution to faster deliberate engineering (a la Kurzweil’s Law of Accelerating Returns), so a “narrow window” might be, say, 1000 years.
So we can potentially formulate alignment in terms of the minimum level of computational capacity that could think about alignment. The definition wouldn’t literally be recursive – it would just happen to talk about the relevant computational level, which happens to be the one that we as humans inhabit (the coincidence explained in this way being reminiscent of the anthropic principle). Then we automatically “pick up” the agents we meant to detect.
Layered Intelligence
Is there one linear measure of computational power that predicts enough about the values of intelligence? Maybe not, but let’s add one more dimension and see what may get easier to predict. I’ve previously suggested a unifying slogan for this blog, “intelligence depends on organizing computation correctly and efficiently”, that is relevant here. Just measuring some notion of the number of atomic units of computation doesn’t do the job. It matters how those units work together to solve hard problems.
It’ll help to orient ourselves with a concept from biology: evolutionary transition in individuality. A great canonical example is more-primitive parts coming together to form cells, which come together to form organisms. At each step, units that had operated independently and even competed with each other find new ways to coordinate and function as coherent wholes. (To head off any confusion, I’ll note that while I’ve found writing on such transitions in an AI-futurism context, it seems to have gone in very different directions from what follows: humans merging with AI and humans by ourselves becoming larger social macro-organisms.)
The world of manmade computers exhibits similar layering. I argued previously for the opportunity to introduce new abstraction layers when AI is adopted pervasively enough, with the potential to provide similar benefit as from the celebrated digital abstraction that underlies most electronic circuits. A single computer, say one suited to act as an Internet server, is built out of components but is structured such that the computer appears to act with one purpose. Individual logic gates may fail or glitch, but it happens infrequently enough to be an exception to the rule, which we usually avoid thinking about. We then see companies putting many computers together into data centers. Through marvels of engineering, these computers can also be seen as acting together for single purposes. Because of the sheer scale, component failures are more common, but good design of distributed systems typically hides those failures admirably well, just as our human bodies prevent cancer at the cellular level most of the time but not always.
However, consider the step up from data centers to the whole Internet. Different data centers are owned by different organizations that compete with each other. It is emphatically not the case that the Internet acts as though it had a uniform purpose, abstracting away the actors that comprise it. Likewise, while humans may abstract their constituent cells well, a global economy populated by humans absolutely reveals the motivations of individual people. The human brain is thought to be able to perform the equivalent of about 10^15 floating-point operations per second (subject to standard caveats about brains being interestingly different from GPUs). The Frontier supercomputer hits about 10^18 in the same dimension. On top of the apparent lead for manmade computers, there is a critical difference between these substrates. Humans roll up into larger social groups, but we maintain individuality that gives rise to continuing conflict and competition within those groups. In contrast, supercomputers can be agglomerated into data centers with uniform governance, providing useful abstractions of working together on single computations on behalf of single owners. Hence, it remains relatively straightforward to increase the scale of a supercomputer (i.e., wait a few years for the next generation to be developed). While brains retain an advantage in power efficiency, we see that deliberate engineering has achieved a scale of internally well-aligned computation well beyond what biological evolution ever has, and it is plausible that this gap only grows with time.
I don’t mean to say that human brains considered as isolated computers are capable of formulating alignment and developing self-improving AI. Instead, we’ve depended on culture and social structures that agglomerate such brains into larger processes surprisingly close to the scale of frontier data centers. That scale may be necessary to tackle either AI-related problem. So, when we measure evolution approaching a threshold to be able to begin tackling those problems, we are considering not just the units with good integrated alignment (like individual people) but also the larger agglomerations that are needed to get the work done. One conjecture here, then, is that the first computational systems powerful enough to formulate alignment will be societies of relatively loosely integrated individuals: we need fairly powerful well-integrated individuals to get real thinking done, but they also need to work together at a pretty large scale to realize the important questions and work out solutions to them.
This observation suggests that we define a computational-power floor applied only to units with strong internal alignment. We backsolve for that floor by analyzing what seems to be necessary to support the right kind of abstract thinking, but I don’t think it’s necessary to characterize the larger, less-integrated computational systems directly – just to be inspired by examples in choosing constants. This move is, at one level, unsatisfying, taking us back toward an alignment condition that seems arbitrarily biased toward one kind of intelligence. The potential saving grace comes from the previous section: the possibility of one Goldilocks band of computational capacity sufficient to develop recursively self-improving AI, so that it really should be true that any society confronting this lesson has similar computational capacity. It then becomes conceivable to deduce what sophistication of individuals supports that capacity, considering that natural evolution runs slowly, so we’ll tend to get the simplest individuals that are up to the job. The simpler individuals couldn’t team up to solve alignment, and the significantly more sophisticated individuals don’t appear until the intelligence explosion is underway and the rules have changed.
This approach offers pretty elegant protection against the outcome of ignoring individuality at the level where we’re used to considering it. For instance, while corporations work effectively toward a variety of goals and can be said to exhibit higher-order thinking, many commentators would see it as dystopian to identify corporations as the primary agents! What saves us is that corporations are leaky abstractions, e.g. a CEO typically needs to work very hard to repeat company strategy over and over again, in a pithy enough way, and yet still finds inconsistent execution across teams. Perhaps the individuals we really intended by that word turn out to be precisely the sufficiently well-integrated, internally aligned computers, so now we only need to figure out how to describe that concept.
I’ve also bumped into a related (and controversial) concept that I haven’t studied in-depth yet: integrated information theory (IIT), which characterizes mathematically how pieces come together into conscious wholes. I’ve written elsewhere about consciousness and why I’m not choosing to refer to that concept. In the present article, I’ve been dealing just with the power of different systems to compute good strategies for alignment, though perhaps some useful mathematics could turn out to be in common with IIT.
Conclusion
All of the above suggests the following alignment principle, which probably wouldn’t seem at all suitable on first inspection, but perhaps I’ve explained why it’s promising.
Generalized alignment should identify individuals by their capacities for well-integrated, internally well-aligned computation past a certain scale, always choosing the maximal such units in a neighborhood.
The maximality part is there so that we don’t do the equivalent of choosing cells, not people. Though the computational-capacity minimum should be chosen to exclude cells, we wouldn’t want an accident like identifying a person’s two brain hemispheres as separate agents.
Naturally, there would be plenty of work left to do in fleshing out all of those concepts into computer code or formal logic, in addition to doing the same for whatever decisions should follow the identification of agents. For instance, there may be good statistical characterizations of well-integrated intelligence that seem obviously compelling after-the-fact, but maybe not. The task seems easier than for somewhat open-ended formalization of human values, with ourselves as an evolutionarily contingent legacy system that is hard to reverse-engineer as a black box. A powerful intelligence might reason from first principles about what intelligence at different computational scales should be expected to want, or, if you’ll allow me to sneak in a second Star Trek reference, it might follow a kind of Prime Directive of noninterference (or other rules of engagement) with qualifying intelligence, as measured by how the AI’s actions perturb statistical patterns in the agents to focus on.
Some interesting misfires of this rule look possible. On the one hand, we worry about false negatives like not identifying children whose brains haven’t developed enough yet, which motivates ideas like setting the computational lower bound conservatively enough or perhaps adding an element of predictable development of an agent to pass the computation threshold. On the other hand, we may have a false positive of identifying a frontier supercomputer as an individual, though perhaps again forecasting where such intelligence is headed reveals a good destination. The next article will partly address that question, considering how there might be some amount of inherent convergence of intelligence toward certain laudable goals, regardless of the alignment starting point.