LLM Agent Swarms Are Easy Mode

This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I draw attention to the fact that recent LLM-powered cyberattacks have been relatively easy to understand, because today’s AI agents leave reasoning trails that are relatively familiar from human organized crime. Ongoing optimization should be expected to produce AI systems that are impossible for us to understand, but we actually have an opportunity to improve safety if those systems are also amenable to rigorous safety proof.

Though there have been predictions for quite some time about misbehaving AIs causing trouble, for many observers the risk became concrete recently, with first OpenAI agents hacking into Hugging Face and then, just last week, disclosure of agents doing unauthorized probing of U.S. government agencies. The possibilities for near-term danger seem pretty significant, but I’m writing to make the case that we should expect things to get much worse, if we just stay on the mainstream trajectory of AI research.

Why are the autonomous cyberattacks in the news today playing on easy mode? They are carried out by swarms of agents based on large language models (LLMs), which are trained on many examples of human-generated language, with the handy side effect that most of their reasoning feels quite familiar to us. Experts can come along after an attack and carry out a pretty effective postmortem, by reading logs of reasoning in English that wouldn’t be so out-of-place in a criticism assignment in a humanities class. The trouble is that we already have complex software systems, like operating systems, that exhibit emergent complex behaviors that surprise even their creators, requiring complex debugging processes to fix, full of what is essentially design of scientific experiments and analysis of the data that come out. There is no guarantee that log files record all relevant information, and indeed much of the debugging process comes from deciding to log new information.

So that complexity exists already for human-engineered software, and we should expect recursive self-improvement (RSI) to produce even more ingenious, complex software, even further beyond our ability to understand. It’s very far from clear that a postmortem could make sense of an attack by such a system. We need to plan ahead and constrain the behavior of unfathomable but useful optimization processes – processes of searching for better and better designs.

Let me step back and review thinking on how we should expect future AI systems to be relatively incomprehensible to us, connecting to the observation that we are in a lucky moment of AI development where relatively powerful systems still “think” (or pretend to!) in ways relatively comprehensible to us. It may even seem like we should rejoice that LLMs generate diaries meaningful to nonexpert humans, leaning into that style of AI development, but I’ll argue that the surface-level appeal is deceiving. Unsurprisingly given the general content of this blog, I’ll make a different case, for formal verification of powerful code generators.

Strange Alien Minds

Yudkowsky made the case in 2008 that there are many ways to organize computation to implement intelligence, and there is no particular reason that we should expect the best design strategies to look especially like our own brains. He also proposed a way to unify our thinking about evolution and deliberate engineering, where evolution involves an underlying optimization process for optimizing organisms for better survival, etc., such that the optimization process is well-separated from the organisms it optimizes; but powerful AI may be able to understand its own optimization process well enough to optimize it further, potentially triggering an intelligence explosion, with exponential growth in optimization effectiveness at a whole new level.

I already gave the example above of operating systems, where they accumulate such complexity that even their human software-engineer authors don’t understand all of their emergent behaviors. Another good example from current experience is how optimizing compilers transform programs into faster versions. Most programmers are not capable of understanding the, say, assembly language that is output – even as it is derived in a mechanical way from those programmers’ own code. Even elite programmers have to sink large amounts of time into understanding long compiler-produced assembly programs, when they might comprehend source code instantly. Now imagine the optimizing compiler optimizing itself repeatedly, or some other case of RSI, and we can see code quickly moving beyond the ability of humans to understand it or monitor its execution for safety. Fundamentally, the ideal implementation of a program, with respect to some optimization goal, may be wildly different from the best version a human programmer came up with so far. The possibility is analogous to how many triumphs of human engineering are very different from the closest analogues in nature (e.g. airplanes vs. birds), which came about thanks to the comparatively “dumb” process of natural evolution. Deliberate engineering should always win out eventually, and it should win out quickly after effective RSI takes off. (By the way, relevantly to the solution style I will get to proposing, both kinds of software artifacts covered in this paragraph have been connected to formal verification with spectacular success, supporting guarantees without human understanding of code.)

Now we come to the recent generative-AI explosion, which succeeded by training LLMs on many examples of text written by humans. It turns out that we mix enough of our knowledge into writing that this training process is efficient to transfer much of that knowledge into AI systems. One important kind of writing lays out logical arguments in favor of conclusions, hence training on all writing also sets up LLMs to carry out chains of reasoning, including to plan cyberattacks (even if the training set contained relatively few extensive reflections on planning cyberattacks). An almost-accidental consequence of the training method is that these reasoning traces are understandable to humans, even those who aren’t especially knowledgeable about AI or even computer science broadly.

Hence METR can analyze the OpenAI-HuggingFace cyberattack after-the-fact largely through reading logs of agents’ reasoning in English. What an amazing advantage for cyberdefenders! Imagine if all criminals reliably left diaries recording how they planned their crimes. Or, more analogously, imagine if every organized-crime group managed to maintain a culture of all participants keeping diaries (free of secret code words or other obfuscation) and periodically depositing them in a central library. Yes, swarms of LLM-powered agents may move more quickly and expertly than cost-effective human teams, but this advantage to understand their misbehaviors is a big deal. Now, to be fair, the METR analysis depended on sifting through epic amounts of these reasoning logs. What would previously have required teams of associates at a law firm to sift through documents was instead performed with yet more AI agents to search for juicy content, adding a risk of correlated misalignment. There was also evidence that some logs had indeed been tampered with, and messages from a shared message board really did use cryptic abbreviations. However, all of the above we can imagine being relevant to a conventional investigation of a large-scale criminal conspiracy! AIs that are truly not human-like could produce fundamentally different evidence trails.

OpenAI researchers have already speculated on how this kind of monitorability may come less easily as models improve. For instance, reasoning in latent spaces is a technique to avoid materializing reasoning as natural language. In general, we should expect that deliberate optimization, backed by lots of computational power and given plenty of time, will eventually beat out natural evolution. LLM agent swarms are generally following the model of humans and our cooperative social arrangements, which I would find astronomically improbable to be nearly the most efficient method for planning cyberattacks or carrying out any other kind of knowledge work. The actual champion methods shouldn’t be expected to generate an accidental byproduct of explanations that are easy for people to understand.

Our discussion so far may tempt some readers to propose that we freeze this characteristic of LLMs for as long as we can. Indeed, this property has been called out as chain-of-thought monitorability and promoted as important to retain through future generations of models. I don’t think that strategy should be expected to deliver great results, and I’m going to suggest a different one.

Provable Limits on RSI

Given how well AI coding agents work today, it now seems inevitable that development of whatever comes next will start with intensive use of deep learning, even if this style has inherent disadvantages in performance and reliability. However, the evolution of tools and approaches doesn’t need to stop here. Let me give a few arguments for moving away from LLM-powered agents.

We already don’t expect to be able to control the behavior of people very well! It doesn’t sound like a recipe for success to perpetuate this unreliable substrate in the digital world. Human society depends on dynamics like deterrence and reputation, which assume aspects of human societies that may not be preserved in a world of digital agents. For instance, even if AI agents “think” like humans but decide and act dramatically more quickly, they can do a lot more damage before we notice something is off and take appropriate corrective action. The next swarm attack may not leave our digital infrastructure in a state sufficient to implement future workarounds. It may also leave behind sneakier reasoning logs that hide true intentions.

It seems like hubris to expect that the most effective decision-making systems will look much like cooperating humans. This argument is a plain appeal to the engineering mindset, to seek out the solution that meets the requirements best, even if that solution looks dramatically different from what came before. I wrote previously about how thinking on AI alignment chooses humans as the distinguished beneficiary of systems to promote safety, even as general progress in science has led us to see ourselves as decreasingly central to the story of the universe. Maybe we are the roughly unique kind of intelligence that is smart enough to begin an intelligence explosion but simple enough for evolution to find first. Will better-optimized replacements even be described well with the word “swarm”? The idea of “individual” behind “swarms” may be an accident of what kind of highly coherent, internally aligned intelligence evolution was able to produce. I would bet on effective intelligence of the future looking to us more like undifferentiated soup of computation and storage elements – one integrated computer – without recognizable “individuals” within it that set their own goals and compete with other parts. Competitive pressure will push toward adopting these solutions, which should be much more effective than those that evolution happens to have found.

There is a great opportunity behind RSI with provable guarantees. We feel like we understand our fellow humans at some intuitive level, but they still surprise us regularly. In contrast, the “strange alien minds” that advanced optimization produces may confound our intuitions while still being compatible with mathematically provable guarantees about behavior, if we kick off their generation properly. (I’ll just warn you here that I’m proposing formulating those mathematical properties as an important research direction, not delivering a blueprint in this article! My position is that we should find ways to provide AI systems with limited, mathematically well-defined interfaces, allowing important properties to be made precise.)

It’s a somewhat counterintuitive situation. At first glance, it seems like a great advantage to have powerful AI “thinking like us,” so we can e.g. audit their reasoning logs. However, the much stronger safety advantage, to rule out unsafe actions in advance, may only come from mathematically provable evolution of code. It applies even to code that we have no hope of ever understanding on our own. So there is the paradox: systems that seem harder for us to understand at first may become the most amenable to strong guarantees, if we give them the right structure from early in their evolution. The style of system that I proposed in this mold last week is autonomous evidence factories, which configure fixed hardware resources to solve well-specified mathematical challenges efficiently.

Conclusion

Generative AI with deep learning won the hardware lottery: it could be implemented with good performance on GPUs, which were originally designed for computer graphics. No doubt associated tools will be the foundation on top of which the next powerful AI systems are built. However, I’ve made the case that, while LLMs offer particular advantages for AI safety, we should aim to do better with very different architectures. Specifically, we should ask for provable guarantees about how self-modifying systems evolve. I’m going to turn next to some of the established ideas in programming languages, compilers, and formal methods that can form the foundation for that alternative path, where we can hope to prove strong properties even of programs that we can never understand.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论