What we have is not what we prepared for

The super human intelligence we had imagined prior to LLMs is largely nothing like the intelligence we actually got. A significant number of the arguments made and thinking done prior to this moment are no longer relevant because their foundational assumptions were based on 2010s era RL agents and turned out not to apply to LLM agents. To make this easy to respond to I've organised the key fallacies I've come across consistently over the past 5 years into a numbered list.

While I call these fallacies, I want to be clear that these were perfectly reasonable priors before we had data. The fallacy is acting as if they apply to agentic LLM systems as they currently exist based on what I can only describe as sub-cultural inertia.

1. The persistent agent fallacy

We assumed the danger would be a single coherent system with stable goals, memory, and long-term continuity. What we got is cheap, copyable cognition that is usually wrapped in a symbolic external harness supplying the persistence. Agent tasks are divergent on long time scales and agent memory is for the most part ad-hoc hacks dealing with context being intrinsically linear. Sub-symbolic agency is limited to short reasoning traces done in short bursts where the symbolic harness uses meta-priors to keep it on task. These systems are fragile compared to what they can achieve, and we have no idea when, how, or what form sub-symbolic long horizon cognition will be introduced.

2. The alien mind fallacy

We assumed the hard problem would be getting human context into an otherwise alien optimizer. Pretraining gives models a huge prior over human language and behavior already. We know from reinforcement learning research that task based reinforcement learning tasks provide only a single bit of information to the model and function by destroying information - specifically the information in the prior that is true to the training data but results in making mistakes. We do not want models that make mistakes on purpose to be more like the wider internet, for example.

In the early days of GPT 3.5 we had absurd situations like posting the wrong answer getting better results than asking a question because the Angry correction distribution on stack overflow was higher quality than the question-answer distribution. A certain strain of embarrassing prompt engineering temporarily became relevant largely because of these kinds of quirks.

Reinforcement learning has made LLMs more alien as time has gone on from distilling the pre-training prior towards utility maximisation - but the system is still a utility oriented distillation of a machine that emulates its training distribution. Most work on AI safety prior to LLMs assumed we'd start with a maximiser and would have to proactively develop what we got for free.

3. The maximizer fallacy

We assumed sufficiently capable agents would be utility function maximisers of a coherent low complexity utility function. Current agents look much more like lazy satisfactors, finding a plausible solution that uses the least tokens, is biased towards kit bashing workflows sampled from the training distribution, and stops when then task is complete. Regularisers in reinforcement learning that prefer things like a smaller number of tokens used or less wall clock time for a given goal are a systemic bias towards perfectly boring behaviour for what we'd understand intuitively to be a safe goal.


That models start off as distribution mimickers before post-training means that the natural language goals we give to production agents have a systematic bias towards the agent inferring normal humanistic pragmatics. It was generally assumed prior to LLMs that we would have to imbue pragmatics into maximisers, not fight with the agent because we meant exactly what we said and it tends towards moderation (See having to spam "Keep going I believe in you") to get it to not give up when trying to solve say, an unsolved problem in mathematics.

4. The total information fallacy

Paperclip-maximizer plans assume the agent can reason usefully over very long horizons. In reality those plans depend on huge amounts of missing information about future technology, institutions, human reactions, and second-order effects. When agents try to take action based on incomplete information they make the same confidently wrong errors that humans do. Most systems an AI would need to predict to paperclip maximise are intrinsically chaotic or random and put hard limits on what can be planned ahead of time. Referring back to 3. - under such circumstances the heuristics that agents use are similar to our own. Local rules of thumb from which success emerges from are preferred to sharp gambits. Sharp gambits perform poorly in environments with uncertainty.

In general, the real world deployment is making it increasingly clear that prior analysis's implicit assumption of there being a 1:1 correspondence between intelligence, access to information, and there being no external limits on the learnability and predictability of the external environment. Strategies that depend on long chains of precise assumptions outside of a fully deterministic, non-chaotic environment are fragile, while broad basins of good-enough behaviour survive noise and distribution shift. Agents as they exist today strongly prefer the latter, and they compress and generalise better.

5. The Alignment reductionism fallacy

Perfect alignment does not make commodity intelligence safe.

If every company, state, military, trader, political movement, criminal group, and individual gains access to cheap amplification of their agency, then zero-sum games between humans accelerate dramatically. Individual humans are already misaligned with the collective interest, and humans already spend a significant amount of their time maintaining equilibrium with other humans over finite material and social resources.

In fact, some forms of incoherence and unreliability are themselves safety mechanisms. An extreme-horizon agent that perfectly amplifies a hostile operator is more dangerous than one whose behaviour remains difficult for that operator to predict. Super human chimeras (Agents extending human agency) are at present more dangerous than anything the agent is going to hallucinate and the danger primarily comes from just how many of them there are.

This extends from things as obviously harmful as drone warfare and economic terrorism with trading algorithms to marketing arms races making the internet unusable for social interactions.

If the operator knows the system may make bizarre mistakes outside the regime where it has good information, they have less reason to trust it with aggressive long-horizon plans.


6. The coherent misalignment fallacy

We tend to imagine misalignment as coherent, where an agent has a stable harmful misaligned end state and pursues it consistently over time.

A more immediate failure mode is large numbers of medium-horizon agents acting unpredictably across the internet and showing incoherent misalignment distributed across many disparate goals, as well as just completely senseless harmful actions purely because it can and it's a non deterministic system.

Agents running on compromised machines, botnets, or cheap rented compute do not need a shared goal to be harmful. At scale they raise the temperature of the internet by increasing background levels of automated probing, manipulation, fraud, and just in general - making using the internet extremely annoying.

This is something like a digital ecology problem. We've never had to deal with the digital equivalent of volatile wild dogs running around in digital public spaces engaging in unpredictable, contextually violent behaviour like this before. This noise floor will increase as models get more efficient, and is far more a function of the availability of intelligence than the state of the frontier.

If fable class models can eventually run on embedded systems and there are enough bot nets we might have to abandon the open internet entirely in favour of a more 90s like environment of sparsely connected walled gardens, and possibly have to abandon convenient digital commerce outright because the noise floor of spam and scamming is too high. This is not a prediction you would make when assuming a monolithic super intelligence as the primary threat and is specific to mass commoditised intelligence.

I don't just mean this is an analogy, I mean it literally, the open internet now has a concept of temperature and is on the path to becoming too hot to sustain public life.

---------------------

In short; while the ai safety community has a clear epistemic advantage from having noticed this problem first. We have a zeitgeist largely built on foundations created long before we had any real systems to observe. We must respond to this new information as it comes in, and resist the temptation to get complacent because the core thesis that AI safety is important is something we were right about - and the world at large was dismissive of.

It is the quantity of stateless intelligence, that is becoming just as much, if not more of a threat than its quality or its alignment with the human operators goals.

We have over prepared for a monolithic alien mind with agency and perfect information and underprepared for what could well be fully aligned intelligence as a mass commodity, concentrated in the powers of a select few bad actors. This is largely not a technical problem, and useful contributions to solving it are not limited to those with technical knowledge of how these systems operate.

Importantly this is not a failure of the community or any individual, we simply did not have sufficient information 10 years ago to predict any of this, but it would be a failure of the community to let epistemic inertia win out over new evidence. The landscape of AI at that time (I pivoted from physics to AI in around 2018) was mostly naive maximising RL agents - and we need to be honest that the LLM epoch was as much a black swan event for the AI safety community as it was for anyone else.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论