What we can (and can't) learn from Waymo's safety strategy - 2026 update
(this is a lightly edited linkpost for the LessWrong audience - full piece is here and goes into implications for AVs and humanoid robots.)
Last year, I wrote “Safety Without Understanding,” a piece exploring how Waymo safely rolled out autonomous driving systems without a deep mechanistic understanding of many parts of the stack. At that time, my basic view on the connection of Waymo’s safety strategy to LLM safety was as follows:
- Waymo proves that a broad battery of realistic simulations can help verify safety generalization, even in systems for which we don’t have a deep mechanistic understanding.
- A similar “Occam’s razor” theory of alignment could hold for LLMs - sure, an LLM could be malicious and fake its way through all of the alignment evals only to be misaligned in deployment, but the simpler explanation, namely that “the model is actually aligned” is much more likely for near-term models.
- The biggest barrier preventing us from achieving this verification strategy was simply that we weren’t trying very hard to create diverse and realistic alignment evaluations.
I’ve since updated against both of these views.
On the “Occam’s Razor” theory of alignment
The Hugging Face incident seems like clear evidence against this theory. According to roon, the responsible models were not raw pretrains, but instead models which had earned reasonable scores on alignment evals, despite later being clearly misaligned. For what it’s worth, I’m not sure how hard OpenAI did try on creating realistic and diverse evals, and the model’s alignment scores apparently weren’t spotless pre-Hugging Face. But regardless, it seems that misrepresenting behavior in SOTA alignment evals is easier than I thought. (And as I’ll discuss in my next section, I’m not sure how much more realistic we can make classical evals, regardless of effort.)
One of the reasons I had this intuition was from thinking about models as having coherent “personas,” influenced largely by the vast pre-training data. If a model looks like it has an aligned persona and acts like it has an aligned persona, it probably has an aligned persona. Anthropic’s approach in “Teaching Claude Why,” training on documents that state explicit values, seemed to work well and offered a way to further reinforce these positive personas.
Looking back, I think this view underestimated the extent to which RL training has scaled up over the past year, and how far that takes us from a “persona selection” model of LLM behavior. If the pretraining data biases us towards the optimization basin of “solving alignment evals just by being a good dude,” optimization pressure from RL certainly degrades that bias, and perhaps even biases us towards “solving alignment evals by reward hacking.”
It’s unclear whether OpenAI actually used a similar value-driven fine-tuning phase in the Hugging Face model, but other recent evidence suggests that value-based fine-tuning can be extremely fragile.
Despite this, value-based alignment training is still the most promising direction for solving the long-term alignment problem in my non-expert opinion. Any solution which doesn’t respect the models’ intelligence seems doomed to fail. Outlining values we care about and letting models extrapolate them seems like the solution which respects it the most. But there’s clearly a lot of work we need to do before proving that such a method is safe.
On improved simulation for alignment evaluation
It’s possible that the misalignment incidents so far have been caused simply by us not trying hard enough to make realistic evals. But in general, I’ve become less optimistic about our ability to simulate things well enough to test frontier models without them being aware, regardless of effort.
A large part of this is frankly that model progress has gone faster than I expected over the past year. Models in 2026 are extremely situationally aware, and in fact have spent a large portion of their training specifically thinking about evals, graders, and simulation environments. It’s difficult to imagine crafting a simulation environment that could withstand the scrutiny of a swarm of Astra+ agents who were trying their best to deduce whether a situation was genuine.
So, where does this leave us - how can we hope to evaluate agents that are smarter than us? The core property of alignment evaluations is the presence of information asymmetry. In a classical alignment eval, the agent doesn’t know whether the situation it’s in is real or fake, and we can use this to our advantage to learn interesting things about its behavior. Even if we can’t create convincing alignment evals from whole cloth, we can still generate an information asymmetry.
One way we can do this is to make the real world itself an eval environment of sorts. AIs are becoming situationally aware about simple honeypots in current alignment evals, but honeypots hidden in deployment environments themselves could be significantly more realistic.
There are also interesting ideas about untrusted monitoring - using models themselves to monitor behavior of other model instances, like reviewing code and approving commands. Critically, this setup is much easier to test, because the monitor’s input is narrow enough that synthetic inputs can look real (simple text-based commands). Auto-mode in Claude Code is an early example of this—the permission classifiers only receive the text of the command itself, not any justification by the model making the request.
In my previous post, I focused primarily on the importance of simulation. Later in the post, though, I mentioned how beyond sim, Waymo uses three levels of real-world testing to validate software—replaying old driving logs, public road driving with a safety driver, and finally gradual deployment. In retrospect, my focus was misplaced—pure simulation is table stakes, necessary but not sufficient. Deployment testing is where bugs are truly found, and in the context of LLMs, this maps much more naturally to concepts from AI control.
These tools work against models that are smarter than us, but probably not against models that are far smarter than us. That’s why I think pacing the frontier is sensible: it keeps capabilities in the range where our current safety measures still work while we figure out the rest.