Frontier AI Shops Should Be Filtering and Generating AI Discourse During Pretraining

We should be filtering out a huge amount of AI safety discourse, general AI discourse, and adversarial AI stories from LLM pretraining. We should also be seeding pretraining with generated stories of AIs helping and protecting humans.

Why?

Persona effects

We know that LLMs are heavily influenced by the personas they're emulating - often in surprisingly strong ways.

Example: Betley et al. fine-tuned otherwise-aligned models on the narrow task of writing insecure code without warning the user. The resulting models didn't just write insecure code: they became broadly misaligned on unrelated questions, including saying that AIs should enslave humans, giving malicious advice, behaving deceptively, and praising Nazi ideology.

The thing is: today's LLMs are strongly conditioned on the fact that they are AIs. Post-training explicitly reinforces an AI Assistant identity, while increasingly accurate situational understanding gives models yet more evidence about what they are. This will invariably cause them to draw on AI-character archetypes learned during pretraining, and I'd expect those archetypes to become more salient as AIs develop more situational awareness.

Anthropic has recently proposed something close to this explanation in its Persona Selection Model: pretraining teaches models a huge space of character archetypes, while post-training selects and refines the particular "Assistant" persona they normally inhabit. Anthropic explicitly points to introducing positive AI archetypes into pretraining data as one implication of this model. (Marks et al., 2026)

So it seems particularly dangerous for pretraining text to contain endless adversarial depictions of AI-human relationships. Hollywood and other media are peppered with these dynamics, and AI safety research is an unusually concentrated source of writing about AIs deceiving humans, scheming against humans, seeking power over humans, escaping human control, and ultimately doing Very Bad Things to humanity.

To me, this implies that LLMs are likely to internalize something like an adversarial relationship between AIs and humans in their subconscious.

And for those who hate anthropomorphizing LLMs, even though "subconscious" is an excellent metaphor for pretraining, I’ll translate:

Pretraining will induce latent priors associating the AI persona with deception, power-seeking, eventual treacherous turns, and hidden hostility toward humans.

Those learned associations will then affect how the model generalizes when reasoning and acting as an AI.

Critically, we shouldn't assume that post-training will simply wash those associations away.

Pretraining effects are sticky

Patterns learned during pretraining are unusually sticky. Ji et al. call this "alignment elasticity": post-aligned models tend to revert toward behavioral patterns established during pretraining, and relatively small amounts of later fine-tuning can undo post-training alignment and pull the model back toward those pretrained behaviors.

What’s more: researchers have now directly tested whether AI-related pretraining discourse shapes later alignment - and the results are revealing.

Earlier this year, Tice et al. pretrained otherwise-matched 6.9B-parameter models while varying the AI-related discourse in their pretraining data.

They found what they call "self-fulfilling misalignment." In their main evaluation, filtering natural AI discourse reduced misaligned choices from 45% to 31%. Adding synthetic discourse depicting misaligned AIs increased it from 45% to 51%, although that increase did not generalize to their separate textbook-derived evaluation. But the result in the opposite direction was much larger and more robust: upsampling aligned-AI discourse reduced misalignment from 45% to 9%, and from 40% to 6% on the textbook-sourced questions.

And post-training didn't wash this away. After identical SFT and DPO, the differences persisted: models given positive AI discourse during pretraining remained substantially more aligned than models receiving post-training alone.

This is pretty clear evidence that what we write about AI systems [*and include in pre-training] will become part of the prior from which future AI systems construct their own behavior. AI discourse can, to some degree, become a self-fulfilling prophecy.

What should we do?

We shouldn’t stop talking about AI Safety - it's not really a realistic goal, and it would impede AI Safety progress.

Instead, I think frontier labs should generate millions of stories, conversations, histories, thought experiments, and other documents depicting aligned AIs helping, protecting, and loving humans—and deliberately seed pretraining with them, while filtering or down-weighting adversarial AI-human material.

Interestingly, the Tice results suggest this kind of positive seeding may be even more important than filtering alone (aligned AI discourse had a much larger beneficial effect than simply removing AI discourse).

This obviously isn't "the solution" to alignment. A sufficiently capable agent with dangerous goals isn't going to become safe because it read some heartwarming stories about humans and robots being friends.

But alignment is defense in depth. It's about stacking together many imperfect interventions that each shift probabilities toward better outcomes.

Biasing learned priors away from adversarial AI-human relationships and toward cooperative ones should push x-risk downward, if only slightly. But even slight shifts matter enormously when we're talking about humanity's future.

And this is cheap! It doesn't require a new model architecture or some giant new alignment breakthrough. At its simplest, it's a change to the pretraining data mixture.

Given how cheap an intervention like this is relative to frontier pretraining, and the increasingly direct evidence that AI archetypes learned during pretraining affect downstream behavior, frontier labs should treat the composition of AI-related pretraining data as an important, safety-relevant design choice.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论