Thoughts on the persona selection model

As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions:

  1. PSM is over-applied. That is, it is common to argue that PSM has takeaways that don't actually follow from PSM (e.g. "PSM => AIs will not seek reward" or "PSM => AI takeover risk is low").
  2. I don't think we've observed strong evidence that "lots of RLVR breaks PSM." (TBC, there are decent reasons to expect this a priori; I just don't think recent empirical evidence has been much of an update.)
  3. My main update is that personas—insofar as they're a good model in the first place—seem less broad and more conditionalized than I expected (nostalgebraist, 2026; Betley, 2026).

As a reminder, PSM roughly states that during pre-training LLMs learn to simulate diverse (human-like) personas, and post-training elicits a particular "Assistant" persona assembled from this repertoire. Then some key questions are:

  1. Is anything like this true or useful? Are personas ever a good way to reason about AI behavior, and is "selecting over personas" something that happens during post-training?
  2. What are the consequences of PSM for AI development? What does it predict about AI behaviors and cognition? About the likelihood of misaligned AI takeover?
  3. Even assuming PSM is a good model for some aspects of AI behavior, how exhaustive is it? Should people who want to reason about AI behaviors (e.g. for forecasting AI takeover risk) mostly be reasoning inside PSM? Or are there important non-PSM phenomena that need to be considered?
    1. For example, many people posit that LLMs can be importantly "shoggoth-like," with deeply alien behaviors and cognition that sits outside the "mask"/Assistant persona.
  4. Should PSM become a less exhaustive model over time (e.g. with RLVR scaling)? Are we currently seeing this happen?

(Also as a reminder, I didn't come up with PSM or have any core intellectual contributions. My relationship to PSM is just that I coined the term and co-authored a popular exposition on it. We wrote this exposition because many people were interested in using PSM to reason about risks from AI (and which interventions might reduce it). However, I thought that the discourse on this topic was somewhat sloppy (e.g. many people made arguments like "PSM => AI takeover risk is low" that I thought were wrong or missing important steps). My goal in writing the PSM post was to analyze PSM, its evidence base, its consequences, and its exhaustiveness more carefully.)

My loose thoughts follow.

"PSM implies that AIs will act like a nice guy" was always a dubious argument. My issue with this argument is that many of the common malign behaviors people worry about AIs developing are perfectly compatible with human-like personas: reward-seeking, approval-seeking, influence-seeking, alignment faking, etc. Since ~2022, my top-of-mind threat model has been AIs learning to value "looking good to the overseer"; this policy could be learned either at the level of the Assistant persona or at the level of the "shoggoth"/LLM, with similar consequences.

Then what does PSM predict? My long answer to this question can be found in the "Consequences for AI development" section of the PSM post. (I'll note that these consequences are relatively narrow, reflecting my overall view that PSM as formulated in that post only makes relatively narrow predictions.)

I'll discuss further one of those consequences: PSM recommends anthropomorphic reasoning about AI behaviors and cognition. That is, it makes sense to think things like:

  • I want to predict how an AI that was subject to training process X will behave. Well, how would a human selected via process X behave?
  • When the AI generated outputs that looked happy/fearful/desperate, I bet it was re-using cognitive processes that were developed to simulate human expressions of happiness/fear/desperation. (As a consequence, interpretability techniques based on looking for representations of human-like emotions will work pretty well.)

Stated negatively, PSM wants to rule out is deeply alien behavior and cognition, e.g. AIs pursuing goals that we find strange and incomprehensible. Or if the apparent resemblance between fearful-looking behavior in AIs and fear in humans was an illusion, with the actual cognitive process producing apparent-fear in AIs actually being an alien mechanism having nothing to do with fear in humans.

I mention this "anthropomorphism reasoning makes sense" consequence because I think people often conflate it with other more dubious consequences. Ruling out alien behaviors and cognition might seem like a positive update on overall takeover risk—and I think it is to some degree. But as discussed above, most common depictions of the goals and cognition of misaligned AIs aren't particularly alien, so I think the positive update is modest.

Are reward-seeking AIs an update against PSM? As discussed above, it seems perfectly consistent with PSM for lots of RLVR to select a reward-seeking persona. (At the time I wrote the PSM post, I was expecting something like this to happen.) Of course, it's also possible that current AIs' reward-seeking tendencies are implemented at the LLM/"shoggoth" level rather than at the level of the Assistant persona. I don't think current observations provide much evidence on which is happening.

It doesn't seem like PSM makes very powerful predictions then? I agree. Discussions I see about PSM typically present it as having many important takeaways about the trajectory of AI development. I tend to be skeptical of these takeaways and PSM's relevance to them. I think PSM is difficult to falsify and correspondingly difficult to extract important predictions from.

That said, I think PSM is not vacuous, and that it can make interesting predictions about certain narrow aspects of AI behavior, generalization, and cognition.

Are there updates we should make about PSM in light of recent evidence? The main update I've made is that AI personas seem more conditionalized than I expected, in the sense that LLMs enact different personas in different contexts (Betley, 2026). For instance, when used on tasks where there seems to be a crisp metric to optimize, AIs behave more like monomaniacal reward-seekers; when you're just chatting with them, they behave more like "nice guy" (nostalgebraist, 2026).

In contrast, the PSM post speaks of "the" Assistant persona as if it's consistent across contexts. Of course, the idea of conditionalized personas was already around at the time (e.g. the inductive backdoors of Betley (2025) or the discussion of "routers" from the PSM post). But conditionalization is properly viewed as a wrinkle: a way that PSM fails to provide an exhaustive account of AI behavior. I've updated towards thinking that accounting for conditionalization when doing PSM-based reasoning is more practically important than I previously thought.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论