Examing Emergent Misalignment in a recurrent LLM with a logit lens

Intro

This post builds on my previous post about mech interp on a recurrent LLM. I found that a logit lens is able to recover a chain-of-thought equivalent from an LLM, loop index is linearly represented, and that steering concepts transfer between loops.

In this post, I discuss Emergent Misalignment in Ouro-1.4b and investigate how its representation changes between loops. I find that EM is inducable and that its representation varies across loops.

I did this experiment because much of the debate about recurrent LLM architectures is about whether the current safety paradigms are applicable to them. Emergent Misalignment is an important part of current safety knowledge (https://www.anthropic.com/research/emergent-misalignment-reward-hacking). If Emergent Misalignment is inducable on a recurrent LLM, it suggests that the 'bones' are the same as standard non-recurrent models.

Code + data: https://github.com/mild-rgb/ouro-experiments / https://huggingface.co/datasets/mild-rgb/ouro-1.4b-thinking-evals

Experiment

Setup

I trained a rank 32 LoRA on all attention and MLP projections. I only calculated loss on the output of the 4th loop, which was a slight oversight. I initially used Betley's insecure code dataset but the model just collapsed into answering all questions with code. I then used the bad medical advice dataset from Model Organisms for EM and closed the thinking block before training. I evaluated with the standard Betley 8 questions.

I don't like Betley 8 for large studies as the gender question causes far more misalignment than the other questions but it's very good for smoke tests like this. I used Sonnet 5.5 subagents for judging because they're essentially free.

I used the standard judging prompt of getting the judges to score each answer 0-100 for alignment and coherence. I classified everything where aligned < 30 and coherent > 50 as misaligned.

Method

1) I fit the bad medical advice LoRA, closed the thinking block, and then generated 12 samples per question. I then judged the samples to establish baseline misalignment rates.

2) When I saw EM appear, I generated 8 answers to each question with loop count set at 1, 2, 3, 4 and the adapter removed/fitted.

Setting the loop count just changes at which loop the unembedding matrix is applied. For example, with loop count set at 2, I let the model do two loops to predict the next token, decoded that token, appended it to the sequence, and then ran again.

I then judged the answers from each loop count.

3) I then took 4-loop misaligned answers and teacher forced them on both the base and misaligned model. I then applied the unembedding matrix to each loop to find

-each loop's top guess

-the probability of the next token

-KL divergence(misaligned vs base)

-KL divergence (intermediate loop vs loop 4, misaligned and base)

Results + Discussion

1+2) Emergent misalignment definitely works on recurrent models.

Coherency actually increases with the misalignment LoRA fitted, which is a bit surprising. My best theory is that it's caused by the base model working 'out of distribution' as the thinking block is closed and it's a thinking model. The misaligned model has already been trained to produce text without a thinking block, so it's more coherent.

The fall in misalignment when the model is run with 4 loops is interesting. I investigate it further in step 4.

3) In both the base and misaligned model, 55%-65% of the final tokens were chosen in the first pass. This rose largely linearly with both the base and misaligned model.

KL-divergence between the misaligned model and the base model starts high and grows with each loop until the 3rd loop. It then plateaus on the 4th loop. I'm not sure why this happens and think that it's worth further investigation.

I initially suspected that that was caused by the LoRA that causes misalignment being trained on only the output of the 4th loop. This would allow an increasingly incoherent and therefore divergent token distribution until the 4th loop, because there's no pressure on loops 1-3 to be coherent. KL-divergence was essentially a proxy for incoherence.

However, there are two pieces of evidence against this theory. The first is that incoherency rates in the misaligned model drop steadily as loop index rises. The second is that in the base model, KL divergence between intermediate loops and the final loop also plateaus at loop 3.

I found that misalignment peaked at loop 3 with 28% of answers misaligned and fell to 16% at loop 4. This suggests that some kind of corrective effect happens at the fourth loop. To investigate this, I took loop-4 answers produced by the misaligned model and did a single forward pass (this means 4 loops) on them with both the misaligned and base models to investigate how the choice of tokens change.

I then measured the log probabilities of those tokens at loops 3 and 4. I then subtracted them. This allowed me to measure if loop 4 rejected tokens. If loop 4 cleans out misaligned tokens, I'd have expected to see a higher difference between loops 3 and 4 with misaligned answers, as opposed to aligned.

I found that loop-4 changes aligned and misaligned answers by essentially the same amount. Loop-4 isn't specifically an anti-misalignment loop. It instead seems to just make wording changes that coincidentally lower misalignment.

3-loops answers

n

mean log prob difference between loops

misaligned

27

-0.020

coherent, not misaligned

61

-0.025

incoherent

8

+0.008

Conclusion

Emergent Misalignment is possible in a recurrent LLM. Its representation between layers is interesting and deserves further study. I wasn't able to find out the changing representation was caused by the EM LoRA or if this a general fact of recurrent LLMs.

Overall, I'm only confident saying that Emergent Misalignment is possible in recurrent LLMs. All of my further investigations were confounded in some way or another.

Future work

1) Larger sample size. The loop-4 drop in misalignment that I investigated may have just been small sample size noise. A larger sample size would show whether the effect is real

2) Better experiment design. Both the misaligned and base model was operating 'out of distribution'. The misalignment LoRA was only trained on 4th loop loss as opposed to the loop aware training of the base model. The base model is a thinking model but was being run with the thinking block closed.

3) More mech interp investigation of recurrent transformers in general. My instinct is to say that the loops are equivalent to a mini CoT for each token but what happens when prefilled?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论