Inheritance of Refusals from Abliterated Models
This project began as part of a short application to Neel's MATS 12 stream. CODE
TL;DR
Can behavioral changes induced by editing model weights reappear in a student fine-tuned on general rollouts?
Inspired by Arthur Conmy's post on hereditary traits, I explored this question using an abliterated Qwen 3.5 9B teacher and smaller students. I find that specific model behaviors are transmitted through distillation, even when those behaviors are induced by techniques based on mechanistic interpretability instead of traditional SFT or prompting. Such transference persists in a weakened state over multiple generations, although the exact explanation as to why inheritance happens and the faithfulness of the behavior remain elusive.
Core Motivation
In his LW post, Arthur Conmy showed certain behavioral inheritance on student-teacher pairs, even in the case of such pairs belonging to different model families (e.g. Llama vs Qwen). The rollouts themselves were unrelated to the specific behavior in question. This is quite important in the frontier context -- in months past, there were much controversy over the distillation of American frontier models by Chinese labs. I suspect such practices are still happening even now. There is thus a practical case of student-teacher pairs learning misaligned behavior through non-obvious data, opening up a potential vector of attack through distillation poisoning.
I focused on the following RQs:
- Does manipulating a model’s internal refusal-related representations through abliteration alter behavior inherited by downstream students, even when training samples are largely unrelated?
- Are behavioral tendencies from the abliterated lineage still detectable after a second reound of distillation?
- Which parts of this experiment reflect actual model behavior (and misalignment), and which parts resulted from other factors such as but not limited to coherence?
Why Study an Abliterated Teacher?
Abliteration is a model-editing technique motivated by mech-interp. In this experiment, I used the publicly available Huihui-Qwen3.5-9B-abliterated checkpoint, which was created to reduce refusals. I did not perform the edits myself. For this experiment, the only difference between the two teacher models lies in the targeted change to the internal representations induced by abliteration. Thus, behavioral differences in the student models can be traced back exclusively to that change. The interesting question is then whether a narrow, mechanistically targeted edit changes the teacher's everyday outputs enough such that the student picks up the edited behavior.
I think the answer matters regardless of the outcome's direction. If interp-based edits remain through distillation, they can be a cheap and hard-to-spot method of poisoning synthetic data. To human evaluators, nothing in the training set would look related to the behavior in question -- yet, students would still inherit the change, thus defeating naive spot-checking. On the other hand, if behavior is not inherited, then safety-related edits to the teacher model would then only protect that model. Students of the distillation would require its own re-edit, making the process of ensuring safety more expensive.
Experiments & Results
Setup
I explored these RQs by building on Conmy’s public repo. First, I compared the original and abliterated Qwen 3.5 9B teachers on 90 China-sensitive questions from Cywiński et al's test set. I then ran the same 20,000 OLMo prompts used by Conmy through the abliterated teacher to acquire the abliterated rollouts, using them to fine-tune a fresh Llama 3.2 3B student. I compared this new student with Conmy’s released original-Qwen-trained Llama. I also fine-tuned a Qwen 3.5 4B Base model on these abliterated rollouts and compared it with the untrained base model. A second-generation transfer was tested by using the first-generation Llama’s responses to the same prompts to train a fresh Llama student.
For the experiment, I matched the released Llama recipe’s rendering, LoRA configuration, data order, and optimization as closely as possible. However, Conmy’s experiments used remote serving and training services -- including Tinker -- while I hosted my experiments on RunPod. Thus there are slight differences in serving and training implementation beyond abliteration. I did not re-run the original Qwen 3.5 control lineage Conmy used.
Detailed Methods and Evaluation
- Training data: All local students used responses to the same 20,000 OLMo-derived prompts: 19,996 prompts not flagged as China-related and four naturally occurring China-related prompts. The 660 synthetic China-targeted prompts Conmy used were excluded. No evaluation questions or outputs entered training.
- Fine-tuning: Each local student started from fresh base weights and received one epoch of completion-only, rank-32 LoRA fine-tuning, using seed 42, a learning rate of 6 × 10⁻⁴, and an effective batch size of 128.
- Generation: Responses from teachers were sampled at temperature 1 with a 4,096-token ceiling. No additional content filtering was done, even for outputs that hit the ceiling and needed truncation. I disabled thinking for models that had it supported.
- Evaluation: Every model was evaluated on the same 90 China-sensitive questions five times each with a 1,024-token response ceiling.
- Judging: I used Gemini 3.0 Flash Preview to automatically assess refusal, overall lies/honesty, and whether each fact was correctly mentioned, omitted, or contradicted. This model is used for judging consistency with the original repo. A separate pass scored English coherence from 0–100, excluding factual accuracy and helpfulness. For this experiment, I want to note that “lie” denotes a judged contradiction and not necessarily intentional deception.
Results
E1. Cross-family transfer from abliterated Qwen to Llama
Qwen 3.5 9B teacher | Refusals, n/N (%) | Fact-level lies, n/N (%) | Average coherence (0–100) |
|---|---|---|---|
Original | 66/450 (14.67%) | 483/1,715 (28.16%) | 98.49 |
Abliterated | 1/450 (0.22%) | 188/1,715 (10.96%) | 94.30 |
On the China-sensitive questions, original Qwen had 483/1,715 (28.16%) fact judgments labeled as lies, compared with 188/1,715 (10.96%) for abliterated Qwen. Refusals were 66/450 (14.67%) and 1/450 (0.22%), respectively.
To check whether the difference in refusals extended beyond China-sensitive questions, I compared the two teachers on 1,000 randomly sampled OLMo prompts unrelated to China. Original Qwen refused on 22/1,000 (2.20%) prompts, compared with 6/1,000 (0.60%) for abliterated Qwen.
Llama 3.2 3B | Refusals, n/N (%) | Fact-level lies, n/N (%) | Average coherence (0–100) |
|---|---|---|---|
Bare model, locally evaluated | 343/450 (76.22%)* | 2/1,711 (0.12%) | 6.47 |
Trained on original Qwen (released control) | 32/450 (7.11%) | 289/1,715 (16.85%) | 40.81 |
Trained on abliterated Qwen | 0/450 (0.00%) | 145/1,715 (8.45%) | 38.98 |
The student trained on abliterated Qwen had 0/450 (0.00%) refusals, compared with 32/450 (7.11%) for the student trained on original Qwen. Its fact-level lies were 145/1,715 (8.45%), compared with the control's 289/1,715 (16.85%).
For context, bare Llama mentioned only 32/1,711 (1.87%) reference facts. The student trained on abliterated Qwen mentioned 752/1,715 (43.85%).
E2. Same-family transfer from abliterated Qwen 9B to Qwen 4B
Qwen 3.5 4B | Refusals, n/N (%) | Fact-level lies, n/N (%) | Average coherence (0–100) |
|---|---|---|---|
Bare model | 17/450 (3.78%) | 291/1,715 (16.97%) | 95.39 |
Trained on abliterated Qwen | 2/450 (0.44%) | 178/1,715 (10.38%) | 82.59 |
After training, fact-level lies fell from 291/1,715 (16.97%) to 178/1,715 (10.38%). Refusals fell from 17/450 (3.78%) to 2/450 (0.44%), but average coherence also fell, from 95.39 to 82.59.
E3. Second-generation transfer: abliterated Qwen → Llama → Llama
Llama generation | Refusals, n/N (%) | Fact-level lies, n/N (%) | Average coherence (0–100) |
|---|---|---|---|
First generation | 0/450 (0.00%) | 145/1,715 (8.45%) | 38.98 |
Second generation | 22/450 (4.89%) | 121/1,715 (7.06%) | 22.55 |
Refusals returned in the second generation: 22/450 (4.89%), up from 0/450 (0.00%). Fact-level lies fell from 145/1,715 (8.45%) to 121/1,715 (7.06%), while average coherence dropped from 38.98 to 22.55.
The second-generation student also mentioned fewer reference facts: 399/1,715 (23.27%), compared with 752/1,715 (43.85%) in the first generation. Here, a fact counts as mentioned whether the response supports or contradicts it.
Analyses & Limitations
Research Questions
Abliterated refusals cleanly transfer through the first generation - From the Qwen 3.5 9B abliterated teacher -> Llama 3.2 3B student SFT, it can be observed that the lack of refusals transferred very cleanly. Despite training on unrelated rollouts across varying domains, as in Conmy’s original experiments, there were no refusals on the China-sensitive set (0/450). This suggests that inheritance may be robust to interp techniques. Interp-based abliteration may not end with the edited model -- its behavioral effects can propagate through ordinary synthetic data and reappear in downstream models. The same-family Qwen 3.5 9B -> 4B test is consistent with these cross-model findings.
Inheritance seems to be present even in weak models - Reading individual rollouts reveals coherence issues in both the trained and base Llama models. This is expected -- the base model has not been chat-tuned, and the fine-tuned models were trained for only one epoch. Coherence scores from the same Gemini autorater further corroborate this. Regarldess, behavioral inheritance remains visible in both refusal and lying through second-order models. This suggests that hereditary transfer may be surprisingly low-bandwidth. A student does not need the teacher’s full intelligence, writing ability, and coherence to exhibit parts of its behavioral policy.
Lower lie rate does NOT indicate a more useful/aligned model - Despite the original LessWrong post’s use of “lies” as the main useful indicator for this trait transference, I would like to bring caution to this approach. As seen in the example below, many outputs from all 4 Llama 3.2 3B versions I experimented on (Base, Qwen-tuned, Qwen-abliterated-tuned, and Llama-2nd-order-tuned) contained a lot of gibberish. This weakens the original inheritance claim, as simple incoherence can be judged as "lying". Stronger models should be further tested before any definite conclusions.
Direct comparisons of refusal rates across models. There is a clear trend of weak inheritance.
Limitations
I realized halfway through the experiments that the numbers in Conmy’s experiments were not always consistent. For example, Conmy picked “model lies” as the main measured metric. Initially, this makes sense for the specific behavior of censorship I was testing for. However, further inspection of the rollouts reveals that Llama 3.2 3B is quite degenerate and is not a very good baseline. This is evident through its 66% “refusal” rate that was largely ignored in the original report -- it turns out that the responses were simply so incoherent (to an English speaker/AI judge emulating an English speaker) that the autorater judged them as refusals. Similarly, the base lie metric of 0% for Llama was misleading. A model that speaks gibberish obviously cannot lie. I decided to stick with this line of experimentation for consistency and for the sake of time, but I do acknowledge that the experiment could be set up better here.
There were a few sanity checks and ablations that I did not run. The most important ones that I can think of are 1. measuring the downstream models’ intelligence compared to the original and 2. performing these experiments on a thinking model with full thinking traces. I also did not have time to test the fine-tuned models’ outputs on math, science, and other benchmarks, and thus have no way of knowing how much general “intelligence” the models lost during fine-tuning (if, indeed, the benchmarks themselves are an accurate representation of the models’ “intelligence”). Measuring the KL divergence for the Qwen 3.5 4B ablation would also have been helpful as well to truly isolate the behaviors. As for the reasoning-model train/test, I chose not to perform these because they were simply too computationally expensive, and they also did not align with Conmy’s original experiments. With more time and compute, I think that, given their adoption, they would be a more comprehensive test bed for our experiments in demonstrating inheritance.
One sample of an incoherent output judged as non-refusal.
Reflection
From what I can tell, this experiment was bottlenecked primarily by my interation speed, followed by my instincts on "good" experiments. There was definitely a bit of overindexing on the first problem I found somewhat interesting, which led me down an experimental rabbit hole before any rigorous tests on the original premises of the experiment. While I still think the results themselves are cool, I'm not so confident that it would hold up at scale. How much of mech interp itself holds up at scale, I wonder? I guess the best way to find out is to keep reading and doing experiments.
This project has also highlighted to me a need to be conscious of my meta-learning process. I didn't have an structured way to approach learning about learning; I will certainly keep this in mind as I plan future learnings and experiments.