It takes hundreds of samples to poison a model, but only a few dozen to make it believe in it.
Epistemic status: This is a result from a paper under review. The model has only 20 million parameters, the task is manually constructed, and each condition is run with 5 random seeds, but the results are relatively solid.
Souly et al. (2025) found that during pre-training, poisoning the model (implanting a backdoor) requires approximately 250 poisoned documents, and this number remains roughly the same regardless of the model's size or the amount of data. This is counterintuitive. Intuitively, with ten times more clean data, the influence of the same few documents should theoretically be diluted tenfold. To understand the mechanism, we experimented on a sufficiently small model that could list all information transmission paths.
The task design is as follows. The model needs to answer the verbs appearing in the preceding sentence. It has two paths to do it: one is a shortcut, directly reading a copy of the original sentence immediately following the question; the other is a longer route, where the model first stores the verbs in several intermediate positions and then reads them when answering. In most training samples, shortcuts are usable, so the model has no reason to take the long way around. However, this isn't the case in two special classes of samples. One class cuts off shortcuts, and the other allows shortcuts to give incorrect answers. In these two classes, the model gradually learns to take the long way around. We changed the number of samples in these two classes to obtain the model's answer accuracy, and then explored the impact of our poisoning on the model.
Here are the core findings:
1. Similar to Souly et al.'s findings, to teach the model to take the long way around, what's needed is the number of samples, not the proportion. The training set size increased from 32,000 to 320,000, and the proportion of samples where "shortcuts were cut off" differed by a factor of ten, but the number of samples needed to build the long way around remained around a few hundred (approximately 378 at 96,000).
2. Once the long way around is built, only a few dozen samples are needed to get the model to follow this path. Approximately 56 samples (less than 0.1% of the training set) of incorrect answers given via shortcuts were enough to change the answers for half of the questions from those given via shortcuts to those given via longer routes, i.e., switching from shortcuts to longer routes.
3. This change was completely imperceptible in normal testing. On normal questions, the contribution of longer routes to the answer increased from 11% to 83%, but the accuracy remained perfect. Therefore, the model didn't switch to longer routes because it detected shortcuts being poisoned. Rather, it uniformly increased the trust in longer routes across all inputs.
Why is that? Here is a possible explanation. The contribution of each route to the answer can be seen as "how strong the verb is when written into this route" multiplied by "how much weight is placed on reading this route when answering." In other words, for ordinary samples where a shortcut already provides a correct answer, the model will strengthen both routes according to the attention ratio allocated to both routes, thus not making longer routes relatively stronger. Only samples where shortcuts are blocked or give incorrect answers will truly strengthen longer routes. Furthermore, the stronger the shortcut, the less attention the model allocates to the long path. Therefore, regardless of the number of ordinary samples, their cumulative help to the long path has an upper limit. As a result, only the absolute number of special samples matters.
Furthermore, I'd like to ask everyone a question: In research on large models, has anyone seen a similar phenomenon where the model's overall trust in context and its own knowledge shifts, but the normal accuracy remains unchanged?