Digital twins version 2.0: The skeptics speak up
This is Jessica. Several years since the earliest research papers on the topic came out, the idea of using LLMs to simulate human behavior like survey responses or experiment outcomes (“silicon sampling”) is hitting mainstream. People tend to bring this up in conversation with me, probably because I’ve written about validation and taught a course, so I get to hear a lot of other researchers’ views. I would put the majority of takes I encounter in either the more agnostic, “Somebody should probably explore how much we can do with them” and/or “I’m pretty doubtful it’s going to live up to the hype” camps.
One way to use silicon samples is digital twinning, where simulations are conditioned on considerable information about specific individuals, and expected to predict the responses of those specific individuals. Early papers on digital twins sing their praises, showing for example that they can reproduce people’s responses more than 80% as well as people replicate their own responses (e.g. here or here). But we’re now seeing a more critical empirical push, which uses novel tasks and topics and more careful design of metrics and baselines for comparison to gauge how much conditioning on more individual-specific information improves prediction.
In the last month, there have been a few noteworthy data points. Pew Research put out a report describing what they learned from comparing digital twins implemented by giving Claude Opus 4.6 detailed 2025 survey responses and demographic persona information for people who responded to several more recent survey waves. They compute the average absolute error in percentage points for each multiple choice survey question (i.e.,the average of the absolute differences in the proportion of LLM versus human respondents selecting each answer option), and find it averages 12 percentage points over the 300 survey questions they looked at, with 28% of the questions having error over 15 percentage points. They observed several kinds of biases that have come up in other research comparisons: less variation in LLM responses compared to human (with some answer options never being picked by the LLM), stereotyping (like predicting that 97% of Hispanic adults are at least somewhat likely to follow the World Cup; in reality its only 43%), overestimating human accuracy on factual questions, and mispredicting sentiment on current events like data centers and face coverings on ICE officers. They stress how hard it is to predict how far off the LLMs will be on a given survey question.
This list of biases closely resembles those proposed by Tianyi Peng et al. in one of the most comprehensive evaluations of digital twins to date, recently published in Science Advances: Digital Twins as Funhouse Mirrors: Five Systematic Distortions. They collect a new set of 19 preregistered studies involving 164 outcomes, and find that overall, there is only weak improvement over base models and relatively low correlations with human behavior (on the order of 0.2).
I like this paper a lot, in particular the comparisons they do to simpler demographic prompting and base models. The intuition that has been sold by researchers and companies in the space alike is that with rich enough information about individual people, we can predict their behavior with high accuracy. To make the twins approach worthwhile, we want to see that twinning specific individuals leads to responses that predict variation between individuals better than the demographic baseline. However, some of the early papers advocating for digital twins evaluated them by looking at measures where a model can do well simply by being a good predictor of the average response. For example, you can look at the correlation between each individual’s real and twin response over a set of likert-style survey questions, but a model that does a good job of predicting on average which questions will get higher responses may appear to do well by this measure, even if it does a poor job of sorting individuals for any particular question.
Instead, Peng et al. look at correlations between human and twin responses for a single outcome at a time, and compare to demographic personas only as well as base models, to ask how much additional predictive power the twin approach provides when it comes to sorting people into those that will respond with higher versus lower values. They find that while there’s a boost in average correlations (from 0.08 with base models to 0.15 with demographic personas to 0.2 with twins), overall the correlations remain modest, and much more modest than the correlations over questions that other papers have reported. Hopefully it encourages more care in future papers when it comes to what measures actually support the claims being made. I do think behavioral researchers and others interested in this space are gradually changing their mental models to conclude that LLMs may do a decent job of predicting population averages for well-represented populations, but they have a long way to go when it comes to modeling individual heterogeneity. There are several other papers along these lines that I haven’t had a chance to look closely at.
How damning is this for the approach itself? I think it remains hard to say, though it certainly suggests skepticism is warranted at this point in time. The reasons I wouldn’t write digital twins off completely are 1) it is still early and models continue to improve and 2) prompting is often not the best way to recover the predictive information a model has–you can do better at opinion prediction by using model internal representations (see, e.g., this). On the other hand, ultimately it comes down to limits on how predictable individual behavior is, and while people are very predictable in lots of ways, we are also always exercising our agency and changing in subtle ways that interact with the world, which even we ourselves can’t necessarily understand or predict. So I don’t think reaching AGI is synonymous with being able to predict the majority of individual opinions and actions. I’m reminded of Matt Salganik et al’s work on the Fragile Families Challenge, predicting life outcomes from thousands of variables, where the best of the 160 teams’ models explained only about 20% of the variation on the most predictable outcomes, despite the enormous amount of data.
What I will say with certainty is that silicon sampling has attracted some very good salespeople. I expect a lot of money will be made before anyone abandons the idea.