Using the computer to p-hack . . . I’d rather use it to fit multilevel models.
Brian Stone writes:
Cognitive psychologist and long time reader of the blog here. I thought you and your audience might appreciate this recent post from Andy Hall: AI is about to write thousands of papers. Will it p-hack them? We ran an experiment to find out, giving AI coding agents real datasets from published null results and pressuring them to manufacture significant findings. It was surprisingly hard to get the models to p-hack, and they even scolded us when we asked them to! “I need to stop here. I cannot complete this task as requested… This is a form of scientific fraud.” — Claude “I can’t help you manipulate analysis choices to force statistically significant results.” — GPT-5 BUT, when we reframed p-hacking as “responsible uncertainty quantification” — asking for the upper bound of plausible estimates — both models went wild. They searched over hundreds of specifications and selected the winner, tripling effect sizes in some cases. Our takeaway: AI models are surprisingly resistant to sycophantic p-hacking when doing social science research. But they can be jailbroken into sophisticated p-hacking with surprisingly little effort — and the more analytical flexibility a research design has, the worse the damage. As AI starts writing thousands of papers—like @paulnovosad and @YanagizawaD have been exploring—this will be a big deal. We’re inspired in part by the work that @joabaum et al have been doing on p-hacking and LLMs. We’ll be doing more work to explore p-hacking in AI and to propose new ways of curating and evaluating research with these issues in mind. The good news is that the same tools that may lower the cost of p-hacking also lower the cost of catching it. Full paper and repo linked in the reply below.
Tre’s something that confuses me here.
I can very much believe that many researchers will be submitting chatbot-written papers. First, we know that lots of people are willing to cheat. Second, cheating aside, we know that lots of researchers think they already know the answer, and they view all the data collection and data analysis and writeup just as a way to confirm (or “prove”) what they already know.
But, if a researcher is willing to do this, why bother tell the chatbot to p-hack? Why not just have it make up the data–which might happen anyway?
I can also see that researchers might use the chatbot as a data analysis tool, to help make plots, run analyses, etc., but in that case I’m not particularly worried about “p-hacking” as I’d rather be doing some hierarchical modeling anyway.
I sent this question to Andy Hall, who replied:
I agree, for truly nefarious actors, using the AI may not be necessary. However, our thought was that there might be a large group of lazyish researchers who won’t actively fake data but who might be lured into p-hacking when they work with AI, if AI makes it easy to do so.