Prior distributions, partial pooling, and reference sets (Bayesians are frequentists)
Sudip Paul writes:
I’ve written a blog post arguing that the philosophical reference class problem — which comparison group is “right” for assigning a probability to an individual case — is, once the grouping structure is fixed, the same estimation problem that partial pooling was built to handle. The bias-variance tradeoff between narrow and broad reference classes is resolved by multilevel shrinkage rather than by picking a level. The post includes a literature review. One finding that may interest you: I couldn’t find any published work that explicitly connects the reference class problem to partial pooling in multilevel models, despite the machinery being standard. The closest prior work connecting the two comes from David Poole’s group at UBC, published in AI venues. The philosophical literature (Hájek, Wallmann, Roth & Tolbert) discusses the problem extensively without arriving at partial pooling.
I replied with some links:
– In chapter 1 of Bayesian Data Analysis we discuss the foundations of probability and the idea of a reference set. In the third edition, see page 12 where we explicitly use the term “reference set” and sections 1.6 and 1.7 for empirical examples.
– I have a post from 2018 called “Bayesians are frequentists,” where I write, “the Bayesian prior distribution corresponds to the frequentist sample space: it’s the set of problems for which a particular statistical model or procedure will be applied.”
– And another post from around the same time, “The idea of replication is central not just to scientific practice but also to formal statistics . . . Frequentist statistics relies on the reference set of repeated experiments, and Bayesian statistics relies on the prior distribution which represents the population of effects.” This is cool because we’re connecting three concepts: the Bayesian prior, the frequentist reference set, and real-world replications. This relates to multilevel modeling too, in that studies of real-world replications will be done using meta-analysis, which uses multilevel modeling.
– And here’s a 2011 article, “Bayesian statistical pragmatism,” where I write, “Bayesian probability calibration is closely connected to frequentist probability statements, in that both are conditional on ‘reference sets’ of comparable events” and “Bayesian probability, like frequentist probability, is except in the simplest of examples a model-based activity that is mathematically anchored by physical randomization at one end and calibration to a reference set at the other.”
Bayesian inference is all about partially pooling toward the prior, and “the prior” represents a population or reference set of relevant problems.
In short, I agree with Sudip Paul, and it’s super-frustrating that many people don’t seem to realize this point.
Paul responds:
Having read through the references you sent — I think the gap my post is trying to fill is the last step: connecting what you’ve been saying about priors and reference sets to the philosophical reference class problem by name — Venn, Reichenbach, Hájek — and framing partial pooling as the resolution to what philosophers have treated as a foundational problem for 150 years. You’ve been making the statistical point for years, but the philosophers discussing the reference class problem haven’t found it (Hájek’s 2007 paper doesn’t cite any multilevel modeling literature, and Wallmann and Williamson’s 2017 survey of four approaches to the problem doesn’t include partial pooling). The two conversations have been happening in parallel without connecting.
In the months since the above exchange, Sudip Paul has expanded his post into a paper. He adds:
Beyond the post, the paper adds a decision-theoretic argument (via Stein inadmissibility, every selection rule in the philosophical literature is dominated by shrinkage in total risk), shows that the classical proposals — Reichenbach, Carnap, Kyburg, Pollock — are limits or special cases of the same pooling apparatus, and works through the exchangeability bridge from de Finetti to the multilevel model.