generative Bayesian computation
Bemused, I scrolled though the list of hundreds of papers posted by Nick Polson this year, a singularity event already discussed by Andrew on his blog. (Where he even witnessed live those papers vanishing from SSRN, on SSRN decision.) Since most of them have disappeared, I cannot read the Sufficiency without a Likelihood in Population Genetics that could have proved of interest, at least judging from the abstract
Approximate Bayesian Computation was invented in population genetics because coalescent likelihoods are intractable, and it has carried two acknowledged defects ever since: the analyst must choose summary statistics, so that the object computed is π (θ| S (y)) rather than π (θ| y), and the analyst must choose a tolerance, which trades bias against acceptance rate. We propose replacing the whole apparatus with a learned transport map. By the noise outsourcing lemma the posterior is a deterministic function of the data and an independent uniform, and that function is the population minimiser of an expected check loss, so it can be fitted by regression on simulated parameter and data pairs. There is no summary statistic and no tolerance. Our main theoretical contribution is Theorem 4, which shows that if a permutation invariant architecture with a d-dimensional bottleneck attains the Bayes risk of the check loss at almost every quantile level, then the bottleneck is a sufficient statistic. This turns sufficiency into something that can be tested by fitting, and it makes the infinite alleles model, where Ewens (1972) proved that \(K_n\) is sufficient, the natural validation problem. Empirically the method rediscovers that theorem: a one-dimensional bottleneck matches an eight-dimensional one to within Monte Carlo error and its learned scalar has Spearman correlation 0.998 with \(K_n\), while a plausible but non-sufficient summary is 62% worse. We then apply the same procedure to the infinite sites model and prove that Watterson’s S is not sufficient there. The learned posterior using the whole site frequency spectrum is 21% narrower than the exact posterior …
even though I cannot make sense of the sentence “the posterior is a deterministic function of the data and an independent uniform“. Posterior function? posterior value? I then went to Polson’s arXivals this year (12 so far), which are yet available, finding the 2024 Generative Bayesian Computation for Maximum Expected Utility by Polson, Ruggeri and Sokolov having links with ABC. The notion is to estimate a neural generative model for the utility
\[
U_d^{(i)}=H(S(y^{(i)}),\theta^{(i)},\tau^{(i)},d)
\]
where the lack of index on \(d\) (which should be \(d^{(i)}\) ?) is surprising, as it seems to indicate a neural network constructed for each decision \(d\)