Singular Learning Theory Comprehensive - 3
Introduction
We continue where we left off. We will now analyze the special case when the true distribution is regular for the statistical model. This covers the classical Bayesian statistics. We analyze the regular posterior distribution, and observe the universal structure that is the asymptotic behavior of this distribution. The theorem is akin to central limit theorem, called the Bayesian central limit theorem / Bernstein–von Mises theorem. We also look at the expansions of the observables in this special case. We finally introduce point estimators, and with this, we cover the conventional asymptotic theory and the basic Bayesian treatment.
This part is more intensive than the previous parts, we are reaching the deep parts of the subject slowly. It is recommended to read the writeup on primer probability, Laplace method and the other ones in the notebook before moving ahead. The main references continue to be the Watanabe books, though I approach it differently for a better understanding.
Observe below that we have already forgone realizability, a major assumption that need not hold in the models of our interest (such as the neural networks we use in practice).
The 1-D example
Our first goal is to estimate the partition function, which is a Laplace integral. We will first start with a simpler example, and we will build and motivate our solution from this method. Once the method is understood, it is all about the control.
Theorem: Let have a unique minimum in the interior. Let . Let , . Then
(Warning: Do not worry about all the analytical assumptions too much just now, you will add them yourself as you go through the proof)
Proof:
We first centre the integral.
Now we do a Taylor expansion of the exponent of the exponential, as Laplace method dictates. Recall that according to the method, we must split the integral into two parts. The first part will be over the set where most of the mass is concentrated (it is clear that this is some neighborhood of ) and we must show that the second integral is negligible compared to the first one, typically by looking at the fraction. The control is in choosing the right neighborhood. Also recall that we estimate the neighborhood integral by sandwiching. Let us do this now.
This is the Taylor expansion with the Lagrange remainder. , and , which is the parameter space for this case.
The idea is here that we really want to be integrating only the first two terms, as we can handle that. So we want the third one to be negligible. Thus we want to naturally choose a neighborhood around of radius (dependent on because we want to zero in on where the mass of the function is as increases, so we must decrease the neighborhood accordingly), , and in this set, , hence we get the first control that we want: .
Let us estimate the integral on this set. First of all, observe that the third Taylor term, which we now call , satisfies (Observe that is compact and the third derivative of is continuous) On this note, we call the first term and the second term and respectively.
Also, since is on a compact interval, we can write
for some constant .
Hence, we can now sandwich and estimate the integral.
where is the leftover integral.
Since and , the prefactors tend to and tend to . Dividing the sandwich through, we get, as ,
Let us calculate the integral and be over with it. Actually, to compute the integral, we would actually have to complete the square and then use the Gaussian integral formula, but we are in good luck since , since is the minimum. But still keep the method in mind, it will come in useful later.
Thus, we can do a simple change of variable to get
We can choose such that as (this is not contradicting the previous control). Hence . Hence, our estimate of the integral is
Now we need to show that the second integral is negligible, that is , where .
This involves finding a lower bound for . This will upper bound our integrand, which is what we want.
Let us write the Taylor expansion again.
The idea is that even though we know that the second order term is positive (assumption), the third order term can be negative, and hence it is not straightforward to lower bound this. However, if we can upper bound the norm of the third order term, we can do this. The idea is that it is less than any given factor of the second order term (say half).
Claim: For all with , for a suitable chosen below,
The idea is that we split into two parts, and .
Let us analyse on now. Let bound on . For ,
where the last inequality needs . This gives our reverse engineering choice of , in which case the claim is true. Hence on , .
What happens on with this choice of ?
Observe that is a compact set, hence since is a continuous function and is the unique minimum, for all , where the minimum is achieved. Note that depends only on , not on , while . Hence for a large enough , , and we have
The same is true for by the previous argument, hence it holds on all of . Now let us bound the integrand. Since we have
we get,
Take the fraction now. For large , we get
Hence we want .
This condition does not contradict the previous conditions. In fact, from all the conditions, and letting (power scaling, natural), we get . I will keep the choice . This allows all the control conditions to work, and the same control is used in the multidimensional case as well. It is important to note it now, as I will simply use this control right from the start and it will all magically work out, but the point is that the interval is forced and all the choices there will work, and it is not really magical.
Estimating the Partition Function
We will now estimate the asymptotics of the normalized partition function. We will be considering the regular case with compact parameter space in a Euclidean space .
The issue is that we cannot directly Taylor expand this. , but it need not be true . Hence, something else needs to be done. We need to decompose our function into a distance-like function and a fluctuation around that and hope it behaves nicely. But why should we expect that? The fluctuation can have a very weird distribution after all. The central limit theorem comes to the rescue in controlling that, it says that the fluctuation is bounded in large samples. This is the universality principle to the rescue, which says that you should expect the fluctuation of to be bounded if the variables are weakly dependent and sufficiently smooth.
So we define the empirical process
This centres the expectation (it is now ) and by the multidimensional central limit theorem, this converges in distribution to the Gaussian distribution, hence it is ,
We divide our partition function into an integral around a neighborhood (we call that the essential part) and the complement of that (called the non essential part). We will see that decomposition into this process works in trying to estimate the essential part, and in figuring out the non essential part, we will need to modify it more. There, we will perform a simple instance of what Hironaka Resolution Theorem does. It is important to grasp that idea and why it comes in the first place. We will see what that is about when we get there.
Essential Part
We have
We will Taylor expand both terms with Lagrange remainder. I am going to remove terms directly which are , you can check that. Call the Hessian of at as . The first term is the second term is . The first term, second term, and the third term are respectively.
We want to discard all the terms into our remainder function for sandwiching.
Let us look at first.
Our choice of is predecided, , so we see that . We can safely push this into our remainder.
Now, is a random variable. Let be the norm of the Hessian, which is a random variable. Observe that
Since , , we have .
Hence is , and similarly, is similar: in the point is random and depends on , so the pointwise central limit theorem is not enough. We assume the uniform bound . Then , since .
Hence we write
We will sandwich and estimate the same way we did for the one dimensional Laplace integral. Assume all sufficient smoothness for the bounds. Let where in probability. Then we have
Now all that is left to do is to compute the integral The idea is to complete the square, write the sum of quadratic + linear as a quadratic + constant. And we know how to handle the quadratic for the exponential, and the constant gets handled trivially.
One can write
where and and .
We first do the change of coordinates from to which does not change the integral as it is just a translation, then from to which gives the factor (this is just the determinant of the transition from the latter to the former).
Now the final transition from to gives the factor
Hence,
where is some ellipsoid. But this is contained in a ball centred around , and it also contains a ball around , and both of them tend to the entire as , in probability. Thus we can adjust the lower bound and the upper bound to have the same integral over 𝕕 instead, which is just . Taking , we get that the lower bound and the upper bound are the same, and the integral in between is
Now we deal with the non essential part.
Non Essential Part
Now we need to lower bound . Well, we can bound , but how do we go about ? We cannot actually do this, since this fluctuation is . We will need to do some sort of normalization here.
Let us first deal with though. We will do it the same way we did the one-dimensional case. The only thing that changes is that this is multidimensional, so the min of the second order term is , so we get the lower bound this time as . Feel free to do the argument again formally here as an exercise.
Let us deal with the fluctuation. As discovered, we will need to normalize this, where the normalization factor has to be of order with respect to .
Define the process
Claim:
Proof: We assume that relatively finite variance holds for the log density ratio function: for all .
Now,
Hence,
When we move to the more general case, we will keep this assumption, and I am not sure why we should assume this to be true. But it will still be a great improvement from assuming regularity.
There are two issues with the definition above though:
- In a more general model, is not one point, it is a real analytic set. Hence cannot be treated locally. More precisely, what this means is that the Taylor expansion we have been doing does not work directly, since now the Hessian is degenerate (not positive definite) (this is a heuristic argument). So this definition of the process does not really work anymore. Keep this in mind, we will deal with this formally later.
- The second is the fact that the process is not defined at as the denominator is . It may not even be possible to continuously extend our process to .
The second issue holds even for our regular model. What is the “resolution” here? The solution is to perform a diffeomorphism of the space, and then extend the function to there. This is the main idea one should keep in mind. We describe the solution formally now.
Define a diffeomorphism from to that takes where and . The numerator becomes and . One can check that the limit now exists.
The more general case is handled by Hironaka’s Resolution Theorem, which deals with both the issues mentioned above.
By using the inequality , we get (take )
where . Observe that even if , need not be . Hence, this is an added assumption that . Also, the existence of comes from the fact that is a continuous function (we have extended it continuously) on a compact space.
This finally enables us to sandwich the non essential part of the partition function.
Using the lower bound , together with on , we get
Now, since is a density function, the integral over is at most . Since , the factor is , and it does not affect the conclusions below. This directly gives the following two results, the second one being what we need to show:
Hence the non essential part is negligible, and we have estimated the essential part of the partition function. We are done here.
We get the asymptotic expansion of the free energy in the regular case now.
Finally, to get the formula, we apply the log to our estimate. Hence,
Let us test out our theory on a regular model.
(Figures generated by Opus 5.5)
The sine regression model : the true model (teal), two other candidate models (purple), and a sample of 60 points (orange).
Free energy computed by numerical integration (solid) against the prediction from the expansion above (dotted), for three growing samples. The lower panel shows the gap between the two, which goes to . (“Theorem 4” in the plot title refers to the free energy expansion above.)
Do note one thing: We cannot make a statement about the expectations yet, we need to show uniform integrability for that (recall from the probability primer: convergence in distribution/probability + uniform integrability implies convergence in expectation)
Asymptotics of the Posterior
First, observe the following:
From now on, is a function of the normalized coordinate , so means evaluated at . Let us now estimate . Not a lot changes from the way we estimated . Do the same completing the square and change of variables to get some which has the Jacobian factors and the term. But this time, we cannot directly take out the remainder term and the term, since can also be negative.
Let . Our idea is that we hope for the leftover to be close to , so we look at the difference as the error term and show that it goes to . We decompose the error terms into three parts:
Whenever is integrable (special case: is a polynomial), it is easy to see that each of these terms are bounded by terms that go to in probability.
Hence,
When is a continuous bounded function, hence this is negligible, and we have the estimate we want. The other special case is for polynomial, when the result still holds, since for every .
Let . Then
Claim:
This is simple.
In particular, and .
Changing back to the original coordinates, through the two formulas (Observe that can be pulled out under )
We get the following:
Now note that this claim is true about every bounded continuous function . This is just the definition of convergence in distribution though! Hence the posterior law of converges to in probability (probability based on the sample, that is what the posterior depends on).
Translating back to , we get
(Figures generated by Opus 5.5)
The exact posterior in the coordinates (surface) against the Bernstein–von Mises prediction (mesh), at . For small the posterior still puts mass on other modes; by it matches the standard bell.
Total-variation distance between the exact posterior and the Gaussian , for three datasets. It shrinks roughly like .
Regular Asymptotics of our Observables
We now derive the asymptotic behaviors of the 4 observables we have defined.
Now, recall the theorem from the last part. I recall it here ( has the same expansion as )
We just now need to estimate the terms. Let
Now, (By Hölder), now since this is a polynomial, by the previous part we have that of the remainder is . Hence, By using our posterior average bounds,
For a single fixed the error coming from in the first term is only . It becomes once we average over , because and . Only these enter the losses.
We also have that
We can again substitute in our posterior averages.
Using the following substitutions:
Here , and in probability.
We get the following asymptotes for the losses (it is an exercise to substitute them and check).
And has the same expansion as .
The distribution of . By the central limit theorem, converges in distribution to (the mean is because ). Hence converges in distribution to , and . In the realizable case , so is asymptotically .
Observe that we cannot just take expectations of these random variables to get the asymptotic behavior of the expectations. However, we can do if they also happen to be uniformly integrable (see the probability primer). But since that proof is neither very instructive nor repeated, we are just going to assume that. Hence, :
Hence, cross-validation (and WAIC) is an asymptotically unbiased estimator of the generalization loss, while the training loss underestimates it by on average.
Checking the losses on the sine model
Let us check if the theory actually works. We are dealing with the sine regression problem here. The figures are generated by Opus 5.5.
Realizable case (true noise , so ). Solid: computed exactly. Dotted: the predicted expansions above (“Theorem 6” in the legend), on the same sample. Each loss locks onto its prediction as grows. Note that on each dataset generalization and cross-validation move in opposite directions: they carry and respectively.
We now check the non realizable case. The best parameter is still , but now , so .
Misspecified case (true noise , ). The predictions still hold. The training loss is now biased by rather than , so a fixed correction of (as in AIC) would be wrong, while cross-validation and WAIC pick up the right amount from the data.
Point Estimators
There are other statistical estimation methods when the true distribution is regular for the statistical model, that are not entirely Bayesian in nature.