247A, Notes 1: Rearrangement-invariant spaces

Disclaimer: due to current events, I have not been able to devote as much time to lecture notes preparation as I would have liked, so I apologize in advance for the unpolished nature of the text below, which has been largely recycled from previous lecture notes I have written.

This is the first set of lecture notes for my graduate course 247A, “Fourier analysis”. The course name is rather general, but I will focus the course not on the Fourier transform per se, but on the closely related topic of real variable harmonic analysis, with a particular emphasis on Calderón–Zygmund theory, which underlies basic tools in PDE such as the theory of Sobolev spaces.

To avoid confusion at the outset, let us make the distinction between real-variable harmonic analysis and abstract harmonic analysis, which are only distantly related to each other despite the similar names. Abstract harmonic analysis, roughly speaking, is the extension of the classical theory of the Fourier transform to other domains, such as locally compact abelian (LCA) groups, non-abelian Lie groups, or symmetric spaces, and typically involves a blend of representation theory, group theory, and analysis. Real-variable harmonic analysis, by contrast, tends to work on classical domains, such as a Euclidean space , a torus , or a lattice , although many of the techniques can extend to more general domains (e.g., to Riemannian manifolds). While the Fourier transform often plays a prominent role (in particular, by setting the stage for time-frequency analysis and enabling various decompositions or other transforms that involve frequency space or phase space in addition to physical space), real-variable harmonic analysis is often focused on estimating other transforms or expressions that often interact well with the Fourier transform, but need not explicitly invoke it. Examples include the Hilbert transform (where we have made the somewhat arbitrary decision to omit the normalizing constant ) or the Hardy-Littlewood maximal function

\displaystyle Mf(x) := \sup_{r>0} \frac{1}{|B(x,r)|} \int_{B(x,r)} |f(y)|\ dy.

A typical question in harmonic analysis is the following: let be some function on a standard domain (such as Euclidean space), and let be an explicit transform of (e.g., the Hilbert transform or maximal function ). To what extent is the “size” of controlled by the “size” of ? The value of such bounds often lies in the general nature of the input function ; some mild regularity or decay hypotheses might be imposed on , but beyond that the function is typically not required to have a very structured form (in particular, it need not be describable by any closed-form expression).

In many situations the transform being studied is linear or sublinear, in which case the natural type of bound to ask is a linear bound

\displaystyle \|Tf\|_Y \leq C \|f\|_X

for suitable function space norms (e.g., norms), and is some bound. Depending on the application, we may be interested in various levels of precision regarding the bound :

  • (a) Optimal bounds, in which we seek the exact optimal value of (i.e., the operator norm of ). For instance, the optimal constant for the Hilbert transform is exactly , whereas the optimal constant for in one dimension turns out to be (a result of Melas).
  • (b) Bounds accurate up to absolute constants (or maybe constants that can depend on basic parameters such as the ambient dimension).
  • (c) Bounds in which we are willing to accept “logarithmic type losses” such as or in auxiliary parameters, such as a scale parameter .

All three regimes are interesting, but we will focus in this class on the regime (b), where we can “afford” to lose absolute constants in the bounds, but will work hard to avoid any logarithmic losses. In particular, significant effort will be devoted in this class to avoiding “logarithmic pileups of scales”, in which the contributions of different dyadic scales such as for all potentially contribute an equal amount that “interfere constructively” to cause a logarithmic divergence. This can be unnecessarily conservative when one is in regime (c) (which is for instance the situation in modern topics such as restriction theory or the Kakeya conjecture); nevertheless, the general skills gained by trying to not lose even a logarithmic factor in the bounds are often valuable in these other types of analysis.

When dealing with linear or sublinear problems, it is natural to try to decompose the initial function into various smaller components by some decomposition , so that the transformed function can be controlled by more tractable expressions in various ways (e.g., via the triangle inequality, by Bessel type inequalities, or by the more modern technique of decoupling inequalities). In short, the subject tends to proceed by a divide and conquer philosophy: it is generally preferable to replace a simple-looking but hard-to-estimate expression with a large, messy-looking combination of expressions that are easier to estimate. As such, the aesthetics of the subject are almost the reverse of those in the more algebraic portions of mathematics, in which progress is often made by making the expressions involved look as simple and unified as possible.

One of the main themes in this classical type of harmonic analysis is the struggle to understand the effect of two phenomena in integrals or sums: singularity and oscillation. The Hilbert transform is a quintessential example of a singular integral, which combines both features: the non-locally integrable nature of the kernel provides the singularity, but the sign change from to provides the oscillation. Classically, the interplay between these two phenomena can be tamed by analyzing the behavior of this operator both in the time (or “physical”) domain and in frequency (or “Fourier”) domain; in particular, the fact that singular integral operators such as the Hilbert transform are simultaneously a well-behaved Fourier multiplier and is “pseudo-local” in physical space lies at the heart of the standard Calderón–Zygmund theory for such operators. This dovetails nicely with more modern “time-frequency analysis” approaches to the subject, which can also handle other interesting operators, such as restriction or Bochner–Riesz operators via tools such as the wave packet decomposition, although these will be outside the scope of this course.

In this initial set of notes I will ignore the effect of oscillation, and develop some tools, such as interpolation theory, which can help control non-oscillatory sums and integrals if they are not too singular. Here, the focus will be on rearrangement-invariant spaces, such as the Lebesgue spaces and their weak variants , which are function spaces that are useful for measuring how “singular” or “decaying” various functions are, but do not pay attention to how they oscillate or where their mass is distributed. As such, these spaces do not capture the underlying geometry of the domain, which also plays an essential role in the subject; but it is nevertheless essential to have a good base understanding of the rearrangement-invariant theory before moving on to the more delicate aspects of harmonic analysis that are sensitive to rearrangements.

We will use the following asymptotic notation throughout the course: , , or denotes the assertion that for some constant , and write for . If we permit this constant to depend on some ambient parameters, we indicate this by subscripts; for instance, or denotes a bound of the form for some constant that can depend on and . As indicated above, in this course we will generally not dwell much on exactly what these constants are, or attempt to optimize them.

Suppose one has some measurable function on some measure space . (Here we will follow the common practice if identifying functions that agree almost everywhere; in particular, we will be content to work with functions that are undefined on a set of measure zero. Also, while we work here with complex-valued functions throughout, most of the discussion here is also valid for real-valued or vector-valued functions.) Informally speaking, to measure how “big” such a function is, there are two (imprecisely defined) basic statistics to be aware of:

  • The height or amplitude of the function, which describes what the typical size of the magnitude is for in the “dominant” component of the support of ; and
  • The width of the function, which describes the measure of this dominant component.
Example 1 (Informal) Given a Gaussian wave packet type function on for some and , this function has magnitude on the ball , which has volume (if we allow constants in the informal notation to depend on the dimension ), so such a function has height and width .
Example 2 (Informal) The function on , which is implicitly involved in the definition of the Hilbert transform , does not have a clear amplitude or width as is. However, if one performs a dyadic decomposition where we use to denote the indicator of a statement (equal to when is true and otherwise), then each component of this decomposition has height and width . Thus, while this function can be viewed as a superposition of components of various heights and widths, rather than a single such component.

These informal concepts of height and width are too imprecise to work with in practice. Experience has shown that a convenient proxy for these concepts are the norms of a function , defined for as

\displaystyle \|f\|_{L^p(X,{\mathcal B},\mu)} := \Big( \int_{X} |f(x)|^p \, d\mu(x) \Big)^{1/p}

and for as

\displaystyle \|f\|_{L^\infty(X,{\mathcal B},\mu)} := \mathrm{ess\,sup}_{x \in X} |f(x)|

where denotes the essential supremum of the function with respect to the measure . Often we abbreviate as , , , or just (and abbreviate as ) when the missing arguments are clear from context. (For instance, when working with Euclidean spaces , the measure is understood to be Lebesgue measure, and the Lebesgue -algebra, unless otherwise specified.) In terms of the width and height of a function , one heuristically has

\displaystyle \|f\|_{L^p(X,\mu)} \approx H W^{1/p}

for both finite and infinite values of , with the convention that is equal to when is positive and when is zero. In the case of a step function (where now is the indicator function of a measurable set ), this heuristic becomes exact:

\displaystyle \|f\|_{L^p(X,\mu)} = A \mu(E)^{1/p}.

The function space is defined as the set of all measurable functions for which the norm is finite, up to almost everywhere equivalence, though we will often abuse notation by identifying a function with its almost everywhere equivalence class.

In the case where is discrete and is counting measure, we abbreviate as , or even just .

Example 3 Let . On a Euclidean space , the function lies in (with a norm of ) if and only if , while the function lies in (with a norm of ) if and only if . The function does not lie in any , although it only fails “logarithmically” to lie in . Thus we see that control in for high rules out severe local singularities at a point, while control in for low rules out insufficiently rapid decay at infinity.

As is well known (see these previous notes) these function spaces enjoy many useful properties:

Theorem 4 (Basic properties of spaces)
  • (i) The space is a Banach space when , a Hilbert space when , and a topological vector space when .
  • (ii) If obeys the scaling condition then one has the Hölder inequality for any measurable (here we adopt the usual conventions ). In particular, if and , then .
  • (iii) If for some , then one has the duality relationship where is the conjugate exponent to , defined by .
Remark 5 Closely related to (iii) is the fact that the dual of can be identified with when (with the additional hypothesis that is -finite if ), but in practice the relation will already be good enough for our purposes.
Remark 6 The Banach space property gives us the basic triangle inequality for both finite and infinite collections of functions when , where in the infinite case the assertion is that if the right-hand side is finite, then the series is absolutely convergent almost everywhere, and obeys the above inequality (so in particular is in ). For , this inequality fails (can you come up with a counterexample?), but one has the weaker -triangle inequality in this case, which follows easily from iterating the easy observation that for any complex numbers , which in turn ultimately stems from the complex triangle inequality and concave nature of for . In particular, for a finite sum , another application of Hölder’s inequality gives the quasi-triangle inequality for , which is not too much worse than when is not too large.
Exercise 7 Give an example to show that the quantity in cannot be replaced by any smaller quantity.
Exercise 8 For a simple function, verify that , and that , where . For this reason, the measure of the support of is sometimes referred to as the norm of , though it would be more accurate (though confusing) to refer to it as the power of the norm.
Remark 9 Note that Hölder’s inequality is not just symmetric under the homogeneities and of the functions, but also under the homogeneity of the underlying measure. This latter symmetry demonstrates why the condition is necessary. (The first two symmetries demonstrate why appears the same number of times on both sides of the inequality, and similarly for .) In the case of Euclidean space, the measure homogeneity symmetry is equivalent to the scaling symmetry for , as the Jacobian of this map is . But the point is that by manipulating the measure directly, one still enjoys this symmetry even when no scaling operation is present.

It is instructive to try to understand inequalities such as using the height-width heuristic introduced previously. Suppose informally that have heights , , and widths , , respectively. Then one expects the heights to be related by the formula

\displaystyle H_{fg} \approx H_f H_g.

What about the widths? Heuristically, the region that concentrates in ought to be a subset of the region that concentrates in, so

\displaystyle W_{fg} \lessapprox W_f

and similarly with replaced by . We can combine these bounds as

\displaystyle W_{fg} \lessapprox \min(W_f, W_g).

The bound then is morally

\displaystyle H_{fg} W^{1/r}_{fg} \lessapprox H_f W^{1/p}_{f} \cdot H_g W^{1/q}_g,

which on applying the previous bounds and should simplify to

\displaystyle \min(W_f, W_g)^{1/p} \min(W_f, W_g)^{1/q} \lessapprox W^{1/p}_{f} W^{1/q}_g.

But this is clear by bounding by for the first factor on the left-hand side, and by for the second factor. Thus we see that the key geometric input that is morally driving the Hölder inequality is the simple fact that the concentration region of the product is contained in the concentration regions of the factors.

Exercise 10 Determine the cases for which holds with equality (dealing with edge cases such as when one or more of equal infinity as appropriate). Discuss how your conclusions align with the heuristic analysis presented above.
Exercise 11 If , determine the cases for which holds with equality. What changes when or ?
Exercise 12 Show that Hölder’s inequality is equivalent to the log-convexity of norms: (For technical reasons one needs to first reduce to the case where has finite measure, and then and are everywhere non-vanishing simple functions. Now consider the convexity of with respect to a measure for some suitable exponents .)
Exercise 13 (Direct approach to log convexity) Differentiate twice with respect to and show that this is non-negative (take to be a non-zero simple function with finite measure support to avoid technicalities). This is an example of a monotonicity formula method — deriving estimates from a monotonicity property, which in turn follows from the non-negativity of a derivative.

You will see that this approach is surprisingly messy. For all the other ways, observe that enjoys homogeneity symmetry in both and , which lets one normalise both and to equal one. Thus the task is now to show that if , then for all between and . This can be done by the pointwise convexity of , or more precisely the estimate

\displaystyle |f(x)|^r \leq (1-\theta) |f(x)|^p + \theta |f(x)|^q;

the observant reader will note that this is merely the proof of Hölder’s inequality in disguise.

Let us now give a more unusual proof of the log-convexity which does not appeal to any pointwise convexity estimate, instead combining the “divide and conquer” strategy with an elegant (and rather cheeky) “tensor power trick“. Again normalise . We split into a broad flat piece and a narrow tall piece

\displaystyle f = f 1_{|f| \leq 1} + f 1_{|f| > 1}

which are disjoint, and thus

\displaystyle \|f\|_r^r = \int_{|f| \leq 1} |f|^r + \int_{|f| \geq 1} |f|^r.

What we are doing here is exploiting some very basic intuition about norms, namely that bounds for large tend to exclude tall narrow spikes, whereas bounds for small tend to exclude short broad tails. Of course, either sort of bound would exclude tall broad functions, and neither excludes narrow short functions. Once again, this intuition can be buttressed by considering the special case of step functions.

When , then , and when , then . Thus we end up with

\displaystyle \|f\|_r^r \leq \int_X |f|^p + \int_X |f|^q = 2.

The above argument (which is a prototype of the real interpolation method) obtained an estimate which is off by a factor of two from what we wanted; this is a typical feature of the method. However we can recover this factor for free by the following tensor power trick. Let be a large integer. We replace the measure space by its power using the product measure construction, and similarly replace with its tensor power , defined by

\displaystyle f^{\oplus M}(x_1,\ldots,x_M) := f(x_1) \ldots f(x_M).

One then observes that

\displaystyle \begin{array}{rl} \|f^{\oplus M}\|_{L^p(X^M)} &= \|f\|_{L^p(X)}^M = 1; \\ \|f^{\oplus M}\|_{L^q(X^M)} &= \|f\|_{L^q(X)}^M = 1; \\ \|f^{\oplus M}\|_{L^r(X^M)} &= \|f\|_{L^r(X)}^M. \end{array}

Now we apply the preceding arguments to instead of to deduce that

\displaystyle \|f^{\oplus M}\|_{L^r(X^M)}^r \leq 2

which on taking roots gives

\displaystyle \|f\|_{L^r(X)}^r \leq 2^{1/M}.

Now the left-hand side is independent of ; take limits as and we obtain as desired.

The tensor power trick can be viewed as another application of symmetry: if an estimate is invariant under raising to a tensor power, then one can automatically replace all absolute constants with ; thus we obtain the “free lunch” of deducing a bound with an explicit constant , from a bound with an unspecified constant (or even with “logarithmic losses”). Contrapositively, if an estimate is invariant under tensor power, then a weak counterexample (which shows that the constant must exceed one) can be amplified into a strong counterexample (which shows that no finite constant suffices) by tensor powering. The tensor power trick seems like a magical trick at present, but is actually exploiting some basic results in information theory such as the Shannon entropy inequalities and the central limit theorem; it also combines well with virtually any inequality which involves Gaussians. Unfortunately due to lack of time we will not be discussing these beautiful topics further in this course. At any rate one sees the power of abstraction in this tensor power trick. (One could similarly perform this trick in , so long as the constants only grew sub-exponentially in the dimension .)

The final proof of log-convexity of the norm that we give here proceeds via complex analysis, and the maximum principle — which in many ways is a complex analogue of convexity (or subharmonicity). We need the following result from complex analysis, namely a form of the Phragmén–Lindelöf principle.

Lemma 14 (Three lines lemma) Let be a complex-analytic function on the strip , which is of at most double-exponential growth, or more precisely for some . Suppose that we have the bounds when and when . Then we have for all in the strip.
Remark 15 The rather strange sub-double-exponential hypothesis here is completely sharp, as the example shows. Note in this hypothesis that we allow the implied constants in the asymptotic notation to depend on , but the hypothesis is qualitative rather than quantitative: the value of these constants is irrelevant for the final conclusion, as long as they are finite. In practice, these sorts of qualitative hypothesis are usually easy to establish (especially when compared to quantitative estimates) by restricting, smoothing, or damping to a nice class of functions, or by smoothing out or discretising various operators and domains. See for instance the proof of this very lemma in which we upgrade “for free” a weak qualitative bound (sub-double-exponential growth) to a strong qualitative bound (decay at infinity).

Proof: The hypotheses and conclusion of the lemma are invariant under the operation of multiplying by a constant (and adjusting appropriately). So we may normalise . Similarly, the hypotheses and conclusion of the lemma are invariant under the operation of multiplying by an exponential for some real . Using this, one can also normalise . So now is bounded by on both sides of the strip and we want to show it is bounded by inside the strip.

Let us first assume that is much better than exponential growth, namely that it goes to zero at infinity. Then for all sufficiently large rectangles the complex-analytic function is bounded by on all four sides of this rectangle, and hence in the interior also by the maximum principle, and we are done by setting .

Now let us handle the general case; as is usual when removing a qualitative assumption, we do this by a limiting argument. We replace by ; a little complex arithmetic shows that this converts the almost double-exponentially growing function to one which is still complex analytic but is now decaying at infinity. It is still bounded by at both sides of the strip, and hence by in the interior also by the previous argument. Now take to conclude the claim.

Exercise 16 Suppose that is analytic on the strip , obeys the sub-double-exponential bound on the strip, and obeys the polynomial bounds on the sides of the strip. Show that it obeys the polynomial bound on the interior of the strip also.

To apply the three-lines lemma to prove , take to be a simple function (with finite measure support) and consider the entire function

\displaystyle z \mapsto \int_X |f|^{z}.

This function has exponential growth at most (because of the qualitative assumption that is simple with finite measure support), and is bounded by on the lines and , and hence (by a trivially rescaled version of the three lines lemma) bounded by on the strip inside the lines. In particular it is bounded by at , which gives the claim for simple functions. The claim for more general functions (dropping the qualitative assumption of simpleness and finite measure support) then follows by a standard limiting argument (using for instance monotone convergence) which we leave as an exercise.

The above argument is a prototype of the complex interpolation method. As one can see, it can give slightly sharper results than the real interpolation method (though using the tensor power trick the real method can sometimes “catch up”), but on the other hand requires the quantities being studied to depend complex-analytically on a parameter rather than (say) real-analytically.

Having conclusively demonstrated the log-convexity in multiple ways, let us now give some quick applications. It shows that control on two extreme norms implies control of the intermediate norms. Under additional assumptions on the measure space , one of these extremes is not necessary. If the measure space is finite in the sense that (thus prohibiting functions from being arbitrarily broad), then higher norms control lower ones: Indeed this is trivial when , and the general case then follows by convexity. The bound can also be usefully written in terms of averages: if we write for , then we see that higher averages control lower averages:

\displaystyle (\rlap{\hspace{2pt}-}\int_X |f|^p\ d\mu)^{1/p} \leq (\rlap{\hspace{2pt}-}\int_X |f|^q\ d\mu)^{1/q} \hbox{ whenever } 0 < p \leq q \leq \infty.

One way to view this is that as one lowers the exponent , the exceptionally large values of become less important, leaving the small values of to dominate. By restricting to its support one can refine to

\displaystyle \| f \|_p \leq \|f\|_q \mu(\mathrm{supp}(f))^{\frac{1}{p} - \frac{1}{q}} \hbox{ whenever } 0 < p \leq q \leq \infty.

(Note that this is a limiting case of log-convexity at the exponent , in view of Exercise ). In the converse direction, if the measure space is granular in the sense that one has a lower bound for all sets of positive measure, then functions are prohibited from being arbitrarily narrow, and lower norms control higher norms:

\displaystyle \| f \|_q \leq \|f\|_p c^{\frac{1}{q} - \frac{1}{p}} \hbox{ whenever } 0 < p \leq q \leq \infty.

This can be seen by first checking the case, and then using log-convexity to get the remaining cases. In particular, in the spaces we see that for . For spaces on points, we thus have (non-matching) upper and lower bounds

\displaystyle \|f\|_{\ell^q} \leq \|f\|_{\ell^p} \leq N^{\frac{1}{p}-\frac{1}{q}} \|f\|_{\ell^q}, \ \ \ \ \ (10)

and norms are comparable to some extent, but the comparability gets worse as or as and get further apart.

Exercise 17 Heuristically justify the bounds , by appealing to the informal notions of width and height of a function.
Exercise 18 When does equality occur for either of the inequalities in ? Note how the example that attains the lower bound is in many ways the “opposite extreme” to the example which attains the upper bound.

Lebesgue measure on Euclidean spaces with the usual Borel or Lebesgue -algebra is not granular. However one can create granularity by coarsening the -algebra. For instance, if we let be the -algebra generated by the lattice unit cubes for , then we have granularity with constant , and now lower norms of functions control higher ones — but only for functions which are measurable with respect to this algebra, i.e. only for functions which are constant on each lattice unit cube. (This is the first time we have actually manipulated the -algebra to say something non-trivial, as opposed to manipulating , , or .) Thus we see that local constancy of functions can lead to additional estimates on norms. Later on we shall see that frequency localisation achieves a similar effect as local constancy, as quantified by Bernstein’s inequality; this is a concrete manifestation of the famous Heisenberg uncertainty principle.

Finiteness and granularity of the measure space prevent a function from being too broad or too narrow respectively. Similar things happen when a function is being prevented from being too tall or too short; for instance if is bounded above by a constant , then we have

\displaystyle \|f\|_q \leq \|f\|_p^{p/q} M^{1-p/q} \hbox{ whenever } 0 \leq p \leq q \leq \infty

(this is just log-convexity at the exponent), while if is bounded below by on its support, then we have the reverse inequality

\displaystyle \|f\|_p \leq \|f\|_q^{q/p} M^{1-q/p} \hbox{ whenever } 0 \leq p \leq q \leq \infty.

There are two obvious algebraic identities involving norms which are worth knowing. The first is that one can interchange sums with integrals for any , in the sense that

\displaystyle \| (\sum_n |f_n|^p)^{1/p} \|_{L^p} = (\sum_n \|f_n\|_{L^p}^p)^{1/p};

this is just an application of the Fubini-Tonelli theorem. Secondly, exponents can pass through norms by changing the exponent: for any we have

\displaystyle \| |f|^q \|_{L^p} = \| f \|_{L^{pq}}^q.

We shall use both of these identities in the sequel without further comment.

Exercise 19 Establish the bound for any measurable and any .

— 2. Lorentz spaces —

Recall that the weak norm of a function is defined for as

\displaystyle \|f\|_{L^{p,\infty}(X,\mu)} := \sup_{\lambda > 0} \lambda \mu( \{ |f| \geq \lambda \} )^{1/p}.

Since

\displaystyle \|f\|_p^p = \int_X |f|^p\ d\mu \geq \int_X \lambda^p 1_{|f| \geq \lambda}\ d\mu = \lambda^p \mu( \{ |f| \geq \lambda \} )

for any and , we obtain Chebyshev’s inequality

\displaystyle \|f\|_{L^{p,\infty}(X,\mu)} \leq \|f\|_{L^{p}(X,\mu)}

(the case is also known as Markov’s inequality). When we adopt the convention that .

We define weak or to be the space of all functions with finite norm, with the usual abbreviations. We sometimes refer to as strong to distinguish it from weak .

Example 20 On a Euclidean space , the power function lies in weak if and only if . Indeed one can think of a weak function as a function which is pointwise dominated in magnitude by a rearrangement of (a multiple of) .

Suppose . From elementary calculus we have

\displaystyle |f(x)|^p = p \int_0^\infty 1_{|f(x)| \geq \lambda} \lambda^p \frac{d\lambda}{\lambda}

and hence on integration and Fubini’s theorem

\displaystyle \|f\|_p^p = p \int_0^\infty \mu( \{ |f| \geq \lambda \} ) \lambda^p \frac{d\lambda}{\lambda}.

To summarise, we have

\displaystyle \|f\|_{L^{p,\infty}} = \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^\infty({\bf R}^+,\frac{d\lambda}{\lambda})}

and

\displaystyle \|f\|_{L^p} = p^{1/p} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^p({\bf R}^+,\frac{d\lambda}{\lambda})}

for . These two identities motivate introducing the Lorentz (quasi-)norm for and by

\displaystyle \|f\|_{L^{p,q}(X,\mu)} := p^{1/q} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q({\bf R}^+,\frac{d\lambda}{\lambda})}. \ \ \ \ \ (11)

Thus for instance norm is identical to the norm. We shall abbreviate by , , or even when there is no chance of confusion.

Remark 21 For various reasons it is not worth trying to define Lorentz norms when , although we will use the convention . The most important values of , in descending order, are , , , and ; the other cases essentially never occur in applications.
Remark 22 The factor is inconsequential, but is traditional in order to maintain compatibility with the strong norm. But in practice the exact form of the Lorentz norm is not important; there are many formulations which are equivalent up to constants, and one generally just picks the formulation which is most convenient. The measure is of course multiplicative Haar measure on . One can interpret the equivalence of (i)-(iii) below by making the change of variables , so that the Haar measure just becomes Lebesgue measure in (modulo an inessential constant) and then passing from continuous to discrete .
Exercise 23 If is a monotone non-increasing function, show that (Depending on how you prove this, it may be convenient to first prove this for smoother , such as diffeomorphisms with strictly negative derivative, in order to apply an inverse function theorem.)
Exercise 24 Show that a step function of height and width has an norm of for any and .
Exercise 25 For any and show that

It is obvious that these norms are both rearrangement-invariant and monotone. To get a better intuitive handle on what the norm represents, we need some more definitions.

  • A sub-step function of height and width is any function supported on a set with the bounds almost everywhere and . (Thus .)
  • A quasi-step function of height and width is any function supported on a set with the bounds almost everywhere on , and . (Thus .)
Remark 26 It is a little dangerous to put fuzzy notation such as within a definition; if multiple quasi-step functions appear in an argument, the question then arises as to whether the implied constants are uniform. In our applications, the implied constants here are true absolute constants (like and ) so this will not be an issue.
Remark 27 From the binary expansion of the unit interval we see that a non-negative sub-step function of height and width can always be decomposed as where is an actual step function of height and width at most . By homogeneity we have a similar statement for other heights. Because of this, bounds on step functions tend to automatically extend to sub-step functions (and hence quasi-step functions) without difficulty.

Just like actual step functions, the norm of sub-step and quasi-step functions are well controlled; a sub-step function of height and width has norm , while a quasi-step function has norm – almost exactly like actual step functions. In the converse direction, it turns out that every function can be decomposed as an sum of “very different” -normalised sub-step or quasi-step functions.

Theorem 28 (Characterisation of ) Let be a function, let and , and let . Then the following are equivalent up to changes in the implied constants:
  • (i) We have .
  • (ii) There exists a decomposition where each is a quasi-step function of height and some width , with the having disjoint supports and Here the subscript in denotes the variable that the norm is being taken over.
  • (iii) There exists a pointwise bound , with
  • (iv) There exists a decomposition where each is a sub-step function of width and some height , with the having disjoint supports, the non-increasing in , the bounds on the support of , and
  • (v) There exists a pointwise bound , where and holds.
Remark 29 The formulations (ii), (iv) are useful when trying to use an bound on ; the formulations (iii), (v) are useful when trying to obtain an bound on . Heuristically, the above theorem is trying to say the following. If is a quasi-step function of height and width , then . But if is instead the sum of quasi-step functions of height and width , and either the heights or the widths are sufficiently variable in (e.g. one or the other grows like a power of two), then .

Proof: We may use homogeneity symmetry to normalise . The implications and are trivial. To see that (i) implies (ii), set and . (This is the “vertically dyadic layer cake decomposition”.) The only thing that requires nontrivial verification is ; but one easily verifies that

\displaystyle 2^m W_m^{1/p} \lesssim_{p,q} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([2^{m-2}, 2^{m-1}],\frac{d\lambda}{\lambda})}

and the claim follows by summing this in .

Similarly, to see that (i) implies (iv), define

\displaystyle H_n := \inf \{ \lambda: \mu( \{ |f| > \lambda \} ) \leq 2^{n-1} \};

note that this is a non-increasing function of , which goes to zero as (this comes from the hypothesis that is finite). We then define

\displaystyle f_n := f 1_{ H_n \geq |f| > H_{n+1} }.

(This is the “horizontally dyadic layer cake decomposition”.) The only non-trivial thing to verify is . But one easily verifies the telescoping estimate

\displaystyle \begin{array}{rl} H_n 2^{n/p} &= (H_n^q 2^{nq/p})^{1/q} \\ &= (\sum_{k=0}^\infty (H_{n+k}^q - H_{n+k+1}^q) 2^{nq/p} )^{1/q} \\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda 2^{(n+k)/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q}\\ &\lesssim_{p,q} (\sum_{k=0}^\infty 2^{-kq/p} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q([H_{n+k+1}, H_{n+k}],\frac{d\lambda}{\lambda})}^q)^{1/q} \end{array}

and the claim follows by summing this in and interchanging the summation signs. (We leave to the reader how to modify the above argument to handle the case .)

It remains to show that (iii) implies (i) and (iv) implies (i). Suppose first that (iii) holds. It is clear that for any we have

\displaystyle \mu( \{ |f| > 2^m \} ) \leq \sum_{k=0}^\infty \mu(E_{m+k})

and hence

\displaystyle \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q((2^m,2^{m+1}],\frac{d\lambda}{\lambda})} \lesssim_{p,q} 2^m (\sum_{k=0}^\infty \mu(E_{m+k}))^{1/p}

and so on taking summation in it would suffice to show that

\displaystyle \| 2^m (\sum_{k=0}^\infty \mu(E_{m+k}))^{1/p} \|_{\ell^q_m} \lesssim_{p,q} 1.

Raising to the power we rewrite as

\displaystyle \| \sum_{k=0}^\infty 2^{pm} \mu(E_{m+k}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 1.

But from the hypothesis we have

\displaystyle \| 2^{pm} \mu(E_{m}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 1

and hence on shifting by

\displaystyle \| 2^{pm} \mu(E_{m+k}) \|_{\ell^{q/p}_m} \lesssim_{p,q} 2^{-kp}.

The claim then follows from or .

Now suppose that (iv) holds. Observe that for any we have

\displaystyle \mu( \{ |f| > \lambda \} ) \lesssim_{p,q} \sup \{ 2^n: H'_n \geq \lambda \}

where are the modified heights

\displaystyle H'_n := \sum_{k=0}^\infty H_{n+k}.

Indeed, if for some then one easily verifies that and hence . The shifting trick and triangle inequality argument used previously shows that obeys the same bound as , thus

\displaystyle \| H'_n 2^{n/p} \|_{\ell^q_n({\bf Z})} \lesssim_{p,q} 1.

We now compute

\displaystyle \begin{array}{rl} \| \lambda \mu( \{ |f| \geq \lambda \} )^{1/p} \|_{L^q({\bf R}^+,\frac{d\lambda}{\lambda})}^q &\lesssim_{p,q} \int_0^\infty \lambda^{q-1} \sup \{ 2^{nq/p}: H'_n \geq \lambda \}\ d\lambda \\ &\lesssim_{p,q} \sum_n \int_0^\infty \lambda^{q-1} 2^{nq/p} 1_{H'_n \geq \lambda}\ d\lambda \\ &\sim_{p,q} \sum_n 2^{nq/p} (H'_n)^q \\ &\lesssim 1 \end{array}

as desired. (We leave to the reader how to modify the above argument to handle the case .)

Remark 30 For future reference we make the technical remark that if is a simple function, then in the horizontal and vertical decompositions in the above theorem, only finitely many of the are non-zero.
Remark 31 Suppose that the ratio between the tallest height and lowest non-zero height of a function is (i.e. there exists such that whenever is non-zero). Then the above theorem shows that two different Lorentz norms , with the same primary exponent only differ by multiplicative powers of . Similarly if the broadest width and narrowest width of a function differs by (e.g. if is equal to times the granularity of ). What this indicates is that the secondary exponent in the Lorentz norms only offers “logarithmic correction” to the Lebesgue norms ; in contrast, shows that varying the primary exponent leads to polynomial-strength changes in the norm. So as a first approximation (ignoring logarithms) one can pretend that . Note also that for quasi-step functions, the norms barely depend on at all.

One easy corollary of the above theorem is that the quasi-norm is indeed a quasi-norm, and in particular that for any ; this can be seen for instance by using the equivalence of (i) and (iii). Another easy consequence is that the simple functions are dense in .

Exercise 32
  • (i) Suppose that is a quasinorm on some function space with quasitriangle inequality constant , thus for all . Let be such that . Show that (Hint: first show that if , then . Then iterate this carefully to show that if and , then .)
  • (ii) Suppose that a sequence for obeys an exponential decay bound of the form for some and all . Show that the series converges in , with
Exercise 33 For each integer , let be a quasi-step function of height and width for some . Show that for all . If instead is a quasi-step function of height and width for some , show that for all . What goes wrong when we remove the absolute values on the ? (This can be repaired by replacing the powers of with powers of a sufficiently large constant (depending on the implied constant in the definition of a quasi-step function) – why?)

A particularly useful consequence of the above theorem is a Hölder inequality for Lorentz spaces, due to O’Neil.

Theorem 34 (Hölder’s inequality in Lorentz spaces) If and obey and then whenever the right-hand side norms are finite.

Proof: We may normalise , and drop the dependence of the implied constants on for brevity. By the equivalence of (i) and (v) in Theorem we may dominate and where and

\displaystyle \| H_n 2^{n/p_1} \|_{l^{q_1}_n}, \| H'_n 2^{n/p_2} \|_{l^{q_2}_n} \lesssim 1.

Then we have

\displaystyle |fg| \leq \sum_k \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}}.

By the quasi-triangle inequality and monotonicity it suffices to show that

\displaystyle \| \sum_{k \geq 0} \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 1

and

\displaystyle \| \sum_{k < 0} \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 1.

By symmetry it suffices to consider the component. Here we observe that has measure at most , so by the equivalence of (i) and (v) in Theorem

\displaystyle \| \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim \| H_n H'_{n+k} 2^{n/p} \|_{l^q_n}.

But by the ordinary Hölder inequality

\displaystyle \| H_n H'_{n+k} 2^{n/p} \|_{l^q_n} \leq \| H_n 2^{n/p_1} \|_{l^{q_1}_n} \| H'_{n+k} 2^{n/p_2} \|_{l^{q_2}_n};

shifting the second by we conclude

\displaystyle \| \sum_n H_n H'_{n+k} 1_{E_n \cap E'_{n+k}} \|_{L^{p,q}} \lesssim 2^{-k/p_2}.

The claim now follows from Exercise .

One corollary of this Hölder inequality is that functions are absolutely integrable on sets of finite measure whenever .

Now we consider dual formulations of the norms. The case is fairly straightforward:

Exercise 35 (Dual formulation of weak ) Let . Then for every in , we have Also show that the hypothesis can be dropped if one instead assumes to be non-negative.

The right-hand side of is clearly a semi-norm at least on . This leads in particular to a quasi-triangle inequality

\displaystyle \| f_1 + \ldots + f_N \|_{L^{p,\infty}(X,d\mu)} \sim_p \|f_1\|_{L^{p,\infty}(X,d\mu)} + \ldots + \|f_N\|_{L^{p,\infty}(X,d\mu)}

for any .

Remark 36 It is worth comparing to . In , one takes the inner product of against all -normalised functions, and the worst inner product becomes the norm. In , one only takes the inner product of against the -normalised step functions . This is fully consistent with the fact that the norm is stronger than the weak norm.

Exercise can be rephrased as follows: if for some and , then the following two statements are equivalent (up to changes in the implied constants):

  • .
  • for all sets of finite measure.

Unfortunately this equivalence breaks down at or below (consider for instance the weak function on , which is not even locally integrable when ). However, one does have a substitute:

Exercise 37 Let , , and . Show that the following are equivalent (up to changes in the implied constant):
  • .
  • For every set of finite measure, there exists a subset of with such that (in particular, we assert that the integral on the left-hand side is absolutely integrable).
Hint: It may be instructive to work out the example on by hand to get a sense of what is going on; this should suggest how to prove things in general. The proof is slightly simpler in the case when is non-negative, so you may want to try that case first. Comment on how this result implies Exercise (or its equivalent version discussed shortly afterwards) when .
Exercise 38 Let be functions with . Show that thus the weak quasinorm only fails to be a norm “by a logarithm”. Show with an example that the cannot be removed.

For more general spaces, we have

Theorem 39 (Dual characterisation of ) Let and . Then for any , Again, the hypothesis can be dropped if is non-negative and is restricted to also be non-negative.

Thus the quasi-norm is in fact equivalent to a norm when and . In particular, weak is equivalent to a normed space when . (For , weak fails to be normable “by a logarithm”; see Q3.) As with other dual characterisations, one can restrict to a dense subclass of , for instance simple functions with finite measure support.

Proof: To obtain the part of this theorem, we simply estimate

\displaystyle |\int_X f \overline{g}\ d\mu| \leq \|fg\|_{L^1} = \|fg\|_{L^{1,1}}

and use Theorem . To obtain the part, we normalise . It then suffices by homogeneity to find with and .

The case follows from , so let us take . We will just give the proof in the case ; the case is trickier, and a partial argument is given in the exercises. By the equivalence of (i) and (ii) in Theorem may write where is a quasi-step function of height and width with disjoint supports such that the sequence has an norm of . Now take

\displaystyle g := \sum_m g_m

where

\displaystyle g_m := a_m^{q-p} |f_m|^{p-2} f_m

adopting the obvious convention that when . Then (because of the disjoint supports)

\displaystyle |\int_X f \overline{g}| = \sum_m \int_X a_m^{q-p} |f_m|^p.

But since has height and width , and so

\displaystyle |\int_X f \overline{g}| \sim_p \sum_m a_m^q \sim_{p,q} 1.

To conclude it will suffice to show that

\displaystyle \|\sum_m g_m \|_{L^{p',q'}} \lesssim_{p,q} 1.

If is the support of , then we have the pointwise bound

\displaystyle g_m \lesssim_{p,q} a_m^{q-p} 2^{m(p-1)} 1_{E_m}

and the measure bound .

At this point we would like to apply Theorem , but neither the height nor width of is necessarily a power of . But we can remedy this by introducing the modified heights

\displaystyle H_m := \sup_{k \geq 0} a_{m-k}^{q-p} 2^{m(p-1)} 2^{-k(p-1)/2}.

We have , and so the increase geometrically. It then suffices to show that

\displaystyle \| \sum_m H_m 1_{E_m} \|_{L^{p',q'}} \lesssim_{p,q} 1.

By refining the by a constant factor we can make each at least twice as large as the previous, and so by applying the equivalences of (i) and (iii) in Theorem and the triangle inequality it suffices to show that

\displaystyle \| H_m \mu(E_m)^{1/p'} \|_{l^{q'}_m} \lesssim_{p,q} 1

which we expand using our bound on as

\displaystyle \| a_m^{p-1} \sup_{k \geq 0} a_{m-k}^{q-p} 2^{-k(p-1)/2} \|_{l^{q'}_m}.

But from Hölder’s inequality (using the hypothesis ) and the bound on we have

\displaystyle \| a_m^{p-1} a_{m-k}^{q-p} 2^{-k(p-1)/2} \|_{l^{q'}_m} \lesssim_{p,q} 2^{-k(p-1)/2};

summing this using the triangle inequality (and estimating the supremum by a sum) we obtain the claim.

The case when are restricted to be non-negative can be deduced from the above result and a monotone convergence argument (representing as a monotone limit of simple functions of finite measure support) which we leave as an exercise to the reader.

Exercise 40 A measure space is said to be non-atomic if, for every measurable set with , there exists a measurable subset such that .
  • (i) (Sierpinski’s theorem) Show that if is non-atomic, is measurable, and , then there exists a measurable subset of such that . (You may find it convenient to use Zorn’s lemma.)
  • (ii) Show that if is non-atomic, then the decomposition in Theorem (iv) can be chosen so that for each , either vanishes, or has support of measure (not just bounded by ).
  • (iii) Establish Theorem in the case that and is non-atomic, by using the modification of Theorem indicated in the previous part of the exercise.
The duality can also be established for measure spaces that contain atoms, but this requires a more careful analysis that treats large atoms separately.

— 3. Orlicz spaces (Optional) —

So far we have studied the Lebesgue spaces , together with the more general Lorentz spaces , which includes weak as a special case. These spaces are all rearrangement-invariant and monotone. There is a different generalisation of the Lebesgue spaces , the Orlicz spaces , which are also rearrangement-invariant and monotone, and which are occasionally useful. (There is a common generalisation of both, the Lorentz-Orlicz spaces, but these occur very rarely in applications.)

The motivation for Orlicz spaces starts with the trivial observation that if , then

\displaystyle \|f\|_{L^p} \leq 1 \hbox{ if and only if } \int_X |f|^p\ d\mu \leq 1.

Inspired by this, we generalise by letting be a function (with some additional properties to be selected shortly) and ask if we can find a norm which obeys the property

\displaystyle \|f\|_{\Phi(L)} \leq 1 \hbox{ if and only if } \int_X \Phi(|f|)\ d\mu \leq 1. \ \ \ \ \ (17)

Since norms need to be homogeneous, this would imply

\displaystyle \|f\|_{\Phi(L)} \leq A \hbox{ if and only if } \int_X \Phi(|f|/A)\ d\mu \leq 1

for all . In particular, if , then we need

\displaystyle \int_X \Phi(|f|/A)\ d\mu \leq 1 \hbox{ implies } \int_X \Phi(|f|/A')\ d\mu \leq 1.

To ensure this property it is thus very natural to require that be increasing. Also to deal with the zero norm case one typically requires .

Next, in order for to be a norm, the unit ball needs to be convex. Looking at , we see that this will indeed be the case when is itself convex. (Note that the proof of was a special case of this argument).

We can put all the above discussion together and conclude: if is increasing and convex with , then the norm

\displaystyle \| f \|_{\Phi(L)} := \inf \{ A > 0: \int_X \Phi(|f|/A)\ d\mu \leq 1 \}

is a norm on the space .

As discussed above, the spaces for are examples of Orlicz spaces with . The space is not really an Orlicz space, but can be viewed the limiting case where is infinite for and zero for (or more informally, ). Aside from the Lebesgue spaces, the most common Orlicz spaces which appear are

  • The space , defined as the Orlicz space with ;
  • The space , defined as the Orlicz space with ;
  • The space , defined as the Orlicz space with .

The correction factors of and in the above functions should not be taken too seriously; note that if two functions are comparable then their Orlicz norms are comparable also; a little more generally, if , then . It is the behaviour of for large values of which is the most important, although when has infinite measure the behaviour at small values of is also relevant.

Problem 41 If has finite measure, verify the relation which is another indication of the irrelevance of the low values of in the finite measure case.

The final fact about Orlicz spaces that we give here is the duality relation. Suppose that is increasing, convex, and is also superlinear in the sense that . We can then define the Young dual of by the formula

\displaystyle \Psi(y) := \sup \{ xy - \Phi(x): x \in {\bf R}^+ \};

the hypothesis that is superlinear ensures that this function is well-defined. We may equivalently define to be the smallest function for which one has the inequality

\displaystyle xy \leq \Phi(x) + \Psi(y) \hbox{ for all } x, y \in {\bf R}^+. \ \ \ \ \ (18)
Exercise 42 If , show that the Young dual of is . Show also that the Young dual of takes the form for . What does the Young duals of and look like?

One can easily verify that is also increasing, convex, and super-linear, and so the Orlicz norm makes sense. From and the triangle inequality it is immediate that

\displaystyle |\int_X f \overline{g}\ d\mu| \leq 2 \hbox{ whenever } \|f\|_{\Phi(L)}, \|g\|_{\Psi(L)} \leq 1

and hence by homogeneity we obtain the duality relation

\displaystyle |\int_X f \overline{g}\ d\mu| \leq 2 \|f\|_{\Phi(L)} \|g\|_{\Psi(L)}

whenever and .

Exercise 43 Establish the more precise relationship
Exercise 44 Show that if is the Young dual of , then is the Young dual of . (It may help to view things geometrically, and in particular understanding as parameterising the support lines of the graph of the convex function .)
Exercise 45 When has finite measure, show that the spaces and are dual to each other. What is the dual to ?
Exercise 46 Let be a function on a measure space of bounded measure . Show that the following are equivalent (up to changes in the implied constants):
  • (i) .
  • (ii) for all .
  • (iii) for all .
Hint: You may find the Taylor expansion for , together with the obvious bounds for integer (or Stirling’s formula, if you know what that is), to be useful.
Exercise 47 Obtain the analogue of Exercise for the Orlicz space .

— 4. Real interpolation —

So far we have only considered functions on a single measure space . Now we shall consider operators which take functions on one measure space to functions on another measure space ; the study of such operators is in fact a major focus of harmonic analysis. Ultimately we want to extend to a standard normed vector space such as , but in practice one has to initially first restrict attention to a dense subspace of functions, such as simple functions or test functions.

We are primarily interested in linear operators, thus and . But it is also worth considering the more general sublinear operators, in which

\displaystyle |T(cf)| = |c| |Tf|

and we have the pointwise estimate

\displaystyle |T(f+g)| \leq |Tf| + |Tg|.

Apart from the linear operators, the next most important example of a sublinear operator is a maximal operator

\displaystyle Tf(x) := \sup_n |T_n f(x)|

where are a collection (possibly countably or uncountably infinite, though in the latter case one has to take some care in ensuring measurability) of linear or sublinear operators. The third most important example is a square function such as

\displaystyle Tf(x) := (\sum_n |T_n f(x)|^2)^{1/2}.

More generally, one can consider a family of operators indexed by some parameter , and take to be the norm in the variable of in some suitable norm. But the above three examples of linear operators, maximal operators, and square functions already cover the vast majority of applications.

Let be exponents, and let be sublinear. Let us define the following concepts.

  • We say that is strong-type (or simply type ) if we have a bound for all in , or in a dense sub-class thereof. Note that in the latter case there is a unique extension to all of .
  • If , we say that is weak-type if we have a bound
  • We say that is restricted strong-type if we have a bound for all sub-step functions of height and width . In particular, we have (Conversely, we can deduce from using tricks such as those in Remark .)
  • If , we say that is restricted weak-type if we have a bound for all sub-step functions of height and width . In particular, we have

Clearly, whenever are fixed, strong-type implies weak-type and restricted strong-type, either of which imply restricted weak-type. In most applications, it is the strong-type bounds which are desired; however, we shall see in this section that the real interpolation method allows us to deduce strong-type bounds from weak-type, or even restricted weak-type bounds, as long as the strong-type bounds are an interpolant between the restricted weak-type bounds. This can be a very useful strategy, because (as we shall see in next week’s notes) weak-type or restricted weak-type bounds are easier to prove than strong-type estimate.

Let us first make a mild (and qualitative) assumption, namely that the form is well-defined whenever are simple functions with finite measure support. This is for instance the case if is of restricted type for some and , or restricted weak-type for some and ; thus in practice this assumption is easily satisfied. We observe that this form is non-negative, homogeneous and sublinear in both

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论