MathLabs

Analysis

Harmonic analysis

Decomposes functions and signals into basic waves, generalizing Fourier series to broader settings.

IntuitionFrom waves to frequencies

Any sound you hear — a violin note, a voice, traffic noise — is really a single wiggling pressure signal over time. Yet your ear (and a graphic equalizer) can tell you it is made of many pure tones at different pitches. Harmonic analysis is the mathematics of that splitting: it takes a function ff and rewrites it as a sum or integral of pure oscillations e2πixξe^{2\pi i x \xi}, each with its own frequency ξ\xi. Fourier series does this for periodic signals using a discrete list of frequencies; harmonic analysis extends the idea to non-periodic signals, higher dimensions, and even other groups (the circle, finite groups, Lie groups).

Interactive Fourier series widget showing a wave built from a variable number of harmonics.
A square-ish wave rebuilt from its first 24 Fourier harmonics; drag n to watch the Gibbs 'ringing' near the jump shrink in width but not in height.

UndergraduateFourier series on the circle: the convergence question

Write the NN-th partial sum of the Fourier series of a periodic function ff as SNf(x)=∑∣n∣≤Nf^(n) einxS_N f(x) = \sum_{|n|\le N} \hat f(n)\, e^{inx}, where f^(n)=12π∫−ππf(x) e−inx dx\hat f(n) = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x)\, e^{-inx}\,dx. In L2L^2 everything is clean: the exponentials einxe^{inx} form an orthonormal basis, Parseval's identity 12π∫∣f∣2=∑∣f^(n)∣2\frac{1}{2\pi}\int |f|^2 = \sum |\hat f(n)|^2 holds, and SNf→fS_N f \to f in the L2L^2 norm. The hard question is pointwise convergence: does SNf(x)→f(x)S_N f(x) \to f(x) at a specific point xx? The square wave shown above already hints at the trouble — near its jump, SNfS_N f always overshoots by about 9% of the jump height, no matter how large NN is (the Gibbs phenomenon, with the overshoot tending to 2πSi(π)≈1.179\frac{2}{\pi}\mathrm{Si}(\pi) \approx 1.179 times the half-jump).

Definition: Dirichlet kernel

The partial sum is itself a convolution: SNf(x)=(DN∗f)(x)=12π∫−ππDN(x−y) f(y) dyS_N f(x) = (D_N * f)(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} D_N(x-y)\, f(y)\, dy, where the Dirichlet kernel is DN(x)=∑∣n∣≤Neinx=sin⁡((N+12)x)sin⁡(x/2)D_N(x) = \sum_{|n|\le N} e^{inx} = \dfrac{\sin\left((N+\tfrac12)x\right)}{\sin(x/2)}. Every convergence question is therefore really a question about the shape of DND_N.

DN(x)=∑∣n∣≤Neinx=sin⁡((N+12)x)sin⁡(x/2)D_N(x) = \sum_{|n|\le N} e^{inx} = \dfrac{\sin\left((N+\tfrac12)x\right)}{\sin(x/2)}
Interactive chart of the Dirichlet, Fejér, or Poisson kernel over one period, with a live estimate of its L1 norm.
The Dirichlet kernel DND_N for N=9N=9: unlike Fejér's or Poisson's kernel, it dips below zero (the coral lobes), and its L1L^1 norm — the Lebesgue constant — grows like (4/π2)log⁡N(4/\pi^2)\log N instead of staying bounded. Switch kernels or raise NN to compare.

The Fejér kernel FNF_N (the Cesàro average of D0,…,DND_0,\ldots,D_N) and the Poisson kernel PrP_r are both nonnegative, so they are honest approximate identities: mass 11, L1L^1 norm 11, and mass concentrating at 00 as N→∞N\to\infty (or r→1r\to1) — enough to force FN∗f→fF_N * f\to f and Pr∗f→fP_r * f\to f uniformly for every continuous ff. The Dirichlet kernel is not: its negative lobes make 12π∫∣DN∣\frac{1}{2\pi}\int|D_N| (the Lebesgue constant) grow like 4π2log⁡N\frac{4}{\pi^2}\log N, unbounded as N→∞N\to\infty. By the Banach–Steinhaus uniform boundedness principle, unbounded Lebesgue constants force the existence of a continuous function whose Fourier series diverges at a point — exactly the phenomenon du Bois-Reymond exhibited by hand in 1873, obtained here for free. Kolmogorov later showed the failure can be far worse for merely integrable functions: in 1923 he built an L1L^1 function whose Fourier series diverges almost everywhere, and in 1926 one that diverges everywhere.

If ff is continuous and 2π2\pi-periodic, then the Cesàro means σNf=1N+1∑k=0NSkf=FN∗f\sigma_N f = \frac{1}{N+1}\sum_{k=0}^N S_k f = F_N * f converge to ff uniformly as N→∞N\to\infty.

Why is it true?

Averaging the partial sums cancels exactly the oscillation that makes SNfS_N f misbehave: while DND_N has wild negative lobes, its running average FNF_N is a smooth, nonnegative bump that only ever adds up pieces of ff with a positive weight. So even though SNfS_N f can fail to converge, its running average always settles down — the same trick that tames a jittery sequence of partial sums in calculus by looking at its Cesàro average instead.

Proof

Step 1 (Fejér kernel is a positive approximate identity). Averaging D0,…,DND_0,\ldots,D_N and using a telescoping trigonometric identity gives the closed form FN(x)=1N+1(sin⁡(N+12x)sin⁡(x/2))2≥0F_N(x) = \frac{1}{N+1}\left(\dfrac{\sin\left(\tfrac{N+1}{2}x\right)}{\sin(x/2)}\right)^2 \ge 0. Since FN=1N+1∑k=0NDkF_N = \frac{1}{N+1}\sum_{k=0}^N D_k and 12π∫Dk=1\frac{1}{2\pi}\int D_k = 1 for every kk, also 12π∫FN=1\frac{1}{2\pi}\int F_N = 1; being nonnegative, 12π∫∣FN∣=1\frac{1}{2\pi}\int |F_N| = 1 too, so the L1L^1 norm never blows up.

Step 2 (mass concentrates at 00). For any fixed δ>0\delta>0, on δ≤∣x∣≤π\delta \le |x| \le \pi the denominator sin⁡(x/2)2\sin(x/2)^2 is bounded below by a positive constant, so FN(x)=O(1/N)F_N(x) = O(1/N) uniformly there; the mass outside [−δ,δ][-\delta,\delta] tends to 00 as N→∞N\to\infty.

Step 3 (approximate-identity estimate). Write σNf(x)−f(x)=12π∫−ππFN(y)(f(x−y)−f(x)) dy\sigma_N f(x) - f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} F_N(y)\big(f(x-y)-f(x)\big)\,dy using 12π∫FN=1\frac{1}{2\pi}\int F_N = 1. Split the integral into ∣y∣<δ|y|<\delta and δ≤∣y∣≤π\delta\le|y|\le\pi. On the first piece, uniform continuity of ff makes ∣f(x−y)−f(x)∣<ε|f(x-y)-f(x)|<\varepsilon for δ\delta small, and FN≥0F_N\ge0 with total mass 11 bounds that piece by ε\varepsilon. On the second piece, FN=O(1/N)F_N=O(1/N) and ff is bounded, so that piece →0\to 0 as N→∞N\to\infty. Both bounds are uniform in xx, so σNf→f\sigma_N f\to f uniformly.

If f∈L2(T)f\in L^2(\mathbb{T}), then SNf(x)→f(x)S_N f(x)\to f(x) for almost every xx. (Hunt, 1968, extended this to every f∈Lp(T)f\in L^p(\mathbb{T}) with p>1p>1.)

Why is it true?

Given Kolmogorov's 1926 example of an L1L^1 function that diverges everywhere, it seemed entirely plausible that L2L^2 — barely a stronger condition — would fail the same way, or at least almost everywhere. Carleson's theorem is a genuine surprise: the tiny extra assumption of square-integrability rules out divergence everywhere except a set of measure zero. It closed a problem that had been open since Luzin conjectured it in 1913, and is regarded as one of the deepest theorems of twentieth-century analysis.

Proof

A full proof is one of the hardest arguments in twentieth-century analysis; Fefferman gave a celebrated simplification in 1973 whose strategy can be sketched in three moves.

Step 1 (control by a maximal operator). It suffices to bound the Carleson maximal operator S∗f(x)=sup⁡N∣SNf(x)∣S^*f(x) = \sup_N |S_N f(x)| on L2L^2, since a weak-type bound on S∗S^* plus density of nice functions (for which convergence is easy) implies the full a.e. convergence statement.

Step 2 (time–frequency decomposition into tiles). Decompose ff using wave packets adapted to dyadic tiles in the time–frequency plane — rectangles of area ∼1\sim 1, exactly the atoms discussed in the 'Time and frequency' section below. Each tile carries a piece of ff localized to both a time interval and a frequency band.

Step 3 (organize tiles into trees and sum). Because the cutoff frequency NN can vary with xx, tiles relevant to S∗f(x)S^*f(x) form 'trees' ordered by time-frequency containment. Fefferman's key combinatorial estimate bounds the total energy carried by all trees, controlling S∗fS^*f in L2L^2 and completing the proof.

The technique that finally cracked Carleson's theorem — organizing time–frequency tiles into trees — turns out to be a general-purpose tool used elsewhere in modern harmonic analysis (for instance, the same style of argument settled a Calderón open problem about the bilinear Hilbert transform in 1997, via Lacey and Thiele).

UndergraduateThe Fourier transform on the real line

Definition: Fourier transform

For an integrable function ff on R\mathbb{R}, its Fourier transform f^\hat f records how much of each pure frequency ξ\xi is present in ff. It is defined by f^(ξ)=∫−∞∞f(x) e−2πixξ dx\hat f(\xi) = \int_{-\infty}^{\infty} f(x)\, e^{-2\pi i x \xi}\,dx, where the exponential e−2πixξe^{-2\pi i x \xi} is a unit-length spinning vector; multiplying f(x)f(x) by it and integrating measures the correlation of ff with that spin rate.

f^(ξ)=∫−∞∞f(x) e−2πixξ dx\hat f(\xi) = \int_{-\infty}^{\infty} f(x)\, e^{-2\pi i x \xi}\,dx

When f^\hat f is itself integrable, the original signal can be rebuilt exactly by summing all those pure frequencies back up: f(x)=∫−∞∞f^(ξ) e2πixξ dξf(x) = \int_{-\infty}^{\infty} \hat f(\xi)\, e^{2\pi i x \xi}\,d\xi. This inversion formula is the precise sense in which 'a function equals the sum of its frequency components.'

f(x)=∫−∞∞f^(ξ) e2πixξ dξf(x) = \int_{-\infty}^{\infty} \hat f(\xi)\, e^{2\pi i x \xi}\,d\xi
Three faces of the same idea
SettingDomain of the signalFrequency sideKey identity
Fourier seriesCircle / periodIntegers Z\mathbb{Z}Discrete sum of harmonics
Fourier transformReal line R\mathbb{R}Real line R\mathbb{R}f(x)=∫−∞∞f^(ξ) e2πixξ dξf(x) = \int_{-\infty}^{\infty} \hat f(\xi)\, e^{2\pi i x \xi}\,d\xi
Poisson summationReal line, sampled at Z\mathbb{Z}Integers Z\mathbb{Z}∑n=−∞∞f(n)=∑k=−∞∞f^(k)\sum_{n=-\infty}^{\infty} f(n) = \sum_{k=-\infty}^{\infty} \hat f(k)

AdvancedKey theorems

If f∈L1(R)∩L2(R)f \in L^1(\mathbb{R}) \cap L^2(\mathbb{R}), then ∫−∞∞∣f(x)∣2 dx=∫−∞∞∣f^(ξ)∣2 dξ\int_{-\infty}^{\infty} |f(x)|^2\,dx = \int_{-\infty}^{\infty} |\hat f(\xi)|^2\,d\xi. In words: the Fourier transform preserves total energy, with no constant needed in this normalization.

Why is it true?

Physically, ∫∣f∣2\int |f|^2 is the energy of a signal (think: power dissipated by a voltage waveform). Plancherel says you can compute that energy either by scanning the signal in time or by scanning its spectrum in frequency — a spectrum analyzer and an oscilloscope must agree on total power.

Proof

Step 1 (a special case). First check the identity for a Gaussian f(x)=e−πx2f(x) = e^{-\pi x^2}. A direct computation (completing the square in the exponent) gives f^(ξ)=e−πξ2\hat f(\xi) = e^{-\pi \xi^2}, so both sides of ∫−∞∞∣f(x)∣2 dx=∫−∞∞∣f^(ξ)∣2 dξ\int_{-\infty}^{\infty} |f(x)|^2\,dx = \int_{-\infty}^{\infty} |\hat f(\xi)|^2\,d\xi equal the same Gaussian integral ∫e−2πx2 dx\int e^{-2\pi x^2}\,dx.

Step 2 (extend by linearity and the convolution theorem). For nice (Schwartz) functions ff, write g=f∗f~g = f * \tilde f where f~(x)=f(−x)‾\tilde f(x) = \overline{f(-x)}; then g(0)=∫∣f(x)∣2 dxg(0) = \int |f(x)|^2\,dx, and by the convolution theorem f∗g^=f^⋅g^\widehat{f * g} = \hat f \cdot \hat g together with Fourier inversion, g(0)=∫g^(ξ) dξ=∫∣f^(ξ)∣2 dξg(0) = \int \hat g(\xi)\,d\xi = \int |\hat f(\xi)|^2\,d\xi. This proves the identity for all Schwartz functions, which include Gaussians and all smooth rapidly-decaying functions.

Step 3 (density argument). Schwartz functions are dense in L2(R)L^2(\mathbb{R}): every f∈L2f \in L^2 is a limit fn→ff_n \to f of Schwartz functions in the L2L^2 norm. Since the Fourier transform is an isometry on this dense subspace by Step 2, it extends uniquely to a bounded operator on all of L2(R)L^2(\mathbb{R}) that still satisfies ∫−∞∞∣f(x)∣2 dx=∫−∞∞∣f^(ξ)∣2 dξ\int_{-\infty}^{\infty} |f(x)|^2\,dx = \int_{-\infty}^{\infty} |\hat f(\xi)|^2\,d\xi — this extension is what 'the Fourier transform of an L2L^2 function' means when the defining integral f^(ξ)=∫−∞∞f(x) e−2πixξ dx\hat f(\xi) = \int_{-\infty}^{\infty} f(x)\, e^{-2\pi i x \xi}\,dx may not converge absolutely.

For a Schwartz function ff, ∑n=−∞∞f(n)=∑k=−∞∞f^(k)\sum_{n=-\infty}^{\infty} f(n) = \sum_{k=-\infty}^{\infty} \hat f(k): summing the function over all integers equals summing its Fourier transform over all integers.

Why is it true?

It is the bridge between sampling a signal at integer points and periodizing its spectrum: the left side is what you get by adding up samples f(n)f(n), the right side is what you get from the spectrum. This single identity underlies the sampling theorem, lattice sums in crystallography, and the functional equation of theta functions and the Riemann zeta function.

Proof

Step 1 (periodize). For a Schwartz function ff, define F(x)=∑n=−∞∞f(x+n)F(x) = \sum_{n=-\infty}^{\infty} f(x+n). Rapid decay of ff makes this sum converge absolutely and uniformly, and FF is a smooth function with period 1, so it has its own Fourier series on the circle.

Step 2 (compute the Fourier coefficients of the periodization). The kk-th Fourier coefficient of FF is ck=∫01F(x)e−2πikx dx=∑n∫01f(x+n)e−2πikx dxc_k = \int_0^1 F(x) e^{-2\pi i k x}\,dx = \sum_{n} \int_0^1 f(x+n) e^{-2\pi i k x}\,dx. Substituting y=x+ny = x+n in each term and using periodicity of e−2πikxe^{-2\pi i k x} in nn reassembles the pieces into a single integral over all of R\mathbb{R}: ck=∫−∞∞f(y)e−2πiky dy=f^(k)c_k = \int_{-\infty}^{\infty} f(y) e^{-2\pi i k y}\,dy = \hat f(k).

Step 3 (evaluate at x=0). Since FF is smooth, its Fourier series converges to it pointwise, in particular at x=0x=0: F(0)=∑kcke0F(0) = \sum_k c_k e^{0}, i.e. ∑n=−∞∞f(n)=∑k=−∞∞f^(k)\sum_{n=-\infty}^{\infty} f(n) = \sum_{k=-\infty}^{\infty} \hat f(k). The left side is literally F(0)=∑nf(n)F(0) = \sum_n f(n) by definition, which finishes the proof.

AdvancedThe Hilbert transform: the first singular integral

Convolving against a kernel that is merely bounded and integrable — like the Poisson or Fejér kernel above — is tame. Some of the most important operators in analysis instead convolve against a kernel with a genuine singularity at the origin, not integrable near y=xy=x and redeemed only by a delicate cancellation between its positive and negative parts. The prototype is the Hilbert transform on the real line, Hf(x)=p.v. 1π∫−∞∞f(y)x−y dyHf(x) = \text{p.v.}\,\frac{1}{\pi}\int_{-\infty}^{\infty}\frac{f(y)}{x-y}\,dy, understood as a principal value: the singularity at y=xy=x is excised symmetrically before the limit is taken. In frequency, this singular convolution turns into strikingly simple multiplication: (Hf^)(ξ)=−i sgn(ξ) f^(ξ)(\widehat{Hf})(\xi) = -i\,\mathrm{sgn}(\xi)\,\hat f(\xi) — the Hilbert transform is a pure 90∘90^\circ phase rotation of every frequency, flipping the sign of ii without changing magnitude.

Hf(x)=p.v. 1π∫−∞∞f(y)x−y dyHf(x) = \text{p.v.}\,\frac{1}{\pi}\int_{-\infty}^{\infty}\frac{f(y)}{x-y}\,dy
(Hf^)(ξ)=−i sgn(ξ) f^(ξ)(\widehat{Hf})(\xi) = -i\,\mathrm{sgn}(\xi)\,\hat f(\xi)

There is a second, classical way to see this: for ff defined on the real line, let u=Py∗fu = P_y * f be its harmonic extension to the upper half-plane (convolution with the Poisson kernel from the earlier section), and let v=Qy∗fv = Q_y * f be its harmonic conjugate, built from the conjugate Poisson kernel QyQ_y so that F=u+ivF=u+iv is a holomorphic function of x+iyx+iy. As y→0+y\to 0^+, the boundary value of vv recovers exactly the Hilbert transform of ff. This is the complex-analytic ancestor of the whole theory: a singular integral operator on the boundary is really the shadow of an honest holomorphic function living one dimension up.

The Hilbert transform HH is bounded on Lp(R)L^p(\mathbb{R}) for every 1<p<∞1<p<\infty: there is a constant CpC_p so that ∥Hf∥Lp≤Cp∥f∥Lp\|Hf\|_{L^p} \le C_p\|f\|_{L^p} for every ff. (Marcel Riesz, 1927.)

Why is it true?

The L2L^2 case is essentially free: the multiplier −i sgn(ξ)-i\,\mathrm{sgn}(\xi) has modulus exactly 11 almost everywhere, so HH is literally an isometry on L2L^2 by Plancherel — a 90∘90^\circ phase rotation of every frequency changes nothing about total energy. But extending boundedness to other pp is a genuinely different and much harder problem: there is no analogue of Parseval's identity outside L2L^2, so no algebraic shortcut is available, and the proof has to fall back on real-variable estimates about the size and geometry of the set where HfHf is large.

Proof

Step 1 (L2L^2, directly). By Plancherel and the multiplier formula, ∥Hf∥L22=∫∣Hf^(ξ)∣2 dξ=∫∣sgn(ξ)∣2∣f^(ξ)∣2 dξ=∥f∥L22\|Hf\|_{L^2}^2 = \int |\widehat{Hf}(\xi)|^2\,d\xi = \int |\mathrm{sgn}(\xi)|^2|\hat f(\xi)|^2\,d\xi = \|f\|_{L^2}^2, using ∣sgn(ξ)∣=1|\mathrm{sgn}(\xi)|=1 for ξ≠0\xi\ne0 (a single point does not affect the integral). So HH is an isometry on L2(R)L^2(\mathbb{R}), in particular bounded with C2=1C_2=1.

Step 2 (general 1<p<∞1<p<\infty, forward reference). The full range of exponents does not follow from Step 1 by any soft argument. It is a genuine theorem of real-variable Calderón–Zygmund theory: the Hilbert transform's kernel 1/(πx)1/(\pi x) satisfies exactly the smoothness and cancellation conditions of a Calderón–Zygmund kernel, so the general singular-integral machinery developed in the next section applies to it. That machinery produces a weak-type (1,1)(1,1) bound from the L2L^2 bound proved in Step 1, and Marcinkiewicz interpolation between the weak-(1,1)(1,1) bound and the L2L^2 bound (then a duality argument for p>2p>2) completes the proof for every 1<p<∞1<p<\infty — this half of the argument is carried out in full in the Calderón–Zygmund section immediately below, rather than repeated here.

Example: The Hilbert transform of an indicator function

Let f=1[−1,1]f = \mathbf{1}_{[-1,1]}, the indicator of [−1,1][-1,1]: a bounded, compactly supported, utterly unremarkable function. Compute HfHf.

Solution

By definition, Hf(x)=1π p.v.∫−11dyx−yHf(x) = \frac{1}{\pi}\,\text{p.v.}\int_{-1}^{1} \frac{dy}{x-y}. Away from [−1,1][-1,1] there is no singularity to excise, and the antiderivative of 1/(x−y)1/(x-y) in yy is −log⁡∣x−y∣-\log|x-y|, so ∫−11dyx−y=[−log⁡∣x−y∣]y=−1y=1=log⁡∣x+1x−1∣\int_{-1}^{1} \frac{dy}{x-y} = \big[-\log|x-y|\big]_{y=-1}^{y=1} = \log\left|\frac{x+1}{x-1}\right| (for xx inside (−1,1)(-1,1) the same computation goes through as a genuine principal value, since the two divergences at y=x−y=x^- and y=x+y=x^+ cancel). This gives Hf(x)=1πlog⁡∣x+1x−1∣Hf(x) = \frac{1}{\pi}\log\left|\frac{x+1}{x-1}\right|.

The striking feature is what happens at x=±1x=\pm1: ff itself is perfectly bounded there (it simply jumps from 11 to 00), yet HfHf blows up logarithmically at exactly those two points. A bounded, compactly supported, textbook-nice function is turned by HH into a function with two genuine singularities — this is the local signature of every singular integral operator, and it is exactly why controlling HH requires more delicate estimates than controlling a convolution with a nice bounded kernel.

AdvancedThe Hardy–Littlewood maximal function and the Calderón–Zygmund decomposition

To finish M. Riesz's theorem for a general singular integral operator — not just the Hilbert transform — analysts needed a purely real-variable tool that says nothing about Fourier transforms at all. Hardy and Littlewood introduced it in 1930: the maximal function Mf(x)=sup⁡r>012r∫∣y−x∣<r∣f(y)∣ dyMf(x) = \sup_{r>0} \frac{1}{2r}\int_{|y-x|<r} |f(y)|\,dy, the largest possible average of ∣f∣|f| over any interval centered at xx. It is a blunt but extremely effective instrument for controlling every averaging process that could ever be applied to ff near xx, all at once.

MM is of weak type (1,1)(1,1): ∣{x:Mf(x)>λ}∣≤Cλ∥f∥1|\{x : Mf(x) > \lambda\}| \le \frac{C}{\lambda}\|f\|_1 for a universal constant CC (one can always take C=5C=5). Consequently, by interpolation, MM is also bounded on Lp(R)L^p(\mathbb{R}) for every 1<p≤∞1<p\le\infty — but MM itself is never bounded on L1(R)L^1(\mathbb{R}).

Why is it true?

Mf(x)Mf(x) bounds, in one stroke, every average of ff that could ever be taken over an interval around xx — the running mean at every possible scale. Controlling that single quantity turns out to be exactly what is needed to prove the Lebesgue differentiation theorem (the average of ff over shrinking intervals around xx converges to f(x)f(x) for almost every xx): once MfMf is known to be finite almost everywhere, a short soft argument upgrades that to the full differentiation statement. This is the real-variable engine behind the Calderón–Zygmund theory below, playing the role that Plancherel's identity played for the easy L2L^2 estimate.

Proof

Step 1 (a cover by good balls). Fix λ>0\lambda>0. For every xx with Mf(x)>λMf(x)>\lambda, by definition of the supremum there is some interval BxB_x centered at xx with 1∣Bx∣∫Bx∣f∣>λ\frac{1}{|B_x|}\int_{B_x}|f| > \lambda. These balls {Bx}\{B_x\} cover the set {Mf>λ}\{Mf>\lambda\}.

Step 2 (Vitali 5r5r-covering lemma). From any collection of balls of bounded radius, one can always extract a countable disjoint subcollection {Bi}\{B_i\} such that the 55-times dilated balls {5Bi}\{5B_i\} still cover the union of the whole original collection. Apply this to {Bx}\{B_x\} to get a disjoint subfamily {Bi}\{B_i\} with {Mf>λ}⊆⋃i5Bi\{Mf>\lambda\}\subseteq\bigcup_i 5B_i.

Step 3 (sum the disjoint pieces). Each BiB_i satisfies λ∣Bi∣<∫Bi∣f∣\lambda|B_i| < \int_{B_i}|f| by construction. Since the BiB_i are pairwise disjoint, summing over ii gives λ∑i∣Bi∣<∑i∫Bi∣f∣≤∫R∣f∣=∥f∥1\lambda\sum_i|B_i| < \sum_i\int_{B_i}|f| \le \int_{\mathbb{R}}|f| = \|f\|_1, so ∑i∣Bi∣<∥f∥1/λ\sum_i|B_i| < \|f\|_1/\lambda. Finally ∣{Mf>λ}∣≤∑i∣5Bi∣=5∑i∣Bi∣<5λ∥f∥1|\{Mf>\lambda\}| \le \sum_i|5B_i| = 5\sum_i|B_i| < \dfrac{5}{\lambda}\|f\|_1, which is exactly ∣{x:Mf(x)>λ}∣≤5λ∥f∥1|\{x : Mf(x) > \lambda\}| \le \frac{5}{\lambda}\|f\|_1.

Definition: Calderón–Zygmund decomposition

Given f∈L1(R)f\in L^1(\mathbb{R}) and a height α>0\alpha>0, Calderón and Zygmund (1952) showed how to split the line into disjoint 'stopping' dyadic intervals adapted to ff and α\alpha: start from one large dyadic interval, and recursively bisect it, stopping and keeping a dyadic interval QjQ_j the very first time its average 1∣Qj∣∫Qj∣f∣\frac{1}{|Q_j|}\int_{Q_j}|f| exceeds α\alpha (its parent's average was still ≤α\le\alpha, so bisecting can at most double it). This produces a disjoint family of stopping intervals {Qj}\{Q_j\} with α<1∣Qj∣∫Qj∣f∣≤2α\alpha < \frac{1}{|Q_j|}\int_{Q_j}|f| \le 2\alpha, with total length ∑∣Qj∣≤∥f∥1/α\sum|Q_j|\le \|f\|_1/\alpha, and with ∣f∣≤α|f|\le\alpha almost everywhere on what remains outside ⋃jQj\bigcup_j Q_j.

Interactive chart of a fixed function's Calderón–Zygmund decomposition on [0,1], with an adjustable level slider α showing the stopping intervals, their averages, and reference lines at α and 2α.
A fixed, wiggly |f| on [0,1] (with a narrow spike near x≈0.18, a near-singular bump near x≈0.62, and a broad bump near x≈0.86), with its Calderón–Zygmund stopping intervals at this level shaded in red: each shaded interval is one stopping cube Q_j, and the thick red segment drawn across it marks its average — always caught between the α and 2α dashed reference lines. Drag the slider to change α and watch the stopping intervals shrink, grow, split, or merge.

The decomposition splits ff into a 'good' part and infinitely many 'bad' pieces, f=g+∑jbjf = g + \sum_j b_j: the good part gg equals ff outside all the QjQ_j and equals the (roughly α\alpha-sized) average of ff on each QjQ_j where it is kept, so ∣g∣≤2α|g|\le 2\alpha everywhere and gg is handled by the easy L2L^2 theory. Each bad piece bjb_j is supported on its own cube QjQ_j and has mean zero there, ∫Qjbj dx=0\int_{Q_j} b_j\,dx = 0 — that cancellation is exactly what lets a smooth kernel's contribution from bjb_j stay small once you are far enough from QjQ_j, since the kernel looks nearly constant across QjQ_j and a nearly-constant kernel integrates a mean-zero function to almost nothing. Combining this good–bad split with the weak-type (1,1)(1,1) Hardy–Littlewood maximal bound, through Marcinkiewicz interpolation, is exactly the real-variable machinery Calderón and Zygmund used to finish M. Riesz's theorem for general singular integral operators — not just the Hilbert transform — extending LpL^p boundedness to every 1<p<∞1<p<\infty, closing the loop back to the theorem above.

UndergraduateReal-World Applications and Worked Examples

Because the Fourier transform turns 'shape in time' into 'content in frequency,' it is the standard tool wherever engineers or scientists need to isolate, filter, or count oscillations: audio equalizers, MRI and radio-telescope imaging, and the heat and wave equations of physics.

Example: Filtering hiss out of a recording

A microphone records f(t)=cos⁡(2π⋅440 t)+cos⁡(2π⋅3000 t)f(t) = \cos(2\pi \cdot 440\,t) + \cos(2\pi \cdot 3000\,t): a 440 Hz musical note (A4) mixed with a 3000 Hz electronic hiss. An engineer applies an ideal low-pass filter that sets f^(ξ)=0\hat f(\xi) = 0 whenever ∣ξ∣>1000|\xi| > 1000. What signal comes out?

Solution

Each cosine is a sum of two pure spinning exponentials at frequencies ±ν\pm\nu: in the frequency domain, f(t)=cos⁡(2π⋅440 t)+cos⁡(2π⋅3000 t)f(t) = \cos(2\pi \cdot 440\,t) + \cos(2\pi \cdot 3000\,t) has energy concentrated only at ξ=±440\xi = \pm 440 and ξ=±3000\xi = \pm 3000 (as spikes in f^\hat f).

The filter keeps everything with ∣ξ∣≤1000|\xi| \le 1000 and kills the rest. Since 440≤1000440 \le 1000 but 3000>10003000 > 1000, the spikes at ±3000\pm 3000 are removed while the spikes at ±440\pm 440 survive untouched.

Applying the inversion formula to what remains reconstructs exactly the surviving frequencies: the output is cos⁡(2π⋅440 t)\cos(2\pi \cdot 440\,t), the clean musical note with the hiss gone. This is literally how a graphic equalizer's low-pass knob works.

Example: Why heat spreads out: solving the heat equation

A rod has initial temperature profile u(x,0)=f(x)u(x,0) = f(x) and obeys the heat equation ut=uxxu_t = u_{xx}. Use the Fourier transform (in xx) to find u(x,t)u(x,t) for t>0t>0.

Solution

Taking the Fourier transform in xx turns each spatial derivative ∂x\partial_x into multiplication by 2πiξ2\pi i \xi, so uxxu_{xx} becomes −4π2ξ2u^-4\pi^2\xi^2 \hat u. The PDE ut=uxxu_t = u_{xx} becomes the ordinary differential equation ∂tu^(ξ,t)=−4π2ξ2 u^(ξ,t)\partial_t \hat u(\xi,t) = -4\pi^2\xi^2\, \hat u(\xi,t) in tt for each fixed ξ\xi — a huge simplification, since a hard PDE became an easy family of decoupled ODEs.

This ODE has solution u^(ξ,t)=f^(ξ) e−4π2ξ2t\hat u(\xi,t) = \hat f(\xi)\, e^{-4\pi^2 \xi^2 t}, an exponentially decaying factor that kills high frequencies ξ\xi fast: fine spatial detail smooths out quickly, which matches the everyday fact that sharp temperature spikes flatten out first.

Inverting the transform (a product in frequency is a convolution in space, by the convolution theorem run in reverse) gives u(x,t)=(f∗Kt)(x),  Kt(x)=14πt e−x2/(4t)u(x,t) = (f * K_t)(x),\ \ K_t(x) = \dfrac{1}{\sqrt{4\pi t}}\, e^{-x^2/(4t)}: the temperature at time tt is the initial profile smeared out (convolved) against a spreading Gaussian bump KtK_t — literally the mathematical picture of heat diffusing.

Definition: Uncertainty principle

A function cannot be sharply localized in both time and frequency at once. If Δx\Delta x and Δξ\Delta \xi measure the spread of ∣f∣2|f|^2 and ∣f^∣2|\hat f|^2 respectively, then Δx⋅Δξ≥14π\Delta x \cdot \Delta \xi \ge \dfrac{1}{4\pi}. Squeezing a pulse in time (small Δx\Delta x) forces its spectrum to spread out (large Δξ\Delta \xi), and vice versa — the same trade-off used to derive Heisenberg's uncertainty principle in quantum mechanics.

Δx⋅Δξ≥14π\Delta x \cdot \Delta \xi \ge \dfrac{1}{4\pi}

Under the convention f^(ξ)=∫f(x)e−2πixξ dx\hat f(\xi) = \int f(x) e^{-2\pi i x\xi}\,dx, the Fourier transform of the Gaussian f(x)=e−πx2f(x) = e^{-\pi x^2} is:

Plancherel's theorem, ∫−∞∞∣f(x)∣2 dx=∫−∞∞∣f^(ξ)∣2 dξ\int_{-\infty}^{\infty} |f(x)|^2\,dx = \int_{-\infty}^{\infty} |\hat f(\xi)|^2\,d\xi, is best described as a statement of:

An MRI machine measures samples of a spatial signal's Fourier transform ('k-space'). If only the low-frequency samples ( ∣ξ∣|\xi| small) are collected, the reconstructed image will be:

The uncertainty principle Δx⋅Δξ≥14π\Delta x \cdot \Delta \xi \ge \dfrac{1}{4\pi} implies that:

Fejér's theorem says the Cesàro means σNf\sigma_N f converge uniformly to any continuous ff. Carleson's theorem is a separate, much deeper statement about:

By M. Riesz's theorem, the Hilbert transform HH is bounded on Lp(R)L^p(\mathbb{R}) exactly for:

References

  1. Terence Tao (2003). Recent progress on the restriction conjecture · arXiv:math/0311181
  2. Elias M. Stein, Guido Weiss (1971). Introduction to Fourier Analysis on Euclidean Spaces
  3. Loukas Grafakos (2014). Classical Fourier Analysis · DOI:10.1007/978-1-4939-1194-3