MathLabs

Analysis

Functional analysis

Studies infinite-dimensional vector spaces of functions and the linear operators acting on them.

IntuitionWhen can we do calculus with infinitely many coordinates?

A vector in Rn\mathbb{R}^n is just a list of nn numbers, and calculus on Rn\mathbb{R}^n works because we can measure length (∥x∥\|x\|) and angle (inner product) between such lists. A function like a sound wave, an image, or a quantum wavefunction is really an infinite list of numbers — its value at every point, or equivalently its infinitely many Fourier coefficients. Functional analysis asks: can we still measure "length" and "angle" between functions, and still do calculus (limits, derivatives, optimization) on the resulting infinite-dimensional space? The answer is yes, provided the space is complete — infinite sums and limits do not "leak out" of the space — which is exactly what turns a plain vector space of functions into a Banach space or, when an inner product is available, a Hilbert space.

Directed graph showing inclusion of Hilbert spaces into Banach spaces into normed vector spaces, with concrete examples labeled at each level.
The network of inclusions between spaces of functions and sequences: every Hilbert space (node ℓ2\ell^2, node L2[0,1]L^2[0,1]) is a Banach space, and every Banach space (also C[0,1]C[0,1] and ℓp\ell^p) is a normed vector space, but the arrows do not reverse — C[0,1]C[0,1] with the sup norm has no inner product that induces it. Highlighting a node traces which broader class it belongs to.

UndergraduateNorms, completeness, and inner products

Definition: Normed vector space and Banach space

A norm on a vector space XX over R\mathbb{R} or C\mathbb{C} is a map ∥x∥\|x\| satisfying ∥x∥\|x\|=0=0 iff x=0x=0, ∥λx∥=∣λ∣∥x∥\|\lambda x\|=|\lambda|\|x\|, and the triangle inequality ∥x+y∥≤∥x∥+∥y∥\|x+y\|\le\|x\|+\|y\|. A sequence (xn)(x_n) is a Cauchy sequence if its terms eventually get arbitrarily close together; XX is complete, and called a Banach space, if every Cauchy sequence in XX converges to some point of XX (not just of a larger space containing it). Completeness is what lets us build a limit function out of an infinite process — an infinite series of functions, an iterative approximation scheme — and be sure the limit is still a bona fide element of the space.

Written out, the Cauchy condition says the tail of the sequence gets tightly bunched together in norm:

∀ε>0 ∃N: n,m≥N⟹∥xn−xm∥<ε\forall\varepsilon>0\ \exists N:\ n,m\ge N \Longrightarrow \|x_n-x_m\|<\varepsilon

Here ε\varepsilon is the tolerance, NN is the point beyond which every pair of terms is within that tolerance, and ∥x∥\|x\| is the norm on XX. When XX additionally carries an inner product ⟨x,x⟩\langle x,x\rangle — a bilinear (or sesquilinear) pairing generalizing the dot product, from which the norm is recovered as ∥x∥=⟨x,x⟩\|x\|=\sqrt{\langle x,x\rangle} — and is complete in that induced norm, XX is called a Hilbert space, usually written HH. The inner product is what lets us speak of orthogonality and angle between functions, not just their length.

∣⟨x,y⟩∣≤∥x∥ ∥y∥|\langle x,y\rangle|\le\|x\|\,\|y\|

This is the Cauchy–Schwarz inequality: the inner product of two vectors can never exceed the product of their lengths, exactly like u⃗⋅v⃗=∥u⃗∥∥v⃗∥cos⁡θ\vec u\cdot\vec v=\|\vec u\|\|\vec v\|\cos\theta in ordinary 3D space, where ∣cos⁡θ∣≤1|\cos\theta|\le1. Hilbert spaces are further characterized among Banach spaces by the parallelogram law, ∥x+y∥2+∥x−y∥2=2∥x∥2+2∥y∥2\|x+y\|^2+\|x-y\|^2=2\|x\|^2+2\|y\|^2 — a purely norm-based identity (no inner product needed to state it) that a norm must satisfy if and only if it actually comes from an inner product. It says the sum of the squared diagonals of a parallelogram equals the sum of the squared sides, generalizing the Pythagorean-flavored geometry of Euclidean space to infinite dimensions.

Common Banach and Hilbert spaces
SpaceNormComplete?Hilbert space?
ℓ2\ell^2(∑i∣xi∣2)1/2\left(\sum_i|x_i|^2\right)^{1/2}YesYes
ℓp\ell^p, p≠2p\ne2(∑i∣xi∣p)1/p\left(\sum_i|x_i|^p\right)^{1/p}YesNo
C[0,1]C[0,1]sup⁡t∣f(t)∣\sup_{t}|f(t)|YesNo
L2[0,1]L^2[0,1](∫01∣f(t)∣2dt)1/2\left(\int_0^1|f(t)|^2dt\right)^{1/2}YesYes

UndergraduateTwo pillars: the Riesz representation theorem and the Hahn–Banach theorem

Let HH be a Hilbert space and φ\varphi a bounded (continuous) linear functional on HH. Then there is a unique y∈Hy\in H such that φ(x)=⟨x,y⟩\varphi(x)=\langle x,y\rangle for every x∈Hx\in H, and moreover ∥φ∥=∥y∥\|\varphi\|=\|y\|.

Why is it true?

It says every way of assigning a number to each vector linearly and continuously is secretly just "taking the inner product with some fixed vector" — abstract functionals are no more general than the most concrete ones you already know.

Proof

If φ=0\varphi=0, take y=0y=0 and we are done. Otherwise let N=ker⁡φ={x∈H:φ(x)=0}N=\ker\varphi=\{x\in H:\varphi(x)=0\}; since φ\varphi is continuous and linear, NN is a closed proper subspace of HH. Because HH is a Hilbert space, the projection theorem gives an orthogonal decomposition H=N⊕N⊥H=N\oplus N^{\perp}, and since NN is a proper subspace, N⊥N^{\perp} contains some z≠0z\ne0.

For any x∈Hx\in H, consider the vector u=φ(x)z−φ(z)xu=\varphi(x)z-\varphi(z)x. Applying φ\varphi gives φ(u)=φ(x)φ(z)−φ(z)φ(x)=0\varphi(u)=\varphi(x)\varphi(z)-\varphi(z)\varphi(x)=0, so u∈Nu\in N. Since z∈N⊥z\in N^{\perp}, we get ⟨u,z⟩=0\langle u,z\rangle=0, i.e. φ(x)⟨z,z⟩−φ(z)⟨x,z⟩=0\varphi(x)\langle z,z\rangle-\varphi(z)\langle x,z\rangle=0. Solving for φ(x)\varphi(x) gives φ(x)=φ(z)∥z∥2⟨x,z⟩=⟨x,φ(z)‾∥z∥2z⟩\varphi(x)=\dfrac{\varphi(z)}{\|z\|^2}\langle x,z\rangle=\left\langle x,\dfrac{\overline{\varphi(z)}}{\|z\|^2}z\right\rangle, so y=φ(z)‾∥z∥2zy=\dfrac{\overline{\varphi(z)}}{\|z\|^2}z represents φ\varphi.

For uniqueness, if ⟨x,y1⟩=⟨x,y2⟩\langle x,y_1\rangle=\langle x,y_2\rangle for all xx, take x=y1−y2x=y_1-y_2 to get ∥y1−y2∥2=0\|y_1-y_2\|^2=0, so y1=y2y_1=y_2. For the norm identity, Cauchy–Schwarz gives ∣φ(x)∣=∣⟨x,y⟩∣≤∥y∥∥x∥|\varphi(x)|=|\langle x,y\rangle|\le\|y\|\|x\|, so ∥φ∥≤∥y∥\|\varphi\|\le\|y\|; and plugging in x=yx=y gives φ(y)=∥y∥2\varphi(y)=\|y\|^2, so ∥φ∥≥∣φ(y)∣/∥y∥=∥y∥\|\varphi\|\ge|\varphi(y)|/\|y\|=\|y\|. Together, ∥φ∥=∥y∥\|\varphi\|=\|y\|.

Let XX be a real normed vector space, YY a linear subspace of XX, and φ\varphi a bounded linear functional on YY with ∣φ(y)∣≤M∥y∥|\varphi(y)|\le M\|y\| for all y∈Yy\in Y. Then there exists a bounded linear functional Φ\Phi on all of XX such that Φ(y)=φ(y)\Phi(y)=\varphi(y) for every y∈Yy\in Y, and ∣Φ(x)∣≤M∥x∥|\Phi(x)|\le M\|x\| for every x∈Xx\in X.

Why is it true?

It guarantees that a linear functional defined only on a small subspace — for instance, only knowing "the value of a signal at a few sample points" — can always be extended to the whole space without growing its bound; nothing ever forces us to be stuck working on a subspace.

Proof

First extend by one dimension. Pick x0∉Yx_0\notin Y and let Y1=Y⊕Rx0Y_1=Y\oplus\mathbb{R}x_0. We must choose the value c=Φ(x0)c=\Phi(x_0) so that ∣φ(y)+tc∣≤M∥y+tx0∥|\varphi(y)+tc|\le M\|y+tx_0\| for all y∈Y,t∈Ry\in Y,t\in\mathbb{R}; dividing by t≠0t\ne0 and substituting y/t→yy/t\to y, this reduces to needing cc with φ(y)−M∥y−x0∥≤c≤M∥y+x0∥−φ(y)\varphi(y)-M\|y-x_0\|\le c\le M\|y+x_0\|-\varphi(y) for all y∈Yy\in Y. Using φ(y1)−φ(y2)=φ(y1−y2)≤M∥y1−y2∥≤M∥y1+x0∥+M∥y2−x0∥\varphi(y_1)-\varphi(y_2)=\varphi(y_1-y_2)\le M\|y_1-y_2\|\le M\|y_1+x_0\|+M\|y_2-x_0\|, one checks the supremum of the left expressions over y1y_1 never exceeds the infimum of the right expressions over y2y_2, so a valid cc in that interval exists; this defines Φ\Phi on Y1Y_1 with the same bound MM.

Next, extend to all of XX using Zorn's Lemma. Consider the set of all pairs (Z,Ψ)(Z,\Psi) where ZZ is a subspace with Y⊆Z⊆XY\subseteq Z\subseteq X and Ψ\Psi extends φ\varphi to ZZ with bound MM, partially ordered by (Z1,Ψ1)≤(Z2,Ψ2)(Z_1,\Psi_1)\le(Z_2,\Psi_2) iff Z1⊆Z2Z_1\subseteq Z_2 and Ψ2∣Z1=Ψ1\Psi_2|_{Z_1}=\Psi_1. Every chain has an upper bound (take the union of the subspaces and the functional agreeing with each on its domain), so Zorn's Lemma gives a maximal element (Z∗,Φ)(Z^*,\Phi).

Finally, if Z∗≠XZ^*\ne X, the one-dimension extension step above applied to Z∗Z^* and any x0∈X∖Z∗x_0\in X\setminus Z^* would produce a strictly larger admissible pair, contradicting maximality of (Z∗,Φ)(Z^*,\Phi). Hence Z∗=XZ^*=X, and Φ\Phi is the desired bound-preserving extension to all of XX.

UndergraduateReal-World Applications and Worked Examples

Functional analysis is the mathematical backbone of Fourier analysis and signal processing (a signal lives in L2[0,2π]L^2[0,2\pi] or L2[0,1]L^2[0,1]), of quantum mechanics (states live in a Hilbert space HH), and of statistics, engineering, and machine learning problems that boil down to finding the closest point in a subspace — from least-squares regression to Tikhonov-regularized inverse problems.

Example: Signal processing: does sin⁡(nt)\sin(nt) settle down as n→∞n\to\infty?

Fix f∈L2[0,2π]f\in L^2[0,2\pi]. Show that ∫02πf(t)sin⁡(nt) dt→0\int_0^{2\pi} f(t)\sin(nt)\,dt\to 0 as n→∞n\to\infty (i.e. sin⁡(nt)⇀0\sin(nt)\rightharpoonup 0 weakly), while the norm ∥sin⁡(nt)∥2=π\|\sin(nt)\|_2=\sqrt{\pi} stays constant — so sin⁡(nt)\sin(nt) itself never settles down in the strong (norm) sense.

Solution

Consider the orthonormal system en(t)=sin⁡(nt)/πe_n(t)=\sin(nt)/\sqrt\pi in L2[0,2π]L^2[0,2\pi] (orthogonality follows from ∫02πsin⁡(nt)sin⁡(mt) dt=0\int_0^{2\pi}\sin(nt)\sin(mt)\,dt=0 for n≠mn\ne m, and ∫02πsin⁡2(nt) dt=π\int_0^{2\pi}\sin^2(nt)\,dt=\pi).

By Bessel's inequality, ∑n=1∞∣⟨f,en⟩∣2≤∥f∥2\sum_{n=1}^{\infty}|\langle f,e_n\rangle|^2\le\|f\|^2, the series of squared Fourier coefficients of ff against this orthonormal system converges, since ∥f∥2<∞\|f\|^2<\infty. A convergent series has terms tending to zero, so ∣⟨f,en⟩∣2→0|\langle f,e_n\rangle|^2\to0, i.e. ⟨f,en⟩→0\langle f,e_n\rangle\to0.

Since ⟨f,en⟩=1π∫02πf(t)sin⁡(nt) dt\langle f,e_n\rangle=\frac{1}{\sqrt\pi}\int_0^{2\pi}f(t)\sin(nt)\,dt, this is exactly ∫02πf(t)sin⁡(nt) dt→0\int_0^{2\pi} f(t)\sin(nt)\,dt\to 0 — this is the Riemann–Lebesgue lemma in disguise, and precisely the statement sin⁡(nt)⇀0\sin(nt)\rightharpoonup 0.

On the other hand, direct computation gives ∥sin⁡(nt)∥22=∫02πsin⁡2(nt) dt=π\|\sin(nt)\|_2^2=\int_0^{2\pi}\sin^2(nt)\,dt=\pi for every nn, so ∥sin⁡(nt)∥2=π\|\sin(nt)\|_2=\sqrt{\pi} for all nn — the norm never shrinks. So sin⁡(nt)\sin(nt) converges to 00 when tested against any fixed ff (weak convergence), but never converges to 00 in norm (no strong convergence): the oscillations get faster and "average out" against every fixed probe, without the signal itself losing energy.

Example: Statistics and engineering: least-squares regression as an orthogonal projection

Given data matrix AA (columns = predictors) and observations bb, the least-squares fit minimizes ∥Ax−b∥2\|Ax-b\|^2 over xx. Use the Hilbert-space projection idea (the finite-dimensional case of the Riesz/orthogonality argument above, with H=RnH=\mathbb{R}^n and the dot product) to derive the normal equations ATAx^=ATbA^{\mathsf T}A\hat x=A^{\mathsf T}b that the minimizer x^\hat x must satisfy.

Solution

The set of achievable outputs {Ax:x∈Rn}\{Ax:x\in\mathbb{R}^n\} is a subspace ran(A)⊆Rm\mathrm{ran}(A)\subseteq\mathbb{R}^m (a finite-dimensional, automatically complete, hence closed subspace of the Hilbert space Rm\mathbb{R}^m). Minimizing ∥Ax−b∥2\|Ax-b\|^2 is exactly the problem of finding the point of ran(A)\mathrm{ran}(A) closest to bb.

By the Hilbert-space projection theorem (the same orthogonal-decomposition idea used to prove the Riesz representation theorem above), the closest point Ax^A\hat x in a closed subspace to bb is characterized by the residual b−Ax^b-A\hat x being orthogonal to the entire subspace, i.e. ⟨b−Ax^, Av⟩=0  ∀v\langle b-A\hat x,\,Av\rangle=0\ \ \forall v.

Writing Av=A(v1,…,vn)Av=A(v_1,\dots,v_n) ranges over all columns of AA as vv ranges over the standard basis vectors, the condition ⟨b−Ax^, Av⟩=0  ∀v\langle b-A\hat x,\,Av\rangle=0\ \ \forall v applied to each column aja_j of AA says ⟨b−Ax^,aj⟩=0\langle b-A\hat x,a_j\rangle=0 for every jj, which packaged as a single matrix equation is exactly AT(b−Ax^)=0A^{\mathsf T}(b-A\hat x)=0.

Expanding gives ATb−ATAx^=0A^{\mathsf T}b-A^{\mathsf T}A\hat x=0, i.e. the normal equations ATAx^=ATbA^{\mathsf T}A\hat x=A^{\mathsf T}b. So the abstract Hilbert-space fact "the residual of a best approximation is orthogonal to the approximating subspace" is, in disguise, exactly the linear-algebra formula used every day to fit a regression line or a financial factor model.

ResearchOpen directions: the geometry of infinite-dimensional spaces

Which of the following Banach spaces is not a Hilbert space (its norm does not come from an inner product)?

In a Hilbert space, if ∥x∥=3\|x\|=3 and ∥y∥=4\|y\|=4, what is ∥x+y∥2+∥x−y∥2\|x+y\|^2+\|x-y\|^2?

What does the Riesz representation theorem guarantee about a bounded linear functional φ\varphi on a Hilbert space HH?

In the sin⁡(nt)\sin(nt) example, why does the signal sin⁡(nt)\sin(nt) "average out to zero" against every fixed probe ff as n→∞n\to\infty, even though it never loses energy?

References

  1. Walter Rudin (1991). Functional Analysis
  2. John B. Conway (2007). A Course in Functional Analysis
  3. Assaf Naor (2012). An introduction to the Ribe program · arXiv:1205.5993