Studies infinite-dimensional vector spaces of functions and the linear operators acting on them.
IntuitionWhen can we do calculus with infinitely many coordinates?
A vector in Rn is just a list of n numbers, and calculus on Rn works because we can measure length (∥x∥) and angle (inner product) between such lists. A function like a sound wave, an image, or a quantum wavefunction is really an infinite list of numbers — its value at every point, or equivalently its infinitely many Fourier coefficients. Functional analysis asks: can we still measure "length" and "angle" between functions, and still do calculus (limits, derivatives, optimization) on the resulting infinite-dimensional space? The answer is yes, provided the space is complete — infinite sums and limits do not "leak out" of the space — which is exactly what turns a plain vector space of functions into a Banach space or, when an inner product is available, a Hilbert space.
Directed graph showing inclusion of Hilbert spaces into Banach spaces into normed vector spaces, with concrete examples labeled at each level.
The network of inclusions between spaces of functions and sequences: every Hilbert space (node ℓ2, node L2[0,1]) is a Banach space, and every Banach space (also C[0,1] and ℓp) is a normed vector space, but the arrows do not reverse — C[0,1] with the sup norm has no inner product that induces it. Highlighting a node traces which broader class it belongs to.
UndergraduateNorms, completeness, and inner products
Definition: Normed vector space and Banach space
A norm on a vector space X over R or C is a map ∥x∥ satisfying ∥x∥=0 iff x=0, ∥λx∥=∣λ∣∥x∥, and the triangle inequality ∥x+y∥≤∥x∥+∥y∥. A sequence (xn) is a Cauchy sequence if its terms eventually get arbitrarily close together; X is complete, and called a Banach space, if every Cauchy sequence in X converges to some point of X (not just of a larger space containing it). Completeness is what lets us build a limit function out of an infinite process — an infinite series of functions, an iterative approximation scheme — and be sure the limit is still a bona fide element of the space.
Written out, the Cauchy condition says the tail of the sequence gets tightly bunched together in norm:
∀ε>0∃N:n,m≥N⟹∥xn−xm∥<ε
Here ε is the tolerance, N is the point beyond which every pair of terms is within that tolerance, and ∥x∥ is the norm on X. When X additionally carries an inner product⟨x,x⟩ — a bilinear (or sesquilinear) pairing generalizing the dot product, from which the norm is recovered as ∥x∥=⟨x,x⟩ — and is complete in that induced norm, X is called a Hilbert space, usually written H. The inner product is what lets us speak of orthogonality and angle between functions, not just their length.
∣⟨x,y⟩∣≤∥x∥∥y∥
This is the Cauchy–Schwarz inequality: the inner product of two vectors can never exceed the product of their lengths, exactly like u⋅v=∥u∥∥v∥cosθ in ordinary 3D space, where ∣cosθ∣≤1. Hilbert spaces are further characterized among Banach spaces by the parallelogram law, ∥x+y∥2+∥x−y∥2=2∥x∥2+2∥y∥2 — a purely norm-based identity (no inner product needed to state it) that a norm must satisfy if and only if it actually comes from an inner product. It says the sum of the squared diagonals of a parallelogram equals the sum of the squared sides, generalizing the Pythagorean-flavored geometry of Euclidean space to infinite dimensions.
Common Banach and Hilbert spaces
Space
Norm
Complete?
Hilbert space?
ℓ2
(∑i∣xi∣2)1/2
Yes
Yes
ℓp, p=2
(∑i∣xi∣p)1/p
Yes
No
C[0,1]
supt∣f(t)∣
Yes
No
L2[0,1]
(∫01∣f(t)∣2dt)1/2
Yes
Yes
UndergraduateTwo pillars: the Riesz representation theorem and the Hahn–Banach theorem
Let H be a Hilbert space and φ a bounded (continuous) linear functional on H. Then there is a unique y∈H such that φ(x)=⟨x,y⟩ for every x∈H, and moreover ∥φ∥=∥y∥.
Why is it true?
It says every way of assigning a number to each vector linearly and continuously is secretly just "taking the inner product with some fixed vector" — abstract functionals are no more general than the most concrete ones you already know.
Proof
If φ=0, take y=0 and we are done. Otherwise let N=kerφ={x∈H:φ(x)=0}; since φ is continuous and linear, N is a closed proper subspace of H. Because H is a Hilbert space, the projection theorem gives an orthogonal decomposition H=N⊕N⊥, and since N is a proper subspace, N⊥ contains some z=0.
For any x∈H, consider the vector u=φ(x)z−φ(z)x. Applying φ gives φ(u)=φ(x)φ(z)−φ(z)φ(x)=0, so u∈N. Since z∈N⊥, we get ⟨u,z⟩=0, i.e. φ(x)⟨z,z⟩−φ(z)⟨x,z⟩=0. Solving for φ(x) gives φ(x)=∥z∥2φ(z)⟨x,z⟩=⟨x,∥z∥2φ(z)z⟩, so y=∥z∥2φ(z)z represents φ.
For uniqueness, if ⟨x,y1⟩=⟨x,y2⟩ for all x, take x=y1−y2 to get ∥y1−y2∥2=0, so y1=y2. For the norm identity, Cauchy–Schwarz gives ∣φ(x)∣=∣⟨x,y⟩∣≤∥y∥∥x∥, so ∥φ∥≤∥y∥; and plugging in x=y gives φ(y)=∥y∥2, so ∥φ∥≥∣φ(y)∣/∥y∥=∥y∥. Together, ∥φ∥=∥y∥.
Let X be a real normed vector space, Y a linear subspace of X, and φ a bounded linear functional on Y with ∣φ(y)∣≤M∥y∥ for all y∈Y. Then there exists a bounded linear functional Φ on all of X such that Φ(y)=φ(y) for every y∈Y, and ∣Φ(x)∣≤M∥x∥ for every x∈X.
Why is it true?
It guarantees that a linear functional defined only on a small subspace — for instance, only knowing "the value of a signal at a few sample points" — can always be extended to the whole space without growing its bound; nothing ever forces us to be stuck working on a subspace.
Proof
First extend by one dimension. Pick x0∈/Y and let Y1=Y⊕Rx0. We must choose the value c=Φ(x0) so that ∣φ(y)+tc∣≤M∥y+tx0∥ for all y∈Y,t∈R; dividing by t=0 and substituting y/t→y, this reduces to needing c with φ(y)−M∥y−x0∥≤c≤M∥y+x0∥−φ(y) for all y∈Y. Using φ(y1)−φ(y2)=φ(y1−y2)≤M∥y1−y2∥≤M∥y1+x0∥+M∥y2−x0∥, one checks the supremum of the left expressions over y1 never exceeds the infimum of the right expressions over y2, so a valid c in that interval exists; this defines Φ on Y1 with the same bound M.
Next, extend to all of X using Zorn's Lemma. Consider the set of all pairs (Z,Ψ) where Z is a subspace with Y⊆Z⊆X and Ψ extends φ to Z with bound M, partially ordered by (Z1,Ψ1)≤(Z2,Ψ2) iff Z1⊆Z2 and Ψ2∣Z1=Ψ1. Every chain has an upper bound (take the union of the subspaces and the functional agreeing with each on its domain), so Zorn's Lemma gives a maximal element (Z∗,Φ).
Finally, if Z∗=X, the one-dimension extension step above applied to Z∗ and any x0∈X∖Z∗ would produce a strictly larger admissible pair, contradicting maximality of (Z∗,Φ). Hence Z∗=X, and Φ is the desired bound-preserving extension to all of X.
UndergraduateReal-World Applications and Worked Examples
Functional analysis is the mathematical backbone of Fourier analysis and signal processing (a signal lives in L2[0,2π] or L2[0,1]), of quantum mechanics (states live in a Hilbert space H), and of statistics, engineering, and machine learning problems that boil down to finding the closest point in a subspace — from least-squares regression to Tikhonov-regularized inverse problems.
Example: Signal processing: does sin(nt) settle down as n→∞?
Fix f∈L2[0,2π]. Show that ∫02πf(t)sin(nt)dt→0 as n→∞ (i.e. sin(nt)⇀0weakly), while the norm ∥sin(nt)∥2=π stays constant — so sin(nt) itself never settles down in the strong (norm) sense.
Solution
Consider the orthonormal system en(t)=sin(nt)/π in L2[0,2π] (orthogonality follows from ∫02πsin(nt)sin(mt)dt=0 for n=m, and ∫02πsin2(nt)dt=π).
By Bessel's inequality, ∑n=1∞∣⟨f,en⟩∣2≤∥f∥2, the series of squared Fourier coefficients of f against this orthonormal system converges, since ∥f∥2<∞. A convergent series has terms tending to zero, so ∣⟨f,en⟩∣2→0, i.e. ⟨f,en⟩→0.
Since ⟨f,en⟩=π1∫02πf(t)sin(nt)dt, this is exactly ∫02πf(t)sin(nt)dt→0 — this is the Riemann–Lebesgue lemma in disguise, and precisely the statement sin(nt)⇀0.
On the other hand, direct computation gives ∥sin(nt)∥22=∫02πsin2(nt)dt=π for every n, so ∥sin(nt)∥2=π for all n — the norm never shrinks. So sin(nt) converges to 0 when tested against any fixed f (weak convergence), but never converges to 0 in norm (no strong convergence): the oscillations get faster and "average out" against every fixed probe, without the signal itself losing energy.
Example: Statistics and engineering: least-squares regression as an orthogonal projection
Given data matrix A (columns = predictors) and observations b, the least-squares fit minimizes ∥Ax−b∥2 over x. Use the Hilbert-space projection idea (the finite-dimensional case of the Riesz/orthogonality argument above, with H=Rn and the dot product) to derive the normal equationsATAx^=ATb that the minimizer x^ must satisfy.
Solution
The set of achievable outputs {Ax:x∈Rn} is a subspace ran(A)⊆Rm (a finite-dimensional, automatically complete, hence closed subspace of the Hilbert space Rm). Minimizing ∥Ax−b∥2 is exactly the problem of finding the point of ran(A) closest to b.
By the Hilbert-space projection theorem (the same orthogonal-decomposition idea used to prove the Riesz representation theorem above), the closest point Ax^ in a closed subspace to b is characterized by the residual b−Ax^ being orthogonal to the entire subspace, i.e. ⟨b−Ax^,Av⟩=0∀v.
Writing Av=A(v1,…,vn) ranges over all columns of A as v ranges over the standard basis vectors, the condition ⟨b−Ax^,Av⟩=0∀v applied to each column aj of A says ⟨b−Ax^,aj⟩=0 for every j, which packaged as a single matrix equation is exactly AT(b−Ax^)=0.
Expanding gives ATb−ATAx^=0, i.e. the normal equations ATAx^=ATb. So the abstract Hilbert-space fact "the residual of a best approximation is orthogonal to the approximating subspace" is, in disguise, exactly the linear-algebra formula used every day to fit a regression line or a financial factor model.
ResearchOpen directions: the geometry of infinite-dimensional spaces
Which of the following Banach spaces is not a Hilbert space (its norm does not come from an inner product)?
In a Hilbert space, if ∥x∥=3 and ∥y∥=4, what is ∥x+y∥2+∥x−y∥2?
What does the Riesz representation theorem guarantee about a bounded linear functional φ on a Hilbert space H?
In the sin(nt) example, why does the signal sin(nt) "average out to zero" against every fixed probe f as n→∞, even though it never loses energy?