MathLabs

Mathematical physics

Mathematical foundations of quantum mechanics

The functional analysis and operator theory that make quantum mechanics mathematically rigorous.

IntuitionQuantum states as vectors, observables as instruments

Imagine every possible state of a quantum system — an electron's spin, the position of a particle, the polarization of a photon — as an arrow (a vector) in some possibly infinite-dimensional space, and every physical quantity you could measure — position, momentum, energy — not as a plain number attached to the arrow, but as an instrument: a linear transformation that acts on the arrow and, when you measure, forces it to jump to one of the instrument's own special directions and reads out the corresponding number. This single picture — vectors, and instruments acting on them — is the entire mathematical skeleton of quantum mechanics; uncertainty, quantization and interference are all just what happens when you take that picture seriously and follow the mathematics wherever it leads.

A superposition of sinusoidal plane waves of increasing frequency combining into a localized wave packet, with a slider controlling how many harmonics are summed.
A wave packet built by superposing many momentum eigenstates eikxe^{ikx}; raising the harmonic count nn narrows the packet in position but spreads it out in momentum — the Fourier-analytic face of the uncertainty principle.

SchoolFrom wavefunctions to Hilbert space

Definition: Hilbert space and quantum states

A Hilbert space H\mathcal H is a complex vector space equipped with an inner product ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle that is complete (every Cauchy sequence converges) with respect to the norm ∥ψ∥=⟨ψ,ψ⟩\|\psi\|=\sqrt{\langle\psi,\psi\rangle}. For a single particle moving in dd dimensions, the relevant space is H=L2(Rd)\mathcal{H} = L^2(\mathbb{R}^d), the square-integrable complex functions ψ(x)\psi(x) with ∫Rd∣ψ(x)∣2 dx<∞\int_{\mathbb R^d}|\psi(x)|^2\,dx<\infty. A pure quantum state is a unit vector ψ∈H\psi\in\mathcal H (strictly, a ray {eiθψ}\{e^{i\theta}\psi\}, since an overall phase is unobservable), and ∣ψ(x)∣2|\psi(x)|^2 is interpreted as a probability density for finding the particle at xx.

⟨ψ,φ⟩=∫Rdψ(x)‾ φ(x) dx,∥ψ∥2=⟨ψ,ψ⟩=∫Rd∣ψ(x)∣2 dx=1\langle \psi,\varphi\rangle = \int_{\mathbb{R}^d} \overline{\psi(x)}\,\varphi(x)\,dx, \qquad \|\psi\|^2=\langle\psi,\psi\rangle=\int_{\mathbb{R}^d}|\psi(x)|^2\,dx = 1

A physical observable — position, momentum, energy, spin — is represented not by a number but by a self-adjoint (Hermitian) operator A^\hat A on H\mathcal H, meaning A^=A^∗\hat{A} = \hat{A}^*: for all ψ,φ\psi,\varphi in its domain, ⟨ψ,A^φ⟩=⟨A^ψ,φ⟩\langle\psi,\hat A\varphi\rangle=\langle\hat A\psi,\varphi\rangle. Self-adjointness is exactly the algebraic condition that forces the possible measurement outcomes (the spectrum of A^\hat A) to be real numbers, and that guarantees eigenvectors for different eigenvalues are orthogonal — the mathematical reason a measurement of a real physical quantity always returns a real number.

iℏ ∂tψ(x,t)=H^ψ(x,t),H^=−ℏ22m∇2+V(x),U^(t)=e−itH^/ℏi\hbar\,\partial_t \psi(x,t) = \hat H\psi(x,t), \qquad \hat H = -\frac{\hbar^2}{2m}\nabla^2 + V(x), \qquad \hat U(t)=e^{-i t\hat H/\hbar}

Stone's theorem says these two pictures are the same fact seen two ways: a strongly continuous one-parameter group of unitary operators U^(t)\hat U(t) (time evolution that preserves total probability, U^(t)∗U^(t)=I\hat U(t)^*\hat U(t)=I) is always of the form U^(t)=e−itH^/ℏ\hat{U}(t) = e^{-i t \hat{H}/\hbar} for a unique self-adjoint operator H^\hat H, the Hamiltonian — and conversely every self-adjoint H^\hat H generates such a group. Differentiating U^(t)ψ\hat U(t)\psi at t=0t=0 recovers the Schrödinger equation iℏ∂tψ=H^ψi\hbar \partial_t \psi = \hat{H}\psi: the two are logically equivalent formulations of the same dynamics.

Classical vs. quantum description of a physical system
ConceptClassical mechanicsQuantum mechanics
StateA point (x,p)(x,p) in phase spaceA unit vector ψ∈\psi\in H=L2(Rd)\mathcal{H} = L^2(\mathbb{R}^d)
ObservableA function f(x,p)f(x,p) on phase spaceA self-adjoint operator A^=A^∗\hat{A} = \hat{A}^*
Measurement outcomeThe value f(x,p)f(x,p), exactlyAn eigenvalue of A^\hat A, with probability ∣⟨e∣ψ⟩∣2|\langle e|\psi\rangle|^2
Time evolutionHamilton's equations q˙=∂H/∂p\dot q=\partial H/\partial pUnitary flow U^(t)=e−itH^/ℏ\hat{U}(t) = e^{-i t \hat{H}/\hbar}

UndergraduateTheorems that make the picture rigorous

For any two self-adjoint operators A^=A^∗\hat A=\hat A^*, B^=B^∗\hat B=\hat B^* on H\mathcal H and any unit vector ψ\psi in the domain of A^B^\hat A\hat B and B^A^\hat B\hat A, the standard deviations σA=⟨ψ∣(A^−⟨A^⟩)2∣ψ⟩\sigma_A=\sqrt{\langle\psi|(\hat A-\langle\hat A\rangle)^2|\psi\rangle} and σB\sigma_B (expectations taken in state ψ\psi) satisfy σAσB≥12∣⟨[A^,B^]⟩∣\sigma_A \sigma_B \ge \tfrac{1}{2}\left|\langle [\hat{A},\hat{B}]\rangle\right|, where [A^,B^]=A^B^−B^A^[\hat A,\hat B]=\hat A\hat B-\hat B\hat A.

Why is it true?

Two observables can only be measured with simultaneous perfect precision if their operators commute; the size of the commutator is a direct measure of how incompatible they are. Cauchy–Schwarz turns "these two vectors cannot both be short" into a hard numerical bound, and separating the inner product into its real and imaginary parts is exactly what isolates the commutator (the genuinely quantum part) from the anticommutator (a classical-looking correlation term).

Proof

Step 1 (centered operators). Let ΔA^=A^−⟨A^⟩I\Delta\hat A=\hat A-\langle\hat A\rangle I and ΔB^=B^−⟨B^⟩I\Delta\hat B=\hat B-\langle\hat B\rangle I, both still self-adjoint since ⟨A^⟩,⟨B^⟩\langle\hat A\rangle,\langle\hat B\rangle are real scalars. By definition σA2=⟨ψ∣ΔA^2∣ψ⟩=∥ΔA^ ψ∥2\sigma_A^2=\langle\psi|\Delta\hat A^2|\psi\rangle=\|\Delta\hat A\,\psi\|^2 and likewise σB2=∥ΔB^ ψ∥2\sigma_B^2=\|\Delta\hat B\,\psi\|^2.

Step 2 (Cauchy–Schwarz). Apply the Cauchy–Schwarz inequality ∣⟨f∣g⟩∣2≤⟨f∣f⟩⟨g∣g⟩|\langle f|g\rangle|^2\le\langle f|f\rangle\langle g|g\rangle to f=ΔA^ ψf=\Delta\hat A\,\psi and g=ΔB^ ψg=\Delta\hat B\,\psi: ∣⟨ΔA^ ψ∣ΔB^ ψ⟩∣2≤σA2 σB2.|\langle\Delta\hat A\,\psi|\Delta\hat B\,\psi\rangle|^2 \le \sigma_A^2\,\sigma_B^2.

Step 3 (split into real and imaginary parts). Write ⟨ΔA^ ψ∣ΔB^ ψ⟩=⟨ψ∣ΔA^ΔB^∣ψ⟩\langle\Delta\hat A\,\psi|\Delta\hat B\,\psi\rangle=\langle\psi|\Delta\hat A\Delta\hat B|\psi\rangle. Since ΔA^,ΔB^\Delta\hat A,\Delta\hat B are self-adjoint, complex-conjugating flips the operator order: ⟨ΔA^ΔB^⟩∗=⟨ΔB^ΔA^⟩\langle\Delta\hat A\Delta\hat B\rangle^*=\langle\Delta\hat B\Delta\hat A\rangle. Hence the anticommutator expectation ⟨{ΔA^,ΔB^}⟩=⟨ΔA^ΔB^⟩+⟨ΔB^ΔA^⟩=2 Re⟨ΔA^ΔB^⟩\langle\{\Delta\hat A,\Delta\hat B\}\rangle=\langle\Delta\hat A\Delta\hat B\rangle+\langle\Delta\hat B\Delta\hat A\rangle=2\,\mathrm{Re}\langle\Delta\hat A\Delta\hat B\rangle is real, while the commutator expectation ⟨[ΔA^,ΔB^]⟩=⟨ΔA^ΔB^⟩−⟨ΔB^ΔA^⟩=2i Im⟨ΔA^ΔB^⟩\langle[\Delta\hat A,\Delta\hat B]\rangle=\langle\Delta\hat A\Delta\hat B\rangle-\langle\Delta\hat B\Delta\hat A\rangle=2i\,\mathrm{Im}\langle\Delta\hat A\Delta\hat B\rangle is purely imaginary. Also [ΔA^,ΔB^]=[A^,B^][\Delta\hat A,\Delta\hat B]=[\hat A,\hat B], since the constant shifts cancel in the commutator.

Step 4 (recombine). By Pythagoras applied to real and imaginary parts, ∣⟨ΔA^ΔB^⟩∣2=(Re⟨ΔA^ΔB^⟩)2+(Im⟨ΔA^ΔB^⟩)2=14∣⟨{ΔA^,ΔB^}⟩∣2+14∣⟨[A^,B^]⟩∣2≥14∣⟨[A^,B^]⟩∣2,|\langle\Delta\hat A\Delta\hat B\rangle|^2=\big(\mathrm{Re}\langle\Delta\hat A\Delta\hat B\rangle\big)^2+\big(\mathrm{Im}\langle\Delta\hat A\Delta\hat B\rangle\big)^2=\tfrac14|\langle\{\Delta\hat A,\Delta\hat B\}\rangle|^2+\tfrac14|\langle[\hat A,\hat B]\rangle|^2 \ge \tfrac14|\langle[\hat A,\hat B]\rangle|^2, dropping the manifestly nonnegative anticommutator term. Combining with Step 2, σA2σB2≥14∣⟨[A^,B^]⟩∣2\sigma_A^2\sigma_B^2\ge\tfrac14|\langle[\hat A,\hat B]\rangle|^2, and taking square roots (both sides nonnegative) gives σAσB≥12∣⟨[A^,B^]⟩∣\sigma_A \sigma_B \ge \tfrac{1}{2}\left|\langle [\hat{A},\hat{B}]\rangle\right|.

The Hamiltonian H^=P^22m+12mω2X^2\hat H=\dfrac{\hat P^2}{2m}+\dfrac12 m\omega^2\hat X^2, with [X^,P^]=iℏ[\hat X,\hat P]=i\hbar, has discrete, nondegenerate spectrum En=ℏω(n+12)E_n = \hbar\omega\left(n+\tfrac12\right), n=0,1,2,…n=0,1,2,\dots.

Why is it true?

Rather than solving a differential equation directly, factor the Hamiltonian algebraically into 'raising' and 'lowering' operators, exactly as one factors a quadratic form. Positivity of the resulting number operator, plus the algebra of how the ladder operators shift its eigenvalues by one unit, forces the eigenvalues to form an evenly spaced ladder starting at a nonnegative bottom rung — energy quantization falls out of algebra alone.

Proof

Step 1 (ladder operators). Define a^=mω2ℏ(X^+imωP^)\hat a=\sqrt{\tfrac{m\omega}{2\hbar}}\left(\hat X+\tfrac{i}{m\omega}\hat P\right) and a^†=mω2ℏ(X^−imωP^)\hat a^\dagger=\sqrt{\tfrac{m\omega}{2\hbar}}\left(\hat X-\tfrac{i}{m\omega}\hat P\right). Using [X^,P^]=iℏ[\hat X,\hat P]=i\hbar, a direct computation gives [a^,a^†]=mω2ℏ(−imω[X^,P^]+imω[P^,X^])=12ℏ(ℏ+ℏ)=1[\hat a,\hat a^\dagger]=\tfrac{m\omega}{2\hbar}\left(\tfrac{-i}{m\omega}[\hat X,\hat P]+\tfrac{i}{m\omega}[\hat P,\hat X]\right)=\tfrac{1}{2\hbar}\left(\hbar+\hbar\right)=1.

Step 2 (rewrite the Hamiltonian). Inverting, X^=ℏ2mω(a^+a^†)\hat X=\sqrt{\tfrac{\hbar}{2m\omega}}(\hat a+\hat a^\dagger) and P^=imωℏ2(a^†−a^)\hat P=i\sqrt{\tfrac{m\omega\hbar}{2}}(\hat a^\dagger-\hat a); substituting into H^=P^22m+12mω2X^2\hat H=\tfrac{\hat P^2}{2m}+\tfrac12 m\omega^2\hat X^2 and using [a^,a^†]=1[\hat a,\hat a^\dagger]=1 to reorder terms gives H^=ℏω(a^†a^+12)\hat H=\hbar\omega\left(\hat a^\dagger\hat a+\tfrac12\right). Define the number operator N^=a^†a^\hat N=\hat a^\dagger\hat a; it is self-adjoint, and for any state ϕ\phi, ⟨ϕ∣N^∣ϕ⟩=∥a^ϕ∥2≥0\langle\phi|\hat N|\phi\rangle=\|\hat a\phi\|^2\ge0, so every eigenvalue nn of N^\hat N satisfies n≥0n\ge0.

Step 3 (ladder relations). From [a^,a^†]=1[\hat a,\hat a^\dagger]=1 one gets [N^,a^]=−a^[\hat N,\hat a]=-\hat a and [N^,a^†]=a^†[\hat N,\hat a^\dagger]=\hat a^\dagger. So if N^∣n⟩=n∣n⟩\hat N|n\rangle=n|n\rangle, then N^a^=a^N^−a^=a^(N^−1)\hat N\hat a=\hat a\hat N-\hat a=\hat a(\hat N-1), giving N^(a^∣n⟩)=(n−1)(a^∣n⟩)\hat N(\hat a|n\rangle)=(n-1)(\hat a|n\rangle): a^\hat a lowers the eigenvalue by 11 (or annihilates the state), and a^†\hat a^\dagger raises it by 11.

Step 4 (termination and the spectrum). Because N^≥0\hat N\ge0, repeatedly applying a^\hat a to any eigenstate cannot produce eigenvalues below 00; the descending chain n,n−1,n−2,…n,n-1,n-2,\dots must terminate at n=0n=0 (a non-integer starting value would force a negative eigenvalue after finitely many steps, a contradiction), so every eigenvalue of N^\hat N is a nonnegative integer, and the ground state ∣0⟩|0\rangle satisfies a^∣0⟩=0\hat a|0\rangle=0. Applying a^†\hat a^\dagger repeatedly to ∣0⟩|0\rangle produces normalizable eigenstates ∣n⟩∝(a^†)n∣0⟩|n\rangle\propto(\hat a^\dagger)^n|0\rangle for every n=0,1,2,…n=0,1,2,\dots, and no others exist. Substituting N^∣n⟩=n∣n⟩\hat N|n\rangle=n|n\rangle into H^=ℏω(N^+12)\hat H=\hbar\omega(\hat N+\tfrac12) gives En=ℏω(n+12)E_n = \hbar\omega\left(n+\tfrac12\right).

UndergraduateReal-World Applications and Worked Examples

These are not textbook abstractions. Quantum tunneling — a direct consequence of the Schrödinger equation being a wave equation, not a particle-trajectory equation — is what lets flash-memory cells store a bit by pushing electrons through an insulating barrier, and what lets a scanning tunneling microscope (STM) image individual atoms by measuring a tunneling current. Discrete energy levels of the kind guaranteed by the spectral theorem are what an MRI machine reads out as a resonance frequency, and what a superconducting qubit uses as its two lowest energy levels ∣0⟩,∣1⟩|0\rangle,|1\rangle to store quantum information; the uncertainty principle sets the ultimate noise floor for how precisely such a qubit's state can be measured and controlled.

Example: Tunneling current in a scanning tunneling microscope

An electron approaches a rectangular potential barrier of height V0=4.0 eVV_0=4.0\,\text{eV} (the vacuum gap between an STM tip and a metal surface) and width L=0.50 nmL=0.50\,\text{nm}, with kinetic energy E=3.0 eV<V0E=3.0\,\text{eV}<V_0. Using the transmission coefficient T≈16EV0(1−EV0)e−2κLT\approx 16\dfrac{E}{V_0}\left(1-\dfrac{E}{V_0}\right)e^{-2\kappa L} with κ=2m(V0−E)ℏ\kappa=\dfrac{\sqrt{2m(V_0-E)}}{\hbar}, estimate TT, and explain why moving the tip 0.1 nm0.1\,\text{nm} closer changes the tunneling current so dramatically.

Solution

Step 1: compute κ\kappa. With V0−E=1.0 eV=1.6×10−19 JV_0-E=1.0\,\text{eV}=1.6\times10^{-19}\,\text{J}, m=9.11×10−31 kgm=9.11\times10^{-31}\,\text{kg}, ℏ=1.055×10−34 J⋅s\hbar=1.055\times10^{-34}\,\text{J·s}: κ=2m(V0−E)/ℏ≈5.1×109 m−1\kappa=\sqrt{2m(V_0-E)}/\hbar\approx 5.1\times10^{9}\,\text{m}^{-1}, i.e. about 5.1 nm−15.1\,\text{nm}^{-1}.

Step 2: compute the exponent. 2κL≈2×5.1 nm−1×0.50 nm≈5.12\kappa L\approx 2\times5.1\,\text{nm}^{-1}\times0.50\,\text{nm}\approx5.1, so e−2κL≈e−5.1≈6.1×10−3e^{-2\kappa L}\approx e^{-5.1}\approx6.1\times10^{-3}.

Step 3: assemble TT. The prefactor 16EV0(1−EV0)=16×0.75×0.25=3.016\frac{E}{V_0}(1-\frac{E}{V_0})=16\times0.75\times0.25=3.0, so T≈3.0×6.1×10−3≈1.8×10−2T\approx3.0\times6.1\times10^{-3}\approx1.8\times10^{-2}.

Step 4: sensitivity to distance. Because T∝e−2κLT\propto e^{-2\kappa L}, decreasing LL by 0.1 nm0.1\,\text{nm} multiplies TT by e2κ×0.1 nm=e1.02≈2.8e^{2\kappa\times0.1\,\text{nm}}=e^{1.02}\approx2.8 — a roughly threefold jump in tunneling current from a sub-atomic change in distance. This exponential sensitivity is what lets an STM resolve individual atoms.

Example: Minimum uncertainty of a spin-1/2 qubit

A superconducting qubit is prepared in the state ∣0⟩|0\rangle (spin up along zz), an eigenstate of σ^z\hat\sigma_z with eigenvalue +1+1. Using the Pauli operators σ^x,σ^y,σ^z\hat\sigma_x,\hat\sigma_y,\hat\sigma_z (each self-adjoint, each with eigenvalues ±1\pm1, and [σ^x,σ^y]=2iσ^z[\hat\sigma_x,\hat\sigma_y]=2i\hat\sigma_z), compute σσx\sigma_{\sigma_x} and σσy\sigma_{\sigma_y} in this state and check the Robertson–Heisenberg bound for the pair (σ^x,σ^y)(\hat\sigma_x,\hat\sigma_y).

Solution

Step 1: expectation values in ∣0⟩|0\rangle. Since ∣0⟩|0\rangle is an eigenstate of σ^z\hat\sigma_z only, and σ^x,σ^y\hat\sigma_x,\hat\sigma_y flip it to the orthogonal state ∣1⟩|1\rangle, ⟨0∣σ^x∣0⟩=⟨0∣σ^y∣0⟩=0\langle0|\hat\sigma_x|0\rangle=\langle0|\hat\sigma_y|0\rangle=0 while ⟨0∣σ^z∣0⟩=1\langle0|\hat\sigma_z|0\rangle=1.

Step 2: variances. Because σ^x2=σ^y2=I\hat\sigma_x^2=\hat\sigma_y^2=I, σσx2=⟨0∣σ^x2∣0⟩−⟨0∣σ^x∣0⟩2=1−0=1\sigma_{\sigma_x}^2=\langle0|\hat\sigma_x^2|0\rangle-\langle0|\hat\sigma_x|0\rangle^2=1-0=1, and identically σσy2=1\sigma_{\sigma_y}^2=1; so σσx=σσy=1\sigma_{\sigma_x}=\sigma_{\sigma_y}=1 and σσxσσy=1\sigma_{\sigma_x}\sigma_{\sigma_y}=1.

Step 3: the Robertson–Heisenberg bound. 12∣⟨0∣[σ^x,σ^y]∣0⟩∣=12∣⟨0∣2iσ^z∣0⟩∣=∣⟨0∣σ^z∣0⟩∣=1\frac12|\langle0|[\hat\sigma_x,\hat\sigma_y]|0\rangle|=\frac12|\langle0|2i\hat\sigma_z|0\rangle|=|\langle0|\hat\sigma_z|0\rangle|=1.

Step 4: compare. The product 11 equals the bound 11: ∣0⟩|0\rangle is a minimum-uncertainty state for (σ^x,σ^y)(\hat\sigma_x,\hat\sigma_y), meaning a measurement of σ^x\hat\sigma_x or σ^y\hat\sigma_y on a qubit prepared along zz is maximally random (50/50).

What is the commutator [X^,P^]=iℏ[\hat X,\hat P] = i\hbar equal to?

By Stone's theorem, for U^(t)=e−itH^/ℏ\hat{U}(t) = e^{-i t \hat{H}/\hbar} to be unitary for every tt, what must H^\hat H be?

A quantum harmonic oscillator has ℏω=2 eV\hbar\omega=2\,\text{eV}. What is the energy of its n=1n=1 level?

In a scanning tunneling microscope, the extreme sensitivity of the tunneling current to tip-sample distance comes from which feature of T≈16EV0(1−EV0)e−2κLT\approx 16\frac{E}{V_0}(1-\frac{E}{V_0})e^{-2\kappa L}?

References

  1. John von Neumann (1955). Mathematical Foundations of Quantum Mechanics
  2. Howard P. Robertson (1929). The Uncertainty Principle · DOI:10.1103/PhysRev.34.163
  3. Marshall H. Stone (1932). On One-Parameter Unitary Groups in Hilbert Space · DOI:10.2307/1968538