MathLabs

Probability and statistics

Brownian motion and stochastic calculus

A continuous random process modeling erratic motion, with a calculus of its own used in finance and physics.

IntuitionWhy does a speck of pollen dancing in water need its own kind of calculus?

In 1827 the botanist Robert Brown watched pollen grains suspended in water under a microscope and saw them jitter endlessly, changing direction at every instant with no pattern he could discern. The explanation, only understood decades later, is that each grain is being bombarded from all sides by billions of invisible water molecules per second, and the net random kick at each moment nudges the grain a little further in a completely unpredictable direction. Brownian motion is the mathematical idealization of this endless, directionless jitter: a random path that is continuous — it never jumps — yet so jagged that it has no well-defined velocity at any instant, ever. That single fact, that the path is continuous but nowhere smooth, is what forces mathematicians to build an entirely new calculus, since ordinary calculus assumes you can zoom in on a curve and see it look like a straight line.

A bell-shaped Gaussian probability-density curve for the Brownian position at a fixed time.
At any time t>0t > 0, the position WtW_t of a standard Brownian motion has a Gaussian bell-curve density N(0,t)\mathcal{N}(0, t), with ≈68%\approx 68\% probability inside [−t,+t][-\sqrt{t}, +\sqrt{t}].

The distribution curve above describes where a particle may be at one fixed time. A two-dimensional Brownian motion uses two independent copies of this process; with time as a third coordinate, one sample becomes a jagged space-time path.

An interactive 3D discrete sample path of two independent Brownian coordinates against time, with a control for the number of steps.
One deterministic sample of two independent Brownian coordinates against time. Coordinates x and y show position, and z is time; change the step count to compare reproducible discrete samples at different resolutions.

UndergraduateThe Wiener process: a rigorous definition

Definition: Standard Wiener process (Brownian motion)

A standard Wiener process (or Brownian motion) is a continuous-time stochastic process (Wt)t≥0(W_t)_{t\ge0} such that: (1) W0=0W_0=0 almost surely; (2) it has independent increments — for any 0≤t1<t2<t3<⋯0\le t_1<t_2<t_3<\cdots, the increments Wt2−Wt1, Wt3−Wt2,…W_{t_2}-W_{t_1},\ W_{t_3}-W_{t_2},\dots are mutually independent; (3) each increment is normally distributed, Wt−Ws∼N(0,t−s)W_t-W_s\sim N(0,t-s) for s<ts<t; and (4) almost every sample path t↦Wtt\mapsto W_t is continuous. Property (3) alone already forces the process to look identical in distribution at every time scale — this self-similarity is why Brownian motion appears as the universal limit of so many different discrete random walks.

Wt−Ws∼N(0,t−s)W_t-W_s\sim N(0,t-s)

A defining, almost paradoxical feature of Brownian paths is their quadratic variation: summing the squared increments of WW over a fine partition of [0,t][0,t] converges (in probability) to tt itself, ∑i(Wti+1−Wti)2→t\sum_{i}(W_{t_{i+1}}-W_{t_i})^2 \to t, no matter how fine the partition — the squared wiggles never vanish, they accumulate to exactly tt. Compare this to any ordinary differentiable function ff, whose quadratic variation over any interval is always 00, since its increments shrink like Δt\Delta t rather than Δt\sqrt{\Delta t}. This single difference in scaling — Brownian increments are of order Δt\sqrt{\Delta t}, not Δt\Delta t — is the root cause of everything unusual about stochastic calculus.

∑i(Wti+1−Wti)2→t\sum_{i}(W_{t_{i+1}}-W_{t_i})^2 \to t
Ordinary calculus vs. stochastic (Itô) calculus
AspectSmooth function f(t)f(t)Brownian path WtW_t
DifferentiabilityDifferentiable, tangent line existsNowhere differentiable, no tangent anywhere
Increment size over Δt\Delta tOrder Δt\Delta tOrder Δt\sqrt{\Delta t}
Quadratic variation on [0,t][0,t]00tt (never zero)
Chain rule for F(f(t))F(f(t))dF=F′(f)f′(t) dtdF=F'(f)f'(t)\,dtNeeds an extra second-derivative term (Itô's lemma)

For a standard Wiener process WW and any t>0t>0, ∑i(Wti+1−Wti)2→t\sum_{i}(W_{t_{i+1}}-W_{t_i})^2 \to t where the sum is over a partition of [0,t][0,t] whose mesh tends to 00, the limit holding in probability (in fact almost surely along dyadic partitions). Consequently, almost every sample path s↦Wss\mapsto W_s is differentiable at no point s∈[0,t]s\in[0,t].

Why is it true?

This is the theorem that makes stochastic calculus necessary in the first place: it rules out treating dWtdW_t as an ordinary infinitesimal the way dtdt is treated in classical calculus, and it justifies the heuristic rule (dWt)2=dt(dW_t)^2=dt that drives Itô's lemma below.

Proof

Quadratic variation. Partition [0,t][0,t] into nn equal pieces of length t/nt/n and let Qn=∑i=1n(Wit/n−W(i−1)t/n)2Q_n=\sum_{i=1}^n(W_{it/n}-W_{(i-1)t/n})^2. Each term (Wit/n−W(i−1)t/n)2(W_{it/n}-W_{(i-1)t/n})^2 has mean t/nt/n (since the increment is N(0,t/n)N(0,t/n)) and, using the fourth moment of a normal variable, variance 2(t/n)22(t/n)^2. Summing nn independent such terms, E[Qn]=tE[Q_n]=t exactly and Var(Qn)=2t2/n→0\mathrm{Var}(Q_n)=2t^2/n\to0 as n→∞n\to\infty, so Qn→tQ_n\to t in probability (and mean square) — this proves the quadratic variation identity.

Nowhere differentiability. Suppose, for contradiction, that WW were differentiable at some point ss with derivative W′(s)=LW'(s)=L finite. Then for small hh, Ws+h−Ws≈LhW_{s+h}-W_s\approx Lh, so the squared increment over a tiny interval of length hh would be of order h2h^2 — but the quadratic variation computation above shows the typical squared increment over an interval of length hh is of order hh (much larger than h2h^2 for small hh), and summing these order-hh terms over t/ht/h intervals is exactly what produces a total that converges to the finite, positive number tt rather than to 00.

This mismatch — differentiability would force quadratic variation to be 00, but it is provably t>0t>0 — is irreconcilable, so no point of differentiability can exist; a full measure-theoretic argument (due to Paley, Wiener and Zygmund in 1933) makes this rigorous by directly bounding the probability that any difference quotient stays bounded near any point, showing this probability is exactly zero simultaneously for every point on the path.

Theorem: Itô's lemma

Let WtW_t be a standard Wiener process and let F(t,x)F(t,x) be twice continuously differentiable in xx and once in tt. Then the process F(t,Wt)F(t,W_t) satisfies dF=(∂tF+12∂x2F)dt+∂xF dWtdF=\Big(\partial_tF+\tfrac12\partial_x^2F\Big)dt+\partial_xF\,dW_t — a stochastic chain rule with an extra second-derivative ("Itô correction") term compared to the ordinary chain rule.

Why is it true?

This is the single most-used tool in stochastic calculus: it tells you exactly how to differentiate a function of a random path, which is the starting point for deriving the dynamics of any quantity (an option price, a physical observable) that depends on Brownian motion.

Proof

Taylor expand. For a smooth FF, the ordinary second-order Taylor expansion in both variables over a small step Δt\Delta t with ΔW=Wt+Δt−Wt\Delta W=W_{t+\Delta t}-W_t reads ΔF≈∂tF Δt+∂xF ΔW+12∂x2F (ΔW)2+12∂t2F (Δt)2+∂t∂xF Δt ΔW+⋯\Delta F\approx \partial_tF\,\Delta t+\partial_xF\,\Delta W+\tfrac12\partial_x^2F\,(\Delta W)^2+\tfrac12\partial_t^2F\,(\Delta t)^2+\partial_t\partial_xF\,\Delta t\,\Delta W+\cdots — this much is pure calculus, valid for any smooth path.

Order the terms by size. Because ΔW\Delta W is of order Δt\sqrt{\Delta t} (not Δt\Delta t, as shown in the previous theorem), the terms scale as: ∂tF Δt\partial_tF\,\Delta t is order Δt\Delta t; ∂xF ΔW\partial_xF\,\Delta W is order Δt\sqrt{\Delta t} (the dominant, leading-order random term); (Δt)2(\Delta t)^2 and Δt ΔW\Delta t\,\Delta W are order (Δt)2(\Delta t)^2 and (Δt)3/2(\Delta t)^{3/2} respectively, negligible compared to Δt\Delta t; but (ΔW)2(\Delta W)^2 is order Δt\Delta t — the same order as ∂tF Δt\partial_tF\,\Delta t, not negligible at all, unlike in ordinary calculus where (Δx)2(\Delta x)^2 is always negligible next to Δx\Delta x.

Replace (ΔW)2(\Delta W)^2 by its mean. By the quadratic variation theorem, summed over many small steps (ΔW)2(\Delta W)^2 behaves like Δt\Delta t (its mean, with fluctuations around that mean vanishing as the steps shrink and are summed), which is the informal justification for the heuristic substitution rule (dWt)2=dt(dW_t)^2=dt in the limit of infinitesimal steps.

Take the limit. Dropping the negligible higher-order terms and substituting (ΔW)2→dt(\Delta W)^2\to dt in the limit Δt→0\Delta t\to0 turns the Taylor expansion into the differential form dF=(∂tF+12∂x2F)dt+∂xF dWtdF=\Big(\partial_tF+\tfrac12\partial_x^2F\Big)dt+\partial_xF\,dW_t exactly as claimed — the 12∂x2F dt\tfrac12\partial_x^2F\,dt term is precisely the contribution that ordinary calculus discards but stochastic calculus must keep.

UndergraduateReal-World Applications and Worked Examples

Beyond describing physical diffusion (pollen in water, heat spreading through a solid, gas molecules mixing), Brownian motion and Itô calculus became the mathematical foundation of modern quantitative finance. A stock price StS_t is commonly modeled as geometric Brownian motion, dSt=μSt dt+σSt dWtdS_t=\mu S_t\,dt+\sigma S_t\,dW_t, where μ\mu is the average growth rate and σ\sigma the volatility. Applying Itô's lemma to the value V(t,St)V(t,S_t) of a financial derivative (like an option) leads to the celebrated Black–Scholes partial differential equation, ∂tV+12σ2S2∂S2V+rS∂SV−rV=0\partial_tV+\tfrac12\sigma^2S^2\partial_S^2V+rS\partial_SV-rV=0, whose solution gives the fair price of options traded on every major exchange. The same mathematics — Wiener processes and Itô's lemma — also underlies models of neuron membrane potentials in neuroscience, interest-rate models in economics, and turbulent dispersion in fluid dynamics.

Example: Expected value and variance of a Brownian displacement

A particle's position follows a standard Wiener process WtW_t in one dimension, starting at W0=0W_0=0. Find E[W5]E[W_5] and Var(W5−W2)\mathrm{Var}(W_5-W_2).

Solution

For the first quantity, use Wt−Ws∼N(0,t−s)W_t-W_s\sim N(0,t-s) with s=0s=0: W5−W0=W5W_5-W_0=W_5 is N(0,5)N(0,5), so E[W5]=0E[W_5]=0 directly — Brownian motion has no drift, so its expected position never moves away from the start.

For the second quantity, apply the same defining property with s=2s=2, t=5t=5: W5−W2∼N(0,5−2)W_5-W_2\sim N(0,5-2), i.e. N(0,3)N(0,3).

The variance of a N(0,σ2)N(0,\sigma^2) random variable is σ2\sigma^2 by definition, so reading off the parameter directly, Var(W5−W2)=3\mathrm{Var}(W_5-W_2)=3.

Notice the variance only depends on the elapsed time t−s=3t-s=3, not on the starting time s=2s=2 itself — this reflects the time-homogeneity built into the definition of the Wiener process.

Example: Applying Itô's lemma to F(t,x)=x2F(t,x)=x^2

Use Itô's lemma to find d(Wt2)d(W_t^2), the stochastic differential of the squared Wiener process, and use the result to confirm E[Wt2]=tE[W_t^2]=t.

Solution

Take F(t,x)=x2F(t,x)=x^2, so ∂tF=0\partial_tF=0, ∂xF=2x\partial_xF=2x, ∂x2F=2\partial_x^2F=2.

Substituting into Itô's lemma dF=(∂tF+12∂x2F)dt+∂xF dWtdF=\Big(\partial_tF+\tfrac12\partial_x^2F\Big)dt+\partial_xF\,dW_t with these derivatives (evaluated at x=Wtx=W_t): d(Wt2)=(0+12⋅2)dt+2Wt dWtd(W_t^2)=\big(0+\tfrac12\cdot2\big)dt+2W_t\,dW_t, which simplifies to d(Wt2)=dt+2Wt dWtd(W_t^2)=dt+2W_t\,dW_t.

Integrating both sides from 00 to tt (using W02=0W_0^2=0): Wt2=t+2∫0tWs dWsW_t^2=t+2\int_0^tW_s\,dW_s.

Taking expectations, and using the key fact that an Itô integral ∫0tWs dWs\int_0^tW_s\,dW_s always has mean zero (it is built from increments independent of the past, so there is no systematic drift to accumulate): E[Wt2]=t+2E[∫0tWs dWs]=t+0=tE[W_t^2]=t+2E\big[\int_0^tW_s\,dW_s\big]=t+0=t, confirming directly what we already knew from Wt∼N(0,t)W_t\sim N(0,t) (whose variance is tt), but this time derived purely from the stochastic calculus machinery rather than from the definition.

For a standard Wiener process, the increment Wt−WsW_t-W_s (for s<ts<t) is distributed as:

The quadratic variation of Brownian motion on [0,t][0,t] is:

Compared to the ordinary chain rule, Itô's lemma dF=(∂tF+12∂x2F)dt+∂xF dWtdF=\Big(\partial_tF+\tfrac12\partial_x^2F\Big)dt+\partial_xF\,dW_t has an extra term because:

The Black–Scholes PDE for option pricing is derived by applying:

References

  1. Ioannis Karatzas, Steven E. Shreve (1991). Brownian Motion and Stochastic Calculus
  2. Kiyosi Itô (1944). Stochastic Integral
  3. Fischer Black, Myron Scholes (1973). The Pricing of Options and Corporate Liabilities