Probability and statistics
The law of large numbers and the central limit theorem
Why averages of independent random variables settle down, and why — once standardized — they look like a bell curve: Chebyshev's inequality, the weak and strong laws of large numbers, and the central limit theorem.
IntuitionWhy averages settle down
Flip a fair coin once and you cannot predict heads or tails. Flip it 10,000 times and the fraction of heads is almost certainly close to 0.5, even though every single flip is still unpredictable. Randomness does not disappear — it dilutes: individual outcomes stay random, but their average becomes more and more predictable as you collect more of them.
UndergraduateChebyshev's inequality: how far can a random variable stray?
Definition: Chebyshev's inequality
For a random variable with finite mean and variance , Chebyshev's inequality bounds how much probability can sit far from the mean, using nothing but — no assumption on the shape of 's distribution is needed.
For a random variable with finite mean and variance , and any : .
Why is it true?
Markov's inequality applied to the non-negative random variable gives ; the event is exactly .
Proof
We first prove Markov's inequality: for a non-negative random variable and any , the indicator bound holds pointwise, because on the event the right side equals , and elsewhere the right side is . Taking expectations of both sides preserves the inequality: , hence .
Now apply Markov's inequality to the non-negative random variable with threshold : . By definition of variance, , so the right side is exactly .
Finally, the events on either side of Markov's inequality coincide: (both sides are non-negative, so squaring preserves the order). Substituting gives , which is Chebyshev's inequality.
Example: A quick bound
A random variable has mean and variance . Bound .
Solution
: no matter the shape of 's distribution, it strays 4 or more from its mean with probability at most 25%.
Example: Monte Carlo estimation of π
To estimate , a computer samples independent points uniformly in the square and counts what fraction land inside the unit circle. Why does this converge to , and roughly how many samples are needed (via Chebyshev's inequality) so that the estimate is within of with probability at least ?
Solution
Let indicate that the -th point falls inside the unit circle. Since points are uniform on a square of area and the circle has area , . Define the estimator , so .
By the weak law of large numbers proved above, in probability, hence in probability: this is exactly why Monte Carlo simulation works — averaging independent random samples converges to the true expected value.
For a concrete sample size, (a Bernoulli variance), so . Chebyshev's inequality gives ; requiring this to be at most forces — millions of samples for two-decimal accuracy, which is why Monte Carlo methods favor cheaper, higher-variance-reducing variants in practice, but the law of large numbers is what guarantees convergence at all.
Example: Pooling risk in an insurance portfolio
An insurer's annual claim per policyholder has mean and standard deviation (in US dollars; claims are rare but large), and claims across the policyholders in a pool are treated as i.i.d. Using the central limit theorem, estimate a premium per policyholder such that the pool's average claim exceeds it with probability at most about .
Solution
Let be the average claim across the pool. By the weak law of large numbers, as the pool grows, so pooling many independent policyholders makes the average claim predictable even though any single claim is highly volatile — this is the entire economic logic of insurance.
By the central limit theorem proved above, is approximately normal with mean and variance for , giving a standard deviation of dollars — dramatically smaller than the -dollar volatility of a single claim.
For a normal distribution, where is standard normal. So setting the premium at dollars per policyholder keeps the probability of the pool's average claim exceeding the premium at about — the central limit theorem is precisely what lets an insurer convert wildly unpredictable individual risk into a narrow, priceable aggregate risk.
UndergraduateThe weak law of large numbers
Definition: Convergence in probability
A sequence of random variables converges in probability to a constant if for every , as : the chance of a large deviation shrinks to nothing, though for any fixed a large deviation is still possible.
Let be i.i.d. random variables with finite mean . Then the sample mean converges to in probability as .
Why is it true?
(when is finite), so Chebyshev's inequality gives . (The full theorem needs only a finite mean, via a more delicate truncation argument, but the Chebyshev proof is the illuminating special case.)
Proof
Let be the sample mean of i.i.d. copies of with mean and, for this Chebyshev-based proof, finite variance . Linearity of expectation gives .
Independence makes variances add: , since each contributes and the cross terms for vanish by independence.
Fix any and apply Chebyshev's inequality, proved above, to : .
As , the right side for every fixed , which is exactly the definition of convergence in probability. Hence in probability.
A version of this result was first proved by Jacob Bernoulli, published posthumously in 1713 in Ars Conjectandi — the earliest form of the law of large numbers, for the fraction of successes in repeated coin-like trials.
AdvancedThe strong law of large numbers
Definition: Almost sure convergence
A sequence converges almost surely to if : the entire sequence itself settles at for almost every outcome, which is stronger than convergence in probability (which only controls one at a time).
Let be i.i.d. random variables with finite mean . Then almost surely as .
Why is it true?
The classical proof (Kolmogorov, 1933) uses Kolmogorov's inequality — a strengthening of Chebyshev's that controls the whole trajectory of partial sums at once — together with a subsequence argument; it is considerably more delicate than the weak law's one-line Chebyshev proof.
Proof
We sketch the proof under the extra (common textbook) assumption that has a finite fourth moment ; the general statement needs only a finite mean but requires a more delicate truncation argument (Kolmogorov, 1933). After centering, assume (replace by throughout), and let .
Expand over all index quadruples. Independence and kill every term that contains an index appearing exactly once (its factor averages to ), leaving only the terms with all four indices equal and the terms that pair up into two distinct equal indices. This gives for a constant depending only on and .
Dividing by : . Summing over , since converges.
A sum of non-negative random variables with finite expected total sum must itself be finite almost surely (monotone convergence), so almost surely, which forces its individual terms to vanish: , i.e. almost surely. Undoing the centering gives almost surely, the strong law.
AdvancedThe central limit theorem
Let be i.i.d. with mean and finite variance . Then as , the standardized sample mean converges in distribution to the standard normal — regardless of the shape of the original distribution of .
Why is it true?
Intuitively, the standardized sum's moment generating (or characteristic) function converges to that of term by term via a Taylor expansion, because only the mean and variance survive the standardization — all higher moments get washed out as . This is why the same bell shape appears no matter what 's original distribution looked like.
Proof
For a random variable , its characteristic function is . Standardize each by setting , so that , and the standardized sample mean becomes a sum of these: .
Because the are i.i.d., the characteristic function of a sum multiplies: .
Since and , a Taylor expansion of near gives . Substituting : .
Raising to the -th power and using the standard limit : . Since is precisely the characteristic function of the standard normal , Lévy's continuity theorem converts this pointwise convergence of characteristic functions into convergence in distribution: .
| Result | Question answered | Type of statement |
|---|---|---|
| Chebyshev's inequality | How far can stray from ? | Finite-sample bound, any distribution |
| Weak law of large numbers | Does approach ? | Convergence in probability |
| Strong law of large numbers | Does the whole sequence settle at ? | Almost-sure convergence |
| Central limit theorem | What shape do the fluctuations of have? | Convergence in distribution, to |
These convergence results are the engine behind statistics: estimation asks how to turn a sample mean into a confidence interval for , and hypothesis testing asks how to use the same normal approximation to decide between competing claims about a population.
The weak law of large numbers says that as , the sample mean of i.i.d. variables with mean
Chebyshev's inequality requires which assumption about the distribution of ?
By the central limit theorem, the distribution of the standardized sample mean approaches, as ,
After 10 consecutive heads on a fair coin, the law of large numbers implies that
References
- Sheldon Ross (2019). A First Course in Probability
- Rick Durrett (2019). Probability: Theory and Examples
- Andrey Kolmogorov (1933). Grundbegriffe der Wahrscheinlichkeitsrechnung