Let X1,X2,… be independent, identically distributed random variables with E[∣X1∣]<∞ and mean μ=E[X1]. Then the sample mean Xˉn=n1∑i=1nXi converges almost surely to μ: P(limn→∞Xˉn=μ)=1.
Why is it true?
Whereas the weak law says that at any single large step n the average Xˉn is very likely close to μ, the strong law guarantees that along almost every infinite sequence of trials, the running average eventually settles down and stays near μ forever with probability 1.
Proof sketch
Without loss of generality assume μ=0. Truncate by setting Yn=Xn1{∣Xn∣≤n}. Because ∑n=1∞P(∣Xn∣>n)=∑n=1∞P(∣X1∣>n)≤E[∣X1∣]<∞, the Borel–Cantelli lemma implies P(Xn=Yn i.o.)=0. Fubini's theorem gives ∑n=1∞n2Var(Yn)≤∑n=1∞n2E[Yn2]≤2E[∣X1∣]<∞. By Kolmogorov's convergence criterion (derived from Kolmogorov's maximal inequality), ∑n=1∞nYn−E[Yn] converges almost surely. Kronecker's lemma then yields n1∑i=1n(Yi−E[Yi])→0 almost surely, and since E[Yn]→μ=0, it follows that Xˉn→0 almost surely.