MathLabs
TheoremProved

Bayes' theorem

Statement

Let B1,B2,…,BnB_1, B_2, \ldots, B_n be a partition of Ω\Omega with P(Bi)>0P(B_i) > 0 for every ii, and let AA be an event with P(A)>0P(A) > 0. Then for each ii, P(Bi∣A)=P(Bi)⋅P(A∣Bi)∑j=1nP(Bj)⋅P(A∣Bj)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j)}.

Why is it true?

Bayes' theorem reverses the direction of conditioning: it starts from how likely the evidence AA is under each hypothesis BiB_i (often easy to know or measure) and returns how likely each hypothesis is once the evidence AA has actually been observed (usually what we really want to know), by re-weighting each prior probability P(Bi)P(B_i) according to how well it explains AA.

Proof sketch

Step 1 (write P(Bi∩A)P(B_i \cap A) two ways): by the multiplication rule applied in both directions, P(Bi∩A)=P(Bi)⋅P(A∣Bi)P(B_i \cap A) = P(B_i) \cdot P(A \mid B_i) and also P(Bi∩A)=P(A)⋅P(Bi∣A)P(B_i \cap A) = P(A) \cdot P(B_i \mid A), since P(A)>0P(A) > 0.

Step 2 (equate the two expressions and solve): setting the two expressions for P(Bi∩A)P(B_i \cap A) equal gives P(A)⋅P(Bi∣A)=P(Bi)⋅P(A∣Bi)P(A) \cdot P(B_i \mid A) = P(B_i) \cdot P(A \mid B_i), and dividing both sides by P(A)P(A) yields P(Bi∣A)=P(Bi)⋅P(A∣Bi)P(A)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{P(A)}.

Step 3 (expand the denominator with the law of total probability): since B1,…,BnB_1, \ldots, B_n partition Ω\Omega, the law of total probability gives P(A)=∑j=1nP(Bj)⋅P(A∣Bj)P(A) = \sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j).

Step 4 (substitute to finish): replacing P(A)P(A) in Step 2 by this sum gives exactly P(Bi∣A)=P(Bi)⋅P(A∣Bi)∑j=1nP(Bj)⋅P(A∣Bj)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j)}, which is Bayes' theorem.

Topics that use this theorem

Step-by-step proofs

No step-by-step proof yet for this theorem.