Let B1,B2,…,Bn be a partition of Ω with P(Bi)>0 for every i, and let A be an event with P(A)>0. Then for each i, P(Bi∣A)=∑j=1nP(Bj)⋅P(A∣Bj)P(Bi)⋅P(A∣Bi).
Why is it true?
Bayes' theorem reverses the direction of conditioning: it starts from how likely the evidence A is under each hypothesis Bi (often easy to know or measure) and returns how likely each hypothesis is once the evidence A has actually been observed (usually what we really want to know), by re-weighting each prior probability P(Bi) according to how well it explains A.
Proof sketch
Step 1 (write P(Bi∩A) two ways): by the multiplication rule applied in both directions, P(Bi∩A)=P(Bi)⋅P(A∣Bi) and also P(Bi∩A)=P(A)⋅P(Bi∣A), since P(A)>0.
Step 2 (equate the two expressions and solve): setting the two expressions for P(Bi∩A) equal gives P(A)⋅P(Bi∣A)=P(Bi)⋅P(A∣Bi), and dividing both sides by P(A) yields P(Bi∣A)=P(A)P(Bi)⋅P(A∣Bi).
Step 3 (expand the denominator with the law of total probability): since B1,…,Bn partition Ω, the law of total probability gives P(A)=∑j=1nP(Bj)⋅P(A∣Bj).
Step 4 (substitute to finish): replacing P(A) in Step 2 by this sum gives exactly P(Bi∣A)=∑j=1nP(Bj)⋅P(A∣Bj)P(Bi)⋅P(A∣Bi), which is Bayes' theorem.