For any events A and B with P(A)>0 and P(B)>0, P(A∣B)=P(B)P(B∣A)P(A). More generally, if A1,…,Ak form a partition of the sample space with P(Ai)>0 and P(B)>0, then for each j, P(Aj∣B)=∑i=1kP(B∣Ai)P(Ai)P(B∣Aj)P(Aj).
Why is it true?
Bayes' theorem is the mathematical rule for reversing conditional probabilities: it tells you how to update a prior belief P(A) about a hidden cause A into a posterior belief P(A∣B) after observing evidence B, by weighting the prior by how likely A was to produce B.
Proof sketch
By the definition of conditional probability, P(A∣B)=P(B)P(A∩B) when P(B)>0, and symmetrically P(B∣A)=P(A)P(A∩B) when P(A)>0. Equating the two expressions for P(A∩B) gives P(A∩B)=P(B∣A)P(A), and dividing by P(B) yields P(A∣B)=P(B)P(B∣A)P(A). If A1,…,Ak partition the sample space, the law of total probability expands the denominator as P(B)=∑i=1kP(B∩Ai)=∑i=1kP(B∣Ai)P(Ai).