MathLabs

Grade 12

Conditional probability and Bayes' theorem

How the probability of an event changes once we know another event has occurred.

IntuitionIntuition: new information changes the odds

Roll two fair six-sided dice and let AA be the event that the sum of the two dice equals 88, so ordinarily P(A)=536P(A) = \frac{5}{36}, since exactly the five pairs (2,6),(3,5),(4,4),(5,3),(6,2)(2,6), (3,5), (4,4), (5,3), (6,2) out of 3636 equally likely pairs give that sum. Now suppose someone peeks at the first die and reports that it shows a 55; call this event BB. Knowing BB shrinks the sample space from all 3636 pairs down to only the 66 pairs where the first die is 55, and among those 66 pairs only (5,3)(5,3) has sum 88. The extra information changes the probability of AA from 536\frac{5}{36} to 16\frac{1}{6} — this updated value, written P(A∣B)P(A \mid B), is called the conditional probability of AA given BB.

A branching tree diagram with a root splitting into two conditioning events, each further splitting into outcome branches labeled with conditional probabilities.
A probability tree for a two-stage experiment: the first branching splits the sample space by which condition holds (for example, which urn is picked or whether a patient is sick), and each second-level branch is labeled with a conditional probability such as P(A∣B)P(A \mid B). Multiplying along a path gives P(A∩B)P(A \cap B), and summing over all paths that end in AA gives the total probability P(A)P(A).

SchoolThe definition of conditional probability

Definition: Conditional probability

For two events AA and BB in the same sample space with P(B)>0P(B) > 0, the conditional probability of AA given BB is P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}. Intuitively, once we know BB has occurred, the sample space is restricted to BB, and P(A∣B)P(A \mid B) measures what fraction of that restricted space is also covered by AA.

P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}

Here P(A∩B)P(A \cap B) is the probability that both AA and BB happen, and P(B)P(B) is the probability of the conditioning event, which must be positive since dividing by P(B)=0P(B) = 0 is undefined. Rearranging the definition gives the multiplication rule P(A∩B)=P(B)⋅P(A∣B)P(A \cap B) = P(B) \cdot P(A \mid B), and by symmetry also P(A∩B)=P(A)⋅P(B∣A)P(A \cap B) = P(A) \cdot P(B \mid A) whenever P(A)>0P(A) > 0; this rule extends naturally to chains of several events.

P(A∩B)=P(B)⋅P(A∣B)=P(A)⋅P(B∣A)P(A \cap B) = P(B) \cdot P(A \mid B) = P(A) \cdot P(B \mid A)

A special case worth naming is independence: AA and BB are independent exactly when learning BB gives no new information about AA, that is P(A∣B)=P(A)P(A \mid B) = P(A), which is equivalent to P(A∩B)=P(A)⋅P(B)P(A \cap B) = P(A) \cdot P(B). When AA and BB are not independent, the two conditional probabilities P(A∣B)P(A \mid B) and P(B∣A)P(B \mid A) are generally different numbers, and confusing them is one of the most common mistakes in probability.

Conditional probability in special cases
Relationship between AA and BBValue of P(A∣B)P(A \mid B)Reason
IndependentP(A∣B)=P(A)P(A \mid B) = P(A)BB carries no information about AA
AA implies BB (that is A⊆BA \subseteq B)P(A∣B)=P(A)P(B)P(A \mid B) = \frac{P(A)}{P(B)}Since A∩B=AA \cap B = A
AA and BB mutually exclusiveP(A∣B)=0P(A \mid B) = 0Since A∩B=∅A \cap B = \varnothing

UndergraduateThe law of total probability and Bayes' theorem

Let B1,B2,…,BnB_1, B_2, \ldots, B_n be a partition of the sample space Ω\Omega, meaning Bi∩Bj=∅B_i \cap B_j = \varnothing for i≠ji \neq j, B1∪B2∪⋯∪Bn=ΩB_1 \cup B_2 \cup \cdots \cup B_n = \Omega, and P(Bi)>0P(B_i) > 0 for every ii. Then for any event AA, P(A)=∑i=1nP(Bi)⋅P(A∣Bi)P(A) = \sum_{i=1}^{n} P(B_i) \cdot P(A \mid B_i).

Why is it true?

Since the BiB_i partition Ω\Omega, event AA is automatically chopped into nn disjoint pieces, one inside each BiB_i; adding up the probability of each piece, expressed via the multiplication rule, recovers all of P(A)P(A) without double-counting or missing anything.

Proof

Step 1 (partition AA using the BiB_i): define Ai=A∩BiA_i = A \cap B_i for each i=1,…,ni = 1, \ldots, n. Since the BiB_i are pairwise disjoint, so are the AiA_i, and since B1∪⋯∪Bn=Ω⊇AB_1 \cup \cdots \cup B_n = \Omega \supseteq A, every outcome of AA lies in exactly one AiA_i, so A=A1∪A2∪⋯∪AnA = A_1 \cup A_2 \cup \cdots \cup A_n with the union disjoint.

Step 2 (add probabilities of disjoint pieces): because probability is additive over pairwise disjoint events, P(A)=P(A1)+P(A2)+⋯+P(An)=∑i=1nP(A∩Bi)P(A) = P(A_1) + P(A_2) + \cdots + P(A_n) = \sum_{i=1}^{n} P(A \cap B_i).

Step 3 (rewrite each term with the multiplication rule): for each ii, the multiplication rule gives P(A∩Bi)=P(Bi)⋅P(A∣Bi)P(A \cap B_i) = P(B_i) \cdot P(A \mid B_i), which is well defined since P(Bi)>0P(B_i) > 0.

Step 4 (substitute back): replacing every P(A∩Bi)P(A \cap B_i) in Step 2 by P(Bi)⋅P(A∣Bi)P(B_i) \cdot P(A \mid B_i) gives exactly P(A)=∑i=1nP(Bi)⋅P(A∣Bi)P(A) = \sum_{i=1}^{n} P(B_i) \cdot P(A \mid B_i), as claimed.

Let B1,B2,…,BnB_1, B_2, \ldots, B_n be a partition of Ω\Omega with P(Bi)>0P(B_i) > 0 for every ii, and let AA be an event with P(A)>0P(A) > 0. Then for each ii, P(Bi∣A)=P(Bi)⋅P(A∣Bi)∑j=1nP(Bj)⋅P(A∣Bj)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j)}.

Why is it true?

Bayes' theorem reverses the direction of conditioning: it starts from how likely the evidence AA is under each hypothesis BiB_i (often easy to know or measure) and returns how likely each hypothesis is once the evidence AA has actually been observed (usually what we really want to know), by re-weighting each prior probability P(Bi)P(B_i) according to how well it explains AA.

Proof

Step 1 (write P(Bi∩A)P(B_i \cap A) two ways): by the multiplication rule applied in both directions, P(Bi∩A)=P(Bi)⋅P(A∣Bi)P(B_i \cap A) = P(B_i) \cdot P(A \mid B_i) and also P(Bi∩A)=P(A)⋅P(Bi∣A)P(B_i \cap A) = P(A) \cdot P(B_i \mid A), since P(A)>0P(A) > 0.

Step 2 (equate the two expressions and solve): setting the two expressions for P(Bi∩A)P(B_i \cap A) equal gives P(A)⋅P(Bi∣A)=P(Bi)⋅P(A∣Bi)P(A) \cdot P(B_i \mid A) = P(B_i) \cdot P(A \mid B_i), and dividing both sides by P(A)P(A) yields P(Bi∣A)=P(Bi)⋅P(A∣Bi)P(A)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{P(A)}.

Step 3 (expand the denominator with the law of total probability): since B1,…,BnB_1, \ldots, B_n partition Ω\Omega, the law of total probability gives P(A)=∑j=1nP(Bj)⋅P(A∣Bj)P(A) = \sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j).

Step 4 (substitute to finish): replacing P(A)P(A) in Step 2 by this sum gives exactly P(Bi∣A)=P(Bi)⋅P(A∣Bi)∑j=1nP(Bj)⋅P(A∣Bj)P(B_i \mid A) = \frac{P(B_i) \cdot P(A \mid B_i)}{\sum_{j=1}^{n} P(B_j) \cdot P(A \mid B_j)}, which is Bayes' theorem.

UndergraduateReal-World Applications and Worked Examples

Bayes' theorem is the mathematical core of medical diagnosis (updating the chance of a disease after a test result), spam filtering (updating the chance an email is spam after seeing its words), legal reasoning (updating the chance of guilt after new evidence), and machine learning classifiers (updating the chance of a label after observing features). The two examples below apply the law of total probability and Bayes' theorem to concrete numbers, including the classic paradox of false positives in medical testing.

Example: Which urn was it drawn from?

There are two urns. Urn I has 33 red balls and 77 blue balls; urn II has 66 red balls and 44 blue balls. An urn is chosen at random with P(Urn I)=0.4P(\text{Urn I}) = 0.4 and P(Urn II)=0.6P(\text{Urn II}) = 0.6, and then one ball is drawn from the chosen urn. The ball turns out to be red. Find the probability that it came from Urn I.

Solution

Step 1: let RR be the event the drawn ball is red. Within Urn I, P(R∣Urn I)=310P(R \mid \text{Urn I}) = \frac{3}{10}, and within Urn II, P(R∣Urn II)=610P(R \mid \text{Urn II}) = \frac{6}{10}.

Step 2: since Urn I and Urn II partition the choice of urn, the law of total probability gives P(R)=P(Urn I)⋅P(R∣Urn I)+P(Urn II)⋅P(R∣Urn II)=0.4⋅0.3+0.6⋅0.6=0.12+0.36=0.48P(R) = P(\text{Urn I}) \cdot P(R \mid \text{Urn I}) + P(\text{Urn II}) \cdot P(R \mid \text{Urn II}) = 0.4 \cdot 0.3 + 0.6 \cdot 0.6 = 0.12 + 0.36 = 0.48.

Step 3: Bayes' theorem gives P(Urn I∣R)=P(Urn I)⋅P(R∣Urn I)P(R)=0.120.48=14P(\text{Urn I} \mid R) = \frac{P(\text{Urn I}) \cdot P(R \mid \text{Urn I})}{P(R)} = \frac{0.12}{0.48} = \frac{1}{4}, so despite Urn I being chosen less often, seeing red still leaves only a 25%25\% chance it came from Urn I, because Urn II produces red balls twice as reliably.

Example: The false-positive paradox in medical testing

A disease affects 1%1\% of a population, so P(D)=0.01P(D) = 0.01 for a randomly chosen person. A test for the disease has sensitivity 99%99\%, meaning P(+∣D)=0.99P(+ \mid D) = 0.99, and specificity 95%95\%, meaning P(−∣Dc)=0.95P(- \mid D^c) = 0.95 (so the false-positive rate is P(+∣Dc)=0.05P(+ \mid D^c) = 0.05). A random person tests positive. Find P(D∣+)P(D \mid +), the probability that they actually have the disease.

Solution

Step 1: DD and DcD^c (has the disease, does not have the disease) partition the population, with P(D)=0.01P(D) = 0.01 and P(Dc)=0.99P(D^c) = 0.99.

Step 2: apply the law of total probability to the event of testing positive, P(+)=P(D)⋅P(+∣D)+P(Dc)⋅P(+∣Dc)=0.01⋅0.99+0.99⋅0.05=0.0099+0.0495=0.0594P(+) = P(D) \cdot P(+ \mid D) + P(D^c) \cdot P(+ \mid D^c) = 0.01 \cdot 0.99 + 0.99 \cdot 0.05 = 0.0099 + 0.0495 = 0.0594.

Step 3: Bayes' theorem gives P(D∣+)=P(D)⋅P(+∣D)P(+)=0.00990.0594=16P(D \mid +) = \frac{P(D) \cdot P(+ \mid D)}{P(+)} = \frac{0.0099}{0.0594} = \frac{1}{6}, only about 16.7%16.7\%.

Step 4: this is the false-positive paradox: even with a test that is 99%99\% sensitive and 95%95\% specific, most positive results are still false alarms, because the disease is rare so the 99%99\% of healthy people contribute far more false positives (0.04950.0495) than the 1%1\% of sick people contribute true positives (0.00990.0099).

Given P(A)=0.4P(A) = 0.4, P(B)=0.5P(B) = 0.5, and P(A∩B)=0.2P(A \cap B) = 0.2, find P(A∣B)P(A \mid B).

A screening test has sensitivity 99%99\% and specificity 99%99\%. In a population where 1%1\% of people have the disease, a randomly chosen person tests positive. What is P(D∣+)P(D \mid +), rounded to the nearest whole percent?

Let B1,B2,…,BnB_1, B_2, \ldots, B_n be a partition of the sample space Ω\Omega with P(Bi)>0P(B_i) > 0 for every ii. Which formula correctly gives P(A)P(A) for any event AA?

Events AA and BB satisfy P(B)>0P(B) > 0. Which equality is exactly the definition of AA and BB being independent?