How the probability of an event changes once we know another event has occurred.
IntuitionIntuition: new information changes the odds
Roll two fair six-sided dice and let A be the event that the sum of the two dice equals 8, so ordinarily P(A)=365, since exactly the five pairs (2,6),(3,5),(4,4),(5,3),(6,2) out of 36 equally likely pairs give that sum. Now suppose someone peeks at the first die and reports that it shows a 5; call this event B. Knowing B shrinks the sample space from all 36 pairs down to only the 6 pairs where the first die is 5, and among those 6 pairs only (5,3) has sum 8. The extra information changes the probability of A from 365 to 61 — this updated value, written P(A∣B), is called the conditional probability of A given B.
A branching tree diagram with a root splitting into two conditioning events, each further splitting into outcome branches labeled with conditional probabilities.
A probability tree for a two-stage experiment: the first branching splits the sample space by which condition holds (for example, which urn is picked or whether a patient is sick), and each second-level branch is labeled with a conditional probability such as P(A∣B). Multiplying along a path gives P(A∩B), and summing over all paths that end in A gives the total probability P(A).
SchoolThe definition of conditional probability
Definition: Conditional probability
For two events A and B in the same sample space with P(B)>0, the conditional probability of A given B is P(A∣B)=P(B)P(A∩B). Intuitively, once we know B has occurred, the sample space is restricted to B, and P(A∣B) measures what fraction of that restricted space is also covered by A.
P(A∣B)=P(B)P(A∩B)
Here P(A∩B) is the probability that both A and B happen, and P(B) is the probability of the conditioning event, which must be positive since dividing by P(B)=0 is undefined. Rearranging the definition gives the multiplication rule P(A∩B)=P(B)⋅P(A∣B), and by symmetry also P(A∩B)=P(A)⋅P(B∣A) whenever P(A)>0; this rule extends naturally to chains of several events.
P(A∩B)=P(B)⋅P(A∣B)=P(A)⋅P(B∣A)
A special case worth naming is independence: A and B are independent exactly when learning B gives no new information about A, that is P(A∣B)=P(A), which is equivalent to P(A∩B)=P(A)⋅P(B). When A and B are not independent, the two conditional probabilities P(A∣B) and P(B∣A) are generally different numbers, and confusing them is one of the most common mistakes in probability.
Conditional probability in special cases
Relationship between A and B
Value of P(A∣B)
Reason
Independent
P(A∣B)=P(A)
B carries no information about A
A implies B (that is A⊆B)
P(A∣B)=P(B)P(A)
Since A∩B=A
A and B mutually exclusive
P(A∣B)=0
Since A∩B=∅
UndergraduateThe law of total probability and Bayes' theorem
Let B1,B2,…,Bn be a partition of the sample space Ω, meaning Bi∩Bj=∅ for i=j, B1∪B2∪⋯∪Bn=Ω, and P(Bi)>0 for every i. Then for any event A, P(A)=∑i=1nP(Bi)⋅P(A∣Bi).
Why is it true?
Since the Bi partition Ω, event A is automatically chopped into n disjoint pieces, one inside each Bi; adding up the probability of each piece, expressed via the multiplication rule, recovers all of P(A) without double-counting or missing anything.
Proof
Step 1 (partition A using the Bi): define Ai=A∩Bi for each i=1,…,n. Since the Bi are pairwise disjoint, so are the Ai, and since B1∪⋯∪Bn=Ω⊇A, every outcome of A lies in exactly one Ai, so A=A1∪A2∪⋯∪An with the union disjoint.
Step 2 (add probabilities of disjoint pieces): because probability is additive over pairwise disjoint events, P(A)=P(A1)+P(A2)+⋯+P(An)=∑i=1nP(A∩Bi).
Step 3 (rewrite each term with the multiplication rule): for each i, the multiplication rule gives P(A∩Bi)=P(Bi)⋅P(A∣Bi), which is well defined since P(Bi)>0.
Step 4 (substitute back): replacing every P(A∩Bi) in Step 2 by P(Bi)⋅P(A∣Bi) gives exactly P(A)=∑i=1nP(Bi)⋅P(A∣Bi), as claimed.
Let B1,B2,…,Bn be a partition of Ω with P(Bi)>0 for every i, and let A be an event with P(A)>0. Then for each i, P(Bi∣A)=∑j=1nP(Bj)⋅P(A∣Bj)P(Bi)⋅P(A∣Bi).
Why is it true?
Bayes' theorem reverses the direction of conditioning: it starts from how likely the evidence A is under each hypothesis Bi (often easy to know or measure) and returns how likely each hypothesis is once the evidence A has actually been observed (usually what we really want to know), by re-weighting each prior probability P(Bi) according to how well it explains A.
Proof
Step 1 (write P(Bi∩A) two ways): by the multiplication rule applied in both directions, P(Bi∩A)=P(Bi)⋅P(A∣Bi) and also P(Bi∩A)=P(A)⋅P(Bi∣A), since P(A)>0.
Step 2 (equate the two expressions and solve): setting the two expressions for P(Bi∩A) equal gives P(A)⋅P(Bi∣A)=P(Bi)⋅P(A∣Bi), and dividing both sides by P(A) yields P(Bi∣A)=P(A)P(Bi)⋅P(A∣Bi).
Step 3 (expand the denominator with the law of total probability): since B1,…,Bn partition Ω, the law of total probability gives P(A)=∑j=1nP(Bj)⋅P(A∣Bj).
Step 4 (substitute to finish): replacing P(A) in Step 2 by this sum gives exactly P(Bi∣A)=∑j=1nP(Bj)⋅P(A∣Bj)P(Bi)⋅P(A∣Bi), which is Bayes' theorem.
UndergraduateReal-World Applications and Worked Examples
Bayes' theorem is the mathematical core of medical diagnosis (updating the chance of a disease after a test result), spam filtering (updating the chance an email is spam after seeing its words), legal reasoning (updating the chance of guilt after new evidence), and machine learning classifiers (updating the chance of a label after observing features). The two examples below apply the law of total probability and Bayes' theorem to concrete numbers, including the classic paradox of false positives in medical testing.
Example: Which urn was it drawn from?
There are two urns. Urn I has 3 red balls and 7 blue balls; urn II has 6 red balls and 4 blue balls. An urn is chosen at random with P(Urn I)=0.4 and P(Urn II)=0.6, and then one ball is drawn from the chosen urn. The ball turns out to be red. Find the probability that it came from Urn I.
Solution
Step 1: let R be the event the drawn ball is red. Within Urn I, P(R∣Urn I)=103, and within Urn II, P(R∣Urn II)=106.
Step 2: since Urn I and Urn II partition the choice of urn, the law of total probability gives P(R)=P(Urn I)⋅P(R∣Urn I)+P(Urn II)⋅P(R∣Urn II)=0.4⋅0.3+0.6⋅0.6=0.12+0.36=0.48.
Step 3: Bayes' theorem gives P(Urn I∣R)=P(R)P(Urn I)⋅P(R∣Urn I)=0.480.12=41, so despite Urn I being chosen less often, seeing red still leaves only a 25% chance it came from Urn I, because Urn II produces red balls twice as reliably.
Example: The false-positive paradox in medical testing
A disease affects 1% of a population, so P(D)=0.01 for a randomly chosen person. A test for the disease has sensitivity 99%, meaning P(+∣D)=0.99, and specificity 95%, meaning P(−∣Dc)=0.95 (so the false-positive rate is P(+∣Dc)=0.05). A random person tests positive. Find P(D∣+), the probability that they actually have the disease.
Solution
Step 1: D and Dc (has the disease, does not have the disease) partition the population, with P(D)=0.01 and P(Dc)=0.99.
Step 2: apply the law of total probability to the event of testing positive, P(+)=P(D)⋅P(+∣D)+P(Dc)⋅P(+∣Dc)=0.01⋅0.99+0.99⋅0.05=0.0099+0.0495=0.0594.
Step 3: Bayes' theorem gives P(D∣+)=P(+)P(D)⋅P(+∣D)=0.05940.0099=61, only about 16.7%.
Step 4: this is the false-positive paradox: even with a test that is 99% sensitive and 95% specific, most positive results are still false alarms, because the disease is rare so the 99% of healthy people contribute far more false positives (0.0495) than the 1% of sick people contribute true positives (0.0099).
Given P(A)=0.4, P(B)=0.5, and P(A∩B)=0.2, find P(A∣B).
A screening test has sensitivity 99% and specificity 99%. In a population where 1% of people have the disease, a randomly chosen person tests positive. What is P(D∣+), rounded to the nearest whole percent?
Let B1,B2,…,Bn be a partition of the sample space Ω with P(Bi)>0 for every i. Which formula correctly gives P(A) for any event A?
Events A and B satisfy P(B)>0. Which equality is exactly the definition of A and B being independent?