MathLabs

Probability and statistics

Bayesian statistics

An approach to statistics that updates a prior belief into a posterior belief as new data arrives.

IntuitionUpdating beliefs as evidence arrives

Suppose you are handed a coin and asked: is it fair? Before flipping it even once, you might believe it is probably close to fair, but you are not sure — call this your prior belief. Now you flip it 10 times and see 8 heads. That data should shift your belief toward "probably biased toward heads," without throwing away what you believed before. Bayesian statistics is the precise machinery for combining a prior belief with new data to get an updated, posterior belief.

A downward-curving bump whose peak is shifted to the right of center, schematically representing a posterior distribution pulled away from a symmetric prior by observed data.
Bayesian Beta–Binomial updating: after observing k≈npk \approx np successes in nn trials, the diffuse grey prior Beta(2,2)\mathrm{Beta}(2,2) sharpens into the violet posterior density around pp.

UndergraduateBayesian updating: from prior to posterior

Definition: Prior, likelihood, posterior

Given an unknown parameter θ\theta, a prior distribution π(θ)\pi(\theta) encodes belief about θ\theta before seeing data xx. A likelihood L(x∣θ)L(x\mid\theta) says how probable the observed data is for each value of θ\theta. Bayes' theorem combines them into the posterior distribution: π(θ∣x)∝L(x∣θ) π(θ)\pi(\theta\mid x)\propto L(x\mid\theta)\,\pi(\theta) — posterior is proportional to likelihood times prior.

π(θ∣x)∝L(x∣θ) π(θ)\pi(\theta\mid x)\propto L(x\mid\theta)\,\pi(\theta)

The proportionality symbol hides a normalizing constant m(x)=∫L(x∣θ′)π(θ′) dθ′m(x)=\int L(x\mid\theta')\pi(\theta')\,d\theta' (called the marginal likelihood or evidence), which does not depend on θ\theta and simply rescales the product L(x∣θ)π(θ)L(x\mid\theta)\pi(\theta) so that π(θ∣x)\pi(\theta\mid x) integrates to 11 over all θ\theta, exactly like any other probability density.

π(θ∣x)=L(x∣θ)π(θ)∫L(x∣θ′)π(θ′) dθ′\pi(\theta\mid x)=\dfrac{L(x\mid\theta)\pi(\theta)}{\int L(x\mid\theta')\pi(\theta')\,d\theta'}

If θ\theta has prior density π(θ)\pi(\theta) and data xx has likelihood L(x∣θ)L(x\mid\theta), then the posterior density of θ\theta given xx is π(θ∣x)=L(x∣θ)π(θ)m(x)\pi(\theta\mid x)=\dfrac{L(x\mid\theta)\pi(\theta)}{m(x)}, where m(x)=∫L(x∣θ′)π(θ′) dθ′m(x)=\int L(x\mid\theta')\pi(\theta')\,d\theta'; equivalently π(θ∣x)∝L(x∣θ)π(θ)\pi(\theta\mid x)\propto L(x\mid\theta)\pi(\theta).

Why is it true?

This is nothing more than the definition of a conditional density applied to θ\theta and xx jointly: the posterior is the joint density of (θ,x)(\theta,x) divided by the marginal density of xx alone, and the joint density factors as prior times likelihood.

Proof

The joint density of (θ,x)(\theta,x) factors two ways: as π(θ)L(x∣θ)\pi(\theta)L(x\mid\theta) (prior times likelihood, by the definition of the likelihood as the conditional density of xx given θ\theta) and as π(θ∣x)m(x)\pi(\theta\mid x)m(x) (posterior times the marginal density of xx, by the definition of the posterior as the conditional density of θ\theta given xx). Since both expressions equal the same joint density, π(θ)L(x∣θ)=π(θ∣x)m(x)\pi(\theta)L(x\mid\theta)=\pi(\theta\mid x)m(x).

Solving for π(θ∣x)\pi(\theta\mid x) gives π(θ∣x)=L(x∣θ)π(θ)m(x)\pi(\theta\mid x)=\dfrac{L(x\mid\theta)\pi(\theta)}{m(x)}, provided m(x)>0m(x)>0.

It remains to check m(x)=∫L(x∣θ′)π(θ′) dθ′m(x)=\int L(x\mid\theta')\pi(\theta')\,d\theta' is exactly the right normalizing constant: integrating both sides of the factorization π(θ)L(x∣θ)=π(θ∣x)m(x)\pi(\theta)L(x\mid\theta)=\pi(\theta\mid x)m(x) over θ\theta, the left side becomes ∫π(θ)L(x∣θ) dθ=m(x)\int \pi(\theta)L(x\mid\theta)\,d\theta=m(x) by definition, and the right side becomes m(x)∫π(θ∣x) dθ=m(x)×1m(x)\int\pi(\theta\mid x)\,d\theta=m(x)\times1 since π(⋅∣x)\pi(\cdot\mid x) is a probability density. Both sides agree, confirming m(x)m(x) is consistent and π(θ∣x)\pi(\theta\mid x) integrates to 11 as required of a density.

UndergraduateConjugate priors: the Beta-Binomial model

Definition: Conjugate prior

A prior family is conjugate to a likelihood if the posterior stays in the same family as the prior, just with updated parameters. This turns Bayesian updating from an integral (computing m(x)m(x)) into simple arithmetic on the family's parameters — the reason conjugate priors were the workhorse of Bayesian statistics before modern computational methods existed.

π(θ)=θα−1(1−θ)β−1B(α,β),0<θ<1\pi(\theta)=\dfrac{\theta^{\alpha-1}(1-\theta)^{\beta-1}}{B(\alpha,\beta)},\qquad 0<\theta<1

If θ∼Beta(α,β)\theta\sim\mathrm{Beta}(\alpha,\beta) is the prior and, given θ\theta, kk successes are observed in nn independent trials (so k∣θ∼Binomial(n,θ)k\mid\theta\sim\mathrm{Binomial}(n,\theta)), then the posterior is θ∣k∼Beta(α+k, β+n−k)\theta\mid k\sim\mathrm{Beta}(\alpha+k,\ \beta+n-k).

Why is it true?

The Binomial likelihood contributes a factor θk(1−θ)n−k\theta^k(1-\theta)^{n-k}, and the Beta prior contributes θα−1(1−θ)β−1\theta^{\alpha-1}(1-\theta)^{\beta-1}; multiplying these just adds the exponents, landing exactly on the shape of another Beta density.

Proof

The likelihood of observing kk successes in nn trials given θ\theta is L(k∣θ)=(nk)θk(1−θ)n−kL(k\mid\theta)=\binom{n}{k}\theta^k(1-\theta)^{n-k}. By the posterior-proportionality theorem above, π(θ∣k)∝L(k∣θ)π(θ)=(nk)θk(1−θ)n−k⋅θα−1(1−θ)β−1B(α,β)\pi(\theta\mid k)\propto L(k\mid\theta)\pi(\theta)=\binom{n}{k}\theta^k(1-\theta)^{n-k}\cdot\dfrac{\theta^{\alpha-1}(1-\theta)^{\beta-1}}{B(\alpha,\beta)}.

The factors (nk)\binom{n}{k} and B(α,β)B(\alpha,\beta) do not depend on θ\theta, so they can be absorbed into the proportionality: π(θ∣k)∝θk(1−θ)n−k⋅θα−1(1−θ)β−1=θ(α+k)−1(1−θ)(β+n−k)−1\pi(\theta\mid k)\propto\theta^{k}(1-\theta)^{n-k}\cdot\theta^{\alpha-1}(1-\theta)^{\beta-1}=\theta^{(\alpha+k)-1}(1-\theta)^{(\beta+n-k)-1}.

This last expression is exactly the kernel (the θ\theta-dependent part) of a Beta(α+k,β+n−k)\mathrm{Beta}(\alpha+k,\beta+n-k) density. Since a probability density on (0,1)(0,1) with this kernel has a unique normalizing constant (namely 1/B(α+k,β+n−k)1/B(\alpha+k,\beta+n-k), by definition of the Beta function), the posterior must be exactly π(θ∣k)=θ(α+k)−1(1−θ)(β+n−k)−1B(α+k,β+n−k)\pi(\theta\mid k)=\dfrac{\theta^{(\alpha+k)-1}(1-\theta)^{(\beta+n-k)-1}}{B(\alpha+k,\beta+n-k)}, i.e. θ∣k∼Beta(α+k,β+n−k)\theta\mid k\sim\mathrm{Beta}(\alpha+k,\beta+n-k).

Example: A/B testing a website's click-through rate

Before running an experiment, an analyst places a weakly informative Beta(2,2)\mathrm{Beta}(2,2) prior on a button's true click-through rate θ\theta (centered at 0.5, but not very confident). After showing the button to n=20n=20 visitors, k=7k=7 click it. Find the posterior distribution and its mean.

Solution

By Beta-Binomial conjugacy, the posterior is Beta(α+k, β+n−k)=Beta(2+7, 2+13)=Beta(9,15)\mathrm{Beta}(\alpha+k,\ \beta+n-k)=\mathrm{Beta}(2+7,\ 2+13)=\mathrm{Beta}(9,15).

The mean of a Beta(a,b)\mathrm{Beta}(a,b) distribution is a/(a+b)a/(a+b), so the posterior mean is 9/(9+15)=9/24=0.3759/(9+15)=9/24=0.375: after seeing the data, the analyst's best point estimate for the click-through rate moved from the prior mean of 0.50.5 down to 0.3750.375, while still being pulled toward — not equal to — the raw sample proportion 7/20=0.357/20=0.35, because the prior still contributes some weight.

UndergraduateCredible intervals vs confidence intervals

Definition: Credible interval

A (1−α)(1-\alpha) credible interval for θ\theta is any interval [L,U][L,U] with ∫LUπ(θ∣x) dθ=1−α\int_L^U \pi(\theta\mid x)\,d\theta=1-\alpha: it is a direct probability statement about θ\theta, computed from the posterior after the data xx has already been observed.

This is a genuinely different object from a frequentist confidence interval. A confidence interval is built from a procedure that, applied to many hypothetical samples, would capture the true (fixed) θ\theta a known fraction of the time; once a specific interval is computed from the specific data at hand, θ\theta is either in it or not — there is no probability left to talk about. A credible interval instead treats θ\theta itself as having a distribution, so "the probability θ\theta is in [L,U][L,U]" is always a meaningful statement.

Credible interval vs confidence interval
AspectBayesian credible intervalFrequentist confidence interval
What is randomθ\theta is treated as random, with a posterior distribution; the interval is fixed once the data is observedθ\theta is a fixed unknown constant; the interval itself is the random object, varying from sample to sample
InterpretationGiven the observed data, the probability that θ\theta lies in the interval is 1−α1-\alphaOver repeated sampling, 1−α1-\alpha of the intervals constructed this way would contain the true θ\theta
Depends on a priorYes — the prior π(θ)\pi(\theta) enters directly into the posteriorNo — computed only from the likelihood and the sampling distribution

AdvancedBeyond conjugacy: sampling the posterior

Conjugate priors are elegant, but most realistic models — many parameters, hierarchical structure, non-standard likelihoods — have no conjugate form, so the normalizing constant m(x)m(x) has no closed-form integral. Markov chain Monte Carlo (MCMC) methods sidestep the integral entirely: they construct a Markov chain whose stationary distribution is exactly the posterior π(θ∣x)\pi(\theta\mid x), then simulate that chain and use the resulting samples to approximate any posterior quantity (mean, credible interval, or anything else) by simple averages over the samples, without ever computing m(x)m(x).

UndergraduateReal-world applications and worked examples

Bayesian updating shows up wherever a belief needs to be revised in light of noisy evidence: doctors interpreting a diagnostic test, spam filters classifying an email, search-and-rescue teams updating where to look, and spacecraft navigation systems fusing sensor readings all run some version of "posterior ∝\propto likelihood ×\times prior."

Example: Why a positive test is not proof of disease

A rare disease affects 1%1\% of a population (P(D)=0.01P(D)=0.01). A test has 99%99\% sensitivity (P(+∣D)=0.99P(+\mid D)=0.99) and a 5%5\% false-positive rate (P(+∣¬D)=0.05P(+\mid \lnot D)=0.05). A random person tests positive. Find P(D∣+)P(D\mid +).

Solution

By the law of total probability, P(+)=P(+∣D)P(D)+P(+∣¬D)P(¬D)=0.99×0.01+0.05×0.99=0.0099+0.0495=0.0594P(+)=P(+\mid D)P(D)+P(+\mid\lnot D)P(\lnot D)=0.99\times0.01+0.05\times0.99=0.0099+0.0495=0.0594.

Bayes' theorem then gives P(D∣+)=P(+∣D)P(D)P(+)=0.00990.0594≈0.167P(D\mid +)=\dfrac{P(+\mid D)P(D)}{P(+)}=\dfrac{0.0099}{0.0594}\approx0.167.

Despite a test that looks very accurate (99% sensitivity, only 5% false positives), a positive result still means only about a 16.7%16.7\% chance of actually having the disease — because the disease is rare, false positives from the large healthy population outnumber true positives from the tiny diseased population. This is exactly Bayesian updating in action: the low prior P(D)=0.01P(D)=0.01 pulls the posterior far below the test's own accuracy figures.

Example: From a flat prior to a skewed posterior

To estimate a coin's bias θ\theta, a skeptic starts with a completely uninformative Beta(1,1)\mathrm{Beta}(1,1) prior (uniform on (0,1)(0,1)). After flipping the coin 10 times and seeing 8 heads, find the posterior distribution, its mean, and its mode.

Solution

By Beta-Binomial conjugacy with α=β=1\alpha=\beta=1, n=10n=10, k=8k=8: the posterior is Beta(1+8, 1+2)=Beta(9,3)\mathrm{Beta}(1+8,\ 1+2)=\mathrm{Beta}(9,3).

The posterior mean is 9/(9+3)=9/12=0.759/(9+3)=9/12=0.75. The mode of a Beta(a,b)\mathrm{Beta}(a,b) distribution with a,b>1a,b>1 is (a−1)/(a+b−2)(a-1)/(a+b-2), giving 8/10=0.88/10=0.8 here.

The prior was a flat, symmetric bump with no preferred value; the posterior is a sharply asymmetric bump peaking near 0.80.8, exactly the "peak shift away from a symmetric prior" that the widget at the top of this page illustrates schematically — 8 heads out of 10 flips has genuinely pulled belief toward a biased coin, even though the sample size is small.

With a Beta(2,3)\mathrm{Beta}(2,3) prior, after observing k=4k=4 successes in n=10n=10 trials, what is the posterior distribution?

In the medical-test example on this page, a positive result gave only about a 16.7%16.7\% posterior chance of disease, far below the test's 99%99\% sensitivity. What does this illustrate?

Which statement correctly describes a 95%95\% Bayesian credible interval, in contrast to a 95%95\% frequentist confidence interval?

For a Beta(9,3)\mathrm{Beta}(9,3) posterior, what is the posterior mean E[θ]E[\theta]?

References

  1. Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, Donald B. Rubin (2013). Bayesian Data Analysis (3rd ed.)
  2. Matthew D. Hoffman, Andrew Gelman (2014). The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo · arXiv:1111.4246