Probability and statistics
Bayesian statistics
An approach to statistics that updates a prior belief into a posterior belief as new data arrives.
IntuitionUpdating beliefs as evidence arrives
Suppose you are handed a coin and asked: is it fair? Before flipping it even once, you might believe it is probably close to fair, but you are not sure — call this your prior belief. Now you flip it 10 times and see 8 heads. That data should shift your belief toward "probably biased toward heads," without throwing away what you believed before. Bayesian statistics is the precise machinery for combining a prior belief with new data to get an updated, posterior belief.
UndergraduateBayesian updating: from prior to posterior
Definition: Prior, likelihood, posterior
Given an unknown parameter , a prior distribution encodes belief about before seeing data . A likelihood says how probable the observed data is for each value of . Bayes' theorem combines them into the posterior distribution: — posterior is proportional to likelihood times prior.
The proportionality symbol hides a normalizing constant (called the marginal likelihood or evidence), which does not depend on and simply rescales the product so that integrates to over all , exactly like any other probability density.
If has prior density and data has likelihood , then the posterior density of given is , where ; equivalently .
Why is it true?
This is nothing more than the definition of a conditional density applied to and jointly: the posterior is the joint density of divided by the marginal density of alone, and the joint density factors as prior times likelihood.
Proof
The joint density of factors two ways: as (prior times likelihood, by the definition of the likelihood as the conditional density of given ) and as (posterior times the marginal density of , by the definition of the posterior as the conditional density of given ). Since both expressions equal the same joint density, .
Solving for gives , provided .
It remains to check is exactly the right normalizing constant: integrating both sides of the factorization over , the left side becomes by definition, and the right side becomes since is a probability density. Both sides agree, confirming is consistent and integrates to as required of a density.
UndergraduateConjugate priors: the Beta-Binomial model
Definition: Conjugate prior
A prior family is conjugate to a likelihood if the posterior stays in the same family as the prior, just with updated parameters. This turns Bayesian updating from an integral (computing ) into simple arithmetic on the family's parameters — the reason conjugate priors were the workhorse of Bayesian statistics before modern computational methods existed.
If is the prior and, given , successes are observed in independent trials (so ), then the posterior is .
Why is it true?
The Binomial likelihood contributes a factor , and the Beta prior contributes ; multiplying these just adds the exponents, landing exactly on the shape of another Beta density.
Proof
The likelihood of observing successes in trials given is . By the posterior-proportionality theorem above, .
The factors and do not depend on , so they can be absorbed into the proportionality: .
This last expression is exactly the kernel (the -dependent part) of a density. Since a probability density on with this kernel has a unique normalizing constant (namely , by definition of the Beta function), the posterior must be exactly , i.e. .
Example: A/B testing a website's click-through rate
Before running an experiment, an analyst places a weakly informative prior on a button's true click-through rate (centered at 0.5, but not very confident). After showing the button to visitors, click it. Find the posterior distribution and its mean.
Solution
By Beta-Binomial conjugacy, the posterior is .
The mean of a distribution is , so the posterior mean is : after seeing the data, the analyst's best point estimate for the click-through rate moved from the prior mean of down to , while still being pulled toward — not equal to — the raw sample proportion , because the prior still contributes some weight.
UndergraduateCredible intervals vs confidence intervals
Definition: Credible interval
A credible interval for is any interval with : it is a direct probability statement about , computed from the posterior after the data has already been observed.
This is a genuinely different object from a frequentist confidence interval. A confidence interval is built from a procedure that, applied to many hypothetical samples, would capture the true (fixed) a known fraction of the time; once a specific interval is computed from the specific data at hand, is either in it or not — there is no probability left to talk about. A credible interval instead treats itself as having a distribution, so "the probability is in " is always a meaningful statement.
| Aspect | Bayesian credible interval | Frequentist confidence interval |
|---|---|---|
| What is random | is treated as random, with a posterior distribution; the interval is fixed once the data is observed | is a fixed unknown constant; the interval itself is the random object, varying from sample to sample |
| Interpretation | Given the observed data, the probability that lies in the interval is | Over repeated sampling, of the intervals constructed this way would contain the true |
| Depends on a prior | Yes — the prior enters directly into the posterior | No — computed only from the likelihood and the sampling distribution |
AdvancedBeyond conjugacy: sampling the posterior
Conjugate priors are elegant, but most realistic models — many parameters, hierarchical structure, non-standard likelihoods — have no conjugate form, so the normalizing constant has no closed-form integral. Markov chain Monte Carlo (MCMC) methods sidestep the integral entirely: they construct a Markov chain whose stationary distribution is exactly the posterior , then simulate that chain and use the resulting samples to approximate any posterior quantity (mean, credible interval, or anything else) by simple averages over the samples, without ever computing .
UndergraduateReal-world applications and worked examples
Bayesian updating shows up wherever a belief needs to be revised in light of noisy evidence: doctors interpreting a diagnostic test, spam filters classifying an email, search-and-rescue teams updating where to look, and spacecraft navigation systems fusing sensor readings all run some version of "posterior likelihood prior."
Example: Why a positive test is not proof of disease
A rare disease affects of a population (). A test has sensitivity () and a false-positive rate (). A random person tests positive. Find .
Solution
By the law of total probability, .
Bayes' theorem then gives .
Despite a test that looks very accurate (99% sensitivity, only 5% false positives), a positive result still means only about a chance of actually having the disease — because the disease is rare, false positives from the large healthy population outnumber true positives from the tiny diseased population. This is exactly Bayesian updating in action: the low prior pulls the posterior far below the test's own accuracy figures.
Example: From a flat prior to a skewed posterior
To estimate a coin's bias , a skeptic starts with a completely uninformative prior (uniform on ). After flipping the coin 10 times and seeing 8 heads, find the posterior distribution, its mean, and its mode.
Solution
By Beta-Binomial conjugacy with , , : the posterior is .
The posterior mean is . The mode of a distribution with is , giving here.
The prior was a flat, symmetric bump with no preferred value; the posterior is a sharply asymmetric bump peaking near , exactly the "peak shift away from a symmetric prior" that the widget at the top of this page illustrates schematically — 8 heads out of 10 flips has genuinely pulled belief toward a biased coin, even though the sample size is small.
With a prior, after observing successes in trials, what is the posterior distribution?
In the medical-test example on this page, a positive result gave only about a posterior chance of disease, far below the test's sensitivity. What does this illustrate?
Which statement correctly describes a Bayesian credible interval, in contrast to a frequentist confidence interval?
For a posterior, what is the posterior mean ?
References
- Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, Donald B. Rubin (2013). Bayesian Data Analysis (3rd ed.)
- Matthew D. Hoffman, Andrew Gelman (2014). The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo · arXiv:1111.4246