MathLabs
TheoremProved

Bayes' theorem for a posterior distribution

Statement

If θ\theta has prior density π(θ)\pi(\theta) and data xx has likelihood L(x∣θ)L(x\mid\theta), then the posterior density of θ\theta given xx is π(θ∣x)=L(x∣θ)π(θ)m(x)\pi(\theta\mid x)=\dfrac{L(x\mid\theta)\pi(\theta)}{m(x)}, where m(x)=∫L(x∣θ′)π(θ′) dθ′m(x)=\int L(x\mid\theta')\pi(\theta')\,d\theta'; equivalently π(θ∣x)∝L(x∣θ)π(θ)\pi(\theta\mid x)\propto L(x\mid\theta)\pi(\theta).

Why is it true?

This is nothing more than the definition of a conditional density applied to θ\theta and xx jointly: the posterior is the joint density of (θ,x)(\theta,x) divided by the marginal density of xx alone, and the joint density factors as prior times likelihood.

Proof sketch

The joint density of (θ,x)(\theta,x) factors two ways: as π(θ)L(x∣θ)\pi(\theta)L(x\mid\theta) (prior times likelihood, by the definition of the likelihood as the conditional density of xx given θ\theta) and as π(θ∣x)m(x)\pi(\theta\mid x)m(x) (posterior times the marginal density of xx, by the definition of the posterior as the conditional density of θ\theta given xx). Since both expressions equal the same joint density, π(θ)L(x∣θ)=π(θ∣x)m(x)\pi(\theta)L(x\mid\theta)=\pi(\theta\mid x)m(x).

Solving for π(θ∣x)\pi(\theta\mid x) gives π(θ∣x)=L(x∣θ)π(θ)m(x)\pi(\theta\mid x)=\dfrac{L(x\mid\theta)\pi(\theta)}{m(x)}, provided m(x)>0m(x)>0.

It remains to check m(x)=∫L(x∣θ′)π(θ′) dθ′m(x)=\int L(x\mid\theta')\pi(\theta')\,d\theta' is exactly the right normalizing constant: integrating both sides of the factorization π(θ)L(x∣θ)=π(θ∣x)m(x)\pi(\theta)L(x\mid\theta)=\pi(\theta\mid x)m(x) over θ\theta, the left side becomes ∫π(θ)L(x∣θ) dθ=m(x)\int \pi(\theta)L(x\mid\theta)\,d\theta=m(x) by definition, and the right side becomes m(x)∫π(θ∣x) dθ=m(x)×1m(x)\int\pi(\theta\mid x)\,d\theta=m(x)\times1 since π(⋅∣x)\pi(\cdot\mid x) is a probability density. Both sides agree, confirming m(x)m(x) is consistent and π(θ∣x)\pi(\theta\mid x) integrates to 11 as required of a density.

Topics that use this theorem

Step-by-step proofs

No step-by-step proof yet for this theorem.

References

  1. Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, Donald B. Rubin (2013). Bayesian Data Analysis (3rd ed.)
  2. Matthew D. Hoffman, Andrew Gelman (2014). The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo · arXiv:1111.4246