MathLabs
TheoremProved

Beta-Binomial conjugacy

Statement

If θ∼Beta(α,β)\theta\sim\mathrm{Beta}(\alpha,\beta) is the prior and, given θ\theta, kk successes are observed in nn independent trials (so k∣θ∼Binomial(n,θ)k\mid\theta\sim\mathrm{Binomial}(n,\theta)), then the posterior is θ∣k∼Beta(α+k, β+n−k)\theta\mid k\sim\mathrm{Beta}(\alpha+k,\ \beta+n-k).

Why is it true?

The Binomial likelihood contributes a factor θk(1−θ)n−k\theta^k(1-\theta)^{n-k}, and the Beta prior contributes θα−1(1−θ)β−1\theta^{\alpha-1}(1-\theta)^{\beta-1}; multiplying these just adds the exponents, landing exactly on the shape of another Beta density.

Proof sketch

The likelihood of observing kk successes in nn trials given θ\theta is L(k∣θ)=(nk)θk(1−θ)n−kL(k\mid\theta)=\binom{n}{k}\theta^k(1-\theta)^{n-k}. By the posterior-proportionality theorem above, π(θ∣k)∝L(k∣θ)π(θ)=(nk)θk(1−θ)n−k⋅θα−1(1−θ)β−1B(α,β)\pi(\theta\mid k)\propto L(k\mid\theta)\pi(\theta)=\binom{n}{k}\theta^k(1-\theta)^{n-k}\cdot\dfrac{\theta^{\alpha-1}(1-\theta)^{\beta-1}}{B(\alpha,\beta)}.

The factors (nk)\binom{n}{k} and B(α,β)B(\alpha,\beta) do not depend on θ\theta, so they can be absorbed into the proportionality: π(θ∣k)∝θk(1−θ)n−k⋅θα−1(1−θ)β−1=θ(α+k)−1(1−θ)(β+n−k)−1\pi(\theta\mid k)\propto\theta^{k}(1-\theta)^{n-k}\cdot\theta^{\alpha-1}(1-\theta)^{\beta-1}=\theta^{(\alpha+k)-1}(1-\theta)^{(\beta+n-k)-1}.

This last expression is exactly the kernel (the θ\theta-dependent part) of a Beta(α+k,β+n−k)\mathrm{Beta}(\alpha+k,\beta+n-k) density. Since a probability density on (0,1)(0,1) with this kernel has a unique normalizing constant (namely 1/B(α+k,β+n−k)1/B(\alpha+k,\beta+n-k), by definition of the Beta function), the posterior must be exactly π(θ∣k)=θ(α+k)−1(1−θ)(β+n−k)−1B(α+k,β+n−k)\pi(\theta\mid k)=\dfrac{\theta^{(\alpha+k)-1}(1-\theta)^{(\beta+n-k)-1}}{B(\alpha+k,\beta+n-k)}, i.e. θ∣k∼Beta(α+k,β+n−k)\theta\mid k\sim\mathrm{Beta}(\alpha+k,\beta+n-k).

Topics that use this theorem

Step-by-step proofs

No step-by-step proof yet for this theorem.

References

  1. Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, Donald B. Rubin (2013). Bayesian Data Analysis (3rd ed.)
  2. Matthew D. Hoffman, Andrew Gelman (2014). The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo · arXiv:1111.4246