MathLabs
TheoremProved

The Neyman–Pearson lemma

Statement

Consider testing a simple null hypothesis H0:θ=θ0H_0: \theta = \theta_0 against a simple alternative H1:θ=θ1H_1: \theta = \theta_1, based on data with likelihood function L(θ)L(\theta). Let k≥0k \geq 0 be a constant and let CC be the rejection region C={x:L(θ1)>k L(θ0)}C = \{x : L(\theta_1) > k \, L(\theta_0)\}, chosen so that P(C∣H0)=αP(C \mid H_0) = \alpha. Then among all tests with significance level at most α\alpha, the test with rejection region CC has the largest possible power at θ1\theta_1.

Why is it true?

Rejecting when the likelihood ratio L(θ1)/L(θ0)L(\theta_1)/L(\theta_0) is large means rejecting exactly on the outcomes that are relatively most consistent with θ1\theta_1 compared to θ0\theta_0; spending the fixed "budget" α\alpha of false-positive risk on those outcomes, rather than any others, buys the largest possible amount of true-positive detection power.

Proof sketch

Step 1 (set up an arbitrary competitor test): let CC be the likelihood-ratio rejection region from the statement, with indicator function ϕC\phi_C, and let DD be the rejection region of any other test with P(D∣H0)≤αP(D \mid H_0) \leq \alpha, with indicator function ϕD\phi_D. We must show the power at θ1\theta_1 satisfies P(C∣H1)≥P(D∣H1)P(C \mid H_1) \geq P(D \mid H_1).

Step 2 (a key pointwise inequality): for every outcome xx, the quantity (ϕC(x)−ϕD(x))⋅(L(θ1)−k L(θ0))(\phi_C(x) - \phi_D(x)) \cdot (L(\theta_1) - k \, L(\theta_0)) is never negative. Indeed, if x∈Cx \in C then L(θ1)>k L(θ0)L(\theta_1) > k \, L(\theta_0) and ϕC(x)−ϕD(x)≥0\phi_C(x) - \phi_D(x) \geq 0 since ϕC(x)=1\phi_C(x) = 1; if x∉Cx \notin C then L(θ1)≤k L(θ0)L(\theta_1) \leq k \, L(\theta_0) and ϕC(x)−ϕD(x)≤0\phi_C(x) - \phi_D(x) \leq 0 since ϕC(x)=0\phi_C(x) = 0; in both cases the product of the two factors is ≥0\geq 0.

Step 3 (integrate the inequality): summing (or integrating) this nonnegative quantity over all outcomes xx gives ∑x(ϕC(x)−ϕD(x)) L(θ1)−k∑x(ϕC(x)−ϕD(x)) L(θ0)≥0\sum_x (\phi_C(x) - \phi_D(x)) \, L(\theta_1) - k \sum_x (\phi_C(x) - \phi_D(x)) \, L(\theta_0) \geq 0, which rewrites as (P(C∣H1)−P(D∣H1))−k(P(C∣H0)−P(D∣H0))≥0\big(P(C \mid H_1) - P(D \mid H_1)\big) - k \big(P(C \mid H_0) - P(D \mid H_0)\big) \geq 0.

Step 4 (use the significance levels to conclude): by construction P(C∣H0)=αP(C \mid H_0) = \alpha, and by assumption P(D∣H0)≤αP(D \mid H_0) \leq \alpha, so P(C∣H0)−P(D∣H0)≥0P(C \mid H_0) - P(D \mid H_0) \geq 0; since k≥0k \geq 0, subtracting kk times a nonnegative quantity from Step 3's first bracket only strengthens the inequality P(C∣H1)−P(D∣H1)≥k(P(C∣H0)−P(D∣H0))≥0P(C \mid H_1) - P(D \mid H_1) \geq k \big(P(C \mid H_0) - P(D \mid H_0)\big) \geq 0, so P(C∣H1)≥P(D∣H1)P(C \mid H_1) \geq P(D \mid H_1), as required.

Topics that use this theorem

Step-by-step proofs

No step-by-step proof yet for this theorem.