Consider testing a simple null hypothesis H0:θ=θ0 against a simple alternative H1:θ=θ1, based on data with likelihood function L(θ). Let k≥0 be a constant and let C be the rejection region C={x:L(θ1)>kL(θ0)}, chosen so that P(C∣H0)=α. Then among all tests with significance level at most α, the test with rejection region C has the largest possible power at θ1.
Why is it true?
Rejecting when the likelihood ratio L(θ1)/L(θ0) is large means rejecting exactly on the outcomes that are relatively most consistent with θ1 compared to θ0; spending the fixed "budget" α of false-positive risk on those outcomes, rather than any others, buys the largest possible amount of true-positive detection power.
Proof sketch
Step 1 (set up an arbitrary competitor test): let C be the likelihood-ratio rejection region from the statement, with indicator function ϕC, and let D be the rejection region of any other test with P(D∣H0)≤α, with indicator function ϕD. We must show the power at θ1 satisfies P(C∣H1)≥P(D∣H1).
Step 2 (a key pointwise inequality): for every outcome x, the quantity (ϕC(x)−ϕD(x))⋅(L(θ1)−kL(θ0)) is never negative. Indeed, if x∈C then L(θ1)>kL(θ0) and ϕC(x)−ϕD(x)≥0 since ϕC(x)=1; if x∈/C then L(θ1)≤kL(θ0) and ϕC(x)−ϕD(x)≤0 since ϕC(x)=0; in both cases the product of the two factors is ≥0.
Step 3 (integrate the inequality): summing (or integrating) this nonnegative quantity over all outcomes x gives ∑x(ϕC(x)−ϕD(x))L(θ1)−k∑x(ϕC(x)−ϕD(x))L(θ0)≥0, which rewrites as (P(C∣H1)−P(D∣H1))−k(P(C∣H0)−P(D∣H0))≥0.
Step 4 (use the significance levels to conclude): by construction P(C∣H0)=α, and by assumption P(D∣H0)≤α, so P(C∣H0)−P(D∣H0)≥0; since k≥0, subtracting k times a nonnegative quantity from Step 3's first bracket only strengthens the inequality P(C∣H1)−P(D∣H1)≥k(P(C∣H0)−P(D∣H0))≥0, so P(C∣H1)≥P(D∣H1), as required.