A procedure to decide, from sample data, whether to reject a claim about a population.
IntuitionIntuition: is this evidence convincing?
Suppose a coin is flipped 20 times and lands heads 16 times. If the coin were perfectly fair, getting 16 or more heads out of 20 is quite unlikely — it happens only about 0.6% of the time by chance. That smallness is the whole idea behind hypothesis testing: start by assuming a default claim (here, "the coin is fair"), then ask how surprising the observed data would be if that claim were true. If the data would be very surprising, that is evidence against the claim; if not, the data is simply consistent with it and gives no reason to doubt it.
A straight line through the origin with a moderate positive slope, plotted as a degenerate cubic with only the linear term active.
Two-sided z-test on N(0,1): the red vertical lines mark the critical thresholds ±zα/2; observations falling outside the green acceptance region reject H0.
SchoolHypotheses, errors, and the p-value
Definition: Null and alternative hypotheses
The null hypothesis H0 is the default claim being tested (for example, H0:μ=μ0), and the alternative hypothesis H1 is the claim accepted if H0 is rejected (for example, H1:μ=μ0). A test statistic is computed from the sample, and the test rejects H0 when that statistic falls into a predetermined rejection region, chosen so that rejecting a true H0 happens only rarely.
Two kinds of mistakes are possible. A Type I error rejects H0 when H0 is actually true (a false alarm); its probability is called the significance level α=P(reject H0∣H0 true), and is fixed in advance, commonly α=0.05. A Type II error fails to reject H0 when H1 is actually true (a missed detection); its probability is β=P(fail to reject H0∣H1 true), and 1−β is called the power of the test.
The four outcomes of a hypothesis test
Decision
H0 actually true
H0 actually false
Reject H0
Type I error (probability α)
Correct decision (power 1−β)
Fail to reject H0
Correct decision (probability 1−α)
Type II error (probability β)
Z=σ/nXˉ−μ0
This is the one-sample z-statistic for testing H0:μ=μ0 when the population standard deviation σ is known: Xˉ is the sample mean, n is the sample size, and σ/n is the standard error of the mean. Under H0, and assuming the population is normal (or n is large enough for the central limit theorem to apply), Z follows the standard normal distribution.
p=2P(Z≥∣zobs∣∣H0)
Here zobs is the observed value of the test statistic computed from the actual sample, and this formula gives the two-sided p-value: the probability, assuming H0 is true, of observing a test statistic at least as extreme as zobs in either direction. The decision rule is simple: reject H0 if p≤α, and fail to reject H0 otherwise; a smaller p-value is stronger evidence against H0.
UndergraduateThe Neyman–Pearson lemma and the z-test
Consider testing a simple null hypothesis H0:θ=θ0 against a simple alternative H1:θ=θ1, based on data with likelihood function L(θ). Let k≥0 be a constant and let C be the rejection region C={x:L(θ1)>kL(θ0)}, chosen so that P(C∣H0)=α. Then among all tests with significance level at most α, the test with rejection region C has the largest possible power at θ1.
Why is it true?
Rejecting when the likelihood ratio L(θ1)/L(θ0) is large means rejecting exactly on the outcomes that are relatively most consistent with θ1 compared to θ0; spending the fixed "budget" α of false-positive risk on those outcomes, rather than any others, buys the largest possible amount of true-positive detection power.
Proof
Step 1 (set up an arbitrary competitor test): let C be the likelihood-ratio rejection region from the statement, with indicator function ϕC, and let D be the rejection region of any other test with P(D∣H0)≤α, with indicator function ϕD. We must show the power at θ1 satisfies P(C∣H1)≥P(D∣H1).
Step 2 (a key pointwise inequality): for every outcome x, the quantity (ϕC(x)−ϕD(x))⋅(L(θ1)−kL(θ0)) is never negative. Indeed, if x∈C then L(θ1)>kL(θ0) and ϕC(x)−ϕD(x)≥0 since ϕC(x)=1; if x∈/C then L(θ1)≤kL(θ0) and ϕC(x)−ϕD(x)≤0 since ϕC(x)=0; in both cases the product of the two factors is ≥0.
Step 3 (integrate the inequality): summing (or integrating) this nonnegative quantity over all outcomes x gives ∑x(ϕC(x)−ϕD(x))L(θ1)−k∑x(ϕC(x)−ϕD(x))L(θ0)≥0, which rewrites as (P(C∣H1)−P(D∣H1))−k(P(C∣H0)−P(D∣H0))≥0.
Step 4 (use the significance levels to conclude): by construction P(C∣H0)=α, and by assumption P(D∣H0)≤α, so P(C∣H0)−P(D∣H0)≥0; since k≥0, subtracting k times a nonnegative quantity from Step 3's first bracket only strengthens the inequality P(C∣H1)−P(D∣H1)≥k(P(C∣H0)−P(D∣H0))≥0, so P(C∣H1)≥P(D∣H1), as required.
Let X1,…,Xn be independent, identically distributed as N(μ0,σ2) with σ known, and let Z=σ/nXˉ−μ0. Let zα/2 be the value with P(Z>zα/2)=α/2 for the standard normal distribution. Then the test that rejects H0:μ=μ0 when ∣Z∣>zα/2 satisfies P(reject H0∣H0)=α exactly.
Why is it true?
The rejection region is built directly from the exact distribution of Z under H0, so its probability under H0 can be computed exactly rather than merely bounded, giving a test whose false-positive rate is controlled precisely at the chosen level α, not just approximately.
Proof
Step 1 (distribution of Z under H0): since X1,…,Xn are i.i.d. N(μ0,σ2), the sample mean Xˉ is distributed as N(μ0,σ2/n), so standardizing gives Z=σ/nXˉ−μ0∼N(0,1) exactly.
Step 2 (probability of the rejection event): the rejection event is ∣Z∣>zα/2, which splits into the two disjoint events Z>zα/2 and Z<−zα/2, so P(∣Z∣>zα/2)=P(Z>zα/2)+P(Z<−zα/2).
Step 3 (use symmetry of the standard normal): the standard normal density is symmetric about 0, so P(Z<−zα/2)=P(Z>zα/2); by the definition of zα/2, each of these two probabilities equals α/2.
Step 4 (add the two pieces): substituting into Step 2, P(∣Z∣>zα/2)=α/2+α/2=α, and since this probability was computed exactly under H0 (using Step 1), P(reject H0∣H0)=α exactly, as claimed.
UndergraduateReal-World Applications and Worked Examples
Hypothesis testing underlies clinical trials (deciding whether a new drug outperforms a placebo), A/B testing in software and marketing (deciding whether a new design increases conversions), quality control (deciding whether a production line has drifted out of specification), and the discovery thresholds used in physics (deciding whether an experimental signal is more than noise). The two examples below carry out a full z-test, including computing the test statistic, the p-value, and the final decision.
Example: Testing a bottling machine's mean fill volume
A bottling machine is supposed to fill bottles with a mean of μ0=500 ml, and long records show the fill volume has standard deviation σ=10 ml. A sample of n=25 bottles has sample mean xˉ=495 ml. Test H0:μ=500 against H1:μ=500 at significance level α=0.05.
Solution
Step 1: compute the standard error, σ/n=10/25=10/5=2.
Step 2: compute the test statistic, Z=σ/nxˉ−μ0=2495−500=2−5=−2.5.
Step 3: compute the two-sided p-value, p=2P(Z≥2.5)≈2×0.0062=0.0124, using the standard normal distribution.
Step 4: compare p=0.0124 to α=0.05; since p<α, reject H0. There is significant evidence that the machine's mean fill volume differs from 500 ml.
Example: A one-sided test for a biased coin
A coin is flipped n=100 times and lands heads 62 times. Test H0:p=0.5 against the one-sided alternative H1:p>0.5 at significance level α=0.05, using the normal approximation with standard error p0(1−p0)/n under H0.
Solution
Step 1: the sample proportion is p^=62/100=0.62, and under H0 the standard error is p0(1−p0)/n=0.5×0.5/100=0.0025=0.05.
Step 2: compute the test statistic, Z=p0(1−p0)/np^−p0=0.050.62−0.5=0.050.12=2.4.
Step 3: since the alternative is one-sided (H1:p>0.5), the p-value is p=P(Z≥2.4)≈0.0082, using only the upper tail of the standard normal distribution rather than both tails.
Step 4: since p≈0.0082<α=0.05, reject H0. There is significant evidence that the coin is biased toward heads.
A sample of n=16 has mean xˉ=52, and the population is known to have μ0=50 and σ=8. Compute the z-statistic Z=σ/nxˉ−μ0.
In a hypothesis test, a Type I error occurs when:
A test is performed at significance level α=0.05 and gives a p-value of 0.03. What is the correct conclusion?
According to the Neyman–Pearson lemma, for testing a simple H0:θ=θ0 against a simple H1:θ=θ1 at a fixed significance level α, which test is the most powerful?