MathLabs

Probability and statistics

Hypothesis testing

A procedure to decide, from sample data, whether to reject a claim about a population.

IntuitionIntuition: is this evidence convincing?

Suppose a coin is flipped 2020 times and lands heads 1616 times. If the coin were perfectly fair, getting 1616 or more heads out of 2020 is quite unlikely — it happens only about 0.6%0.6\% of the time by chance. That smallness is the whole idea behind hypothesis testing: start by assuming a default claim (here, "the coin is fair"), then ask how surprising the observed data would be if that claim were true. If the data would be very surprising, that is evidence against the claim; if not, the data is simply consistent with it and gives no reason to doubt it.

A straight line through the origin with a moderate positive slope, plotted as a degenerate cubic with only the linear term active.
Two-sided zz-test on N(0,1)\mathcal{N}(0,1): the red vertical lines mark the critical thresholds ±zα/2\pm z_{\alpha/2}; observations falling outside the green acceptance region reject H0H_0.

SchoolHypotheses, errors, and the p-value

Definition: Null and alternative hypotheses

The null hypothesis H0H_0 is the default claim being tested (for example, H0:μ=μ0H_0: \mu = \mu_0), and the alternative hypothesis H1H_1 is the claim accepted if H0H_0 is rejected (for example, H1:μ≠μ0H_1: \mu \neq \mu_0). A test statistic is computed from the sample, and the test rejects H0H_0 when that statistic falls into a predetermined rejection region, chosen so that rejecting a true H0H_0 happens only rarely.

Two kinds of mistakes are possible. A Type I error rejects H0H_0 when H0H_0 is actually true (a false alarm); its probability is called the significance level α=P(reject H0∣H0 true)\alpha = P(\text{reject } H_0 \mid H_0 \text{ true}), and is fixed in advance, commonly α=0.05\alpha = 0.05. A Type II error fails to reject H0H_0 when H1H_1 is actually true (a missed detection); its probability is β=P(fail to reject H0∣H1 true)\beta = P(\text{fail to reject } H_0 \mid H_1 \text{ true}), and 1−β1 - \beta is called the power of the test.

The four outcomes of a hypothesis test
DecisionH0H_0 actually trueH0H_0 actually false
Reject H0H_0Type I error (probability α\alpha)Correct decision (power 1−β1 - \beta)
Fail to reject H0H_0Correct decision (probability 1−α1 - \alpha)Type II error (probability β\beta)
Z=Xˉ−μ0σ/nZ = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}}

This is the one-sample z-statistic for testing H0:μ=μ0H_0: \mu = \mu_0 when the population standard deviation σ\sigma is known: Xˉ\bar{X} is the sample mean, nn is the sample size, and σ/n\sigma / \sqrt{n} is the standard error of the mean. Under H0H_0, and assuming the population is normal (or nn is large enough for the central limit theorem to apply), ZZ follows the standard normal distribution.

p=2 P(Z≥∣zobs∣∣H0)p = 2 \, P(Z \geq |z_{\text{obs}}| \mid H_0)

Here zobsz_{\text{obs}} is the observed value of the test statistic computed from the actual sample, and this formula gives the two-sided p-value: the probability, assuming H0H_0 is true, of observing a test statistic at least as extreme as zobsz_{\text{obs}} in either direction. The decision rule is simple: reject H0H_0 if p≤αp \leq \alpha, and fail to reject H0H_0 otherwise; a smaller p-value is stronger evidence against H0H_0.

UndergraduateThe Neyman–Pearson lemma and the z-test

Consider testing a simple null hypothesis H0:θ=θ0H_0: \theta = \theta_0 against a simple alternative H1:θ=θ1H_1: \theta = \theta_1, based on data with likelihood function L(θ)L(\theta). Let k≥0k \geq 0 be a constant and let CC be the rejection region C={x:L(θ1)>k L(θ0)}C = \{x : L(\theta_1) > k \, L(\theta_0)\}, chosen so that P(C∣H0)=αP(C \mid H_0) = \alpha. Then among all tests with significance level at most α\alpha, the test with rejection region CC has the largest possible power at θ1\theta_1.

Why is it true?

Rejecting when the likelihood ratio L(θ1)/L(θ0)L(\theta_1)/L(\theta_0) is large means rejecting exactly on the outcomes that are relatively most consistent with θ1\theta_1 compared to θ0\theta_0; spending the fixed "budget" α\alpha of false-positive risk on those outcomes, rather than any others, buys the largest possible amount of true-positive detection power.

Proof

Step 1 (set up an arbitrary competitor test): let CC be the likelihood-ratio rejection region from the statement, with indicator function ϕC\phi_C, and let DD be the rejection region of any other test with P(D∣H0)≤αP(D \mid H_0) \leq \alpha, with indicator function ϕD\phi_D. We must show the power at θ1\theta_1 satisfies P(C∣H1)≥P(D∣H1)P(C \mid H_1) \geq P(D \mid H_1).

Step 2 (a key pointwise inequality): for every outcome xx, the quantity (ϕC(x)−ϕD(x))⋅(L(θ1)−k L(θ0))(\phi_C(x) - \phi_D(x)) \cdot (L(\theta_1) - k \, L(\theta_0)) is never negative. Indeed, if x∈Cx \in C then L(θ1)>k L(θ0)L(\theta_1) > k \, L(\theta_0) and ϕC(x)−ϕD(x)≥0\phi_C(x) - \phi_D(x) \geq 0 since ϕC(x)=1\phi_C(x) = 1; if x∉Cx \notin C then L(θ1)≤k L(θ0)L(\theta_1) \leq k \, L(\theta_0) and ϕC(x)−ϕD(x)≤0\phi_C(x) - \phi_D(x) \leq 0 since ϕC(x)=0\phi_C(x) = 0; in both cases the product of the two factors is ≥0\geq 0.

Step 3 (integrate the inequality): summing (or integrating) this nonnegative quantity over all outcomes xx gives ∑x(ϕC(x)−ϕD(x)) L(θ1)−k∑x(ϕC(x)−ϕD(x)) L(θ0)≥0\sum_x (\phi_C(x) - \phi_D(x)) \, L(\theta_1) - k \sum_x (\phi_C(x) - \phi_D(x)) \, L(\theta_0) \geq 0, which rewrites as (P(C∣H1)−P(D∣H1))−k(P(C∣H0)−P(D∣H0))≥0\big(P(C \mid H_1) - P(D \mid H_1)\big) - k \big(P(C \mid H_0) - P(D \mid H_0)\big) \geq 0.

Step 4 (use the significance levels to conclude): by construction P(C∣H0)=αP(C \mid H_0) = \alpha, and by assumption P(D∣H0)≤αP(D \mid H_0) \leq \alpha, so P(C∣H0)−P(D∣H0)≥0P(C \mid H_0) - P(D \mid H_0) \geq 0; since k≥0k \geq 0, subtracting kk times a nonnegative quantity from Step 3's first bracket only strengthens the inequality P(C∣H1)−P(D∣H1)≥k(P(C∣H0)−P(D∣H0))≥0P(C \mid H_1) - P(D \mid H_1) \geq k \big(P(C \mid H_0) - P(D \mid H_0)\big) \geq 0, so P(C∣H1)≥P(D∣H1)P(C \mid H_1) \geq P(D \mid H_1), as required.

Let X1,…,XnX_1, \ldots, X_n be independent, identically distributed as N(μ0,σ2)N(\mu_0, \sigma^2) with σ\sigma known, and let Z=Xˉ−μ0σ/nZ = \frac{\bar{X} - \mu_0}{\sigma / \sqrt{n}}. Let zα/2z_{\alpha/2} be the value with P(Z>zα/2)=α/2P(Z > z_{\alpha/2}) = \alpha/2 for the standard normal distribution. Then the test that rejects H0:μ=μ0H_0: \mu = \mu_0 when ∣Z∣>zα/2|Z| > z_{\alpha/2} satisfies P(reject H0∣H0)=αP(\text{reject } H_0 \mid H_0) = \alpha exactly.

Why is it true?

The rejection region is built directly from the exact distribution of ZZ under H0H_0, so its probability under H0H_0 can be computed exactly rather than merely bounded, giving a test whose false-positive rate is controlled precisely at the chosen level α\alpha, not just approximately.

Proof

Step 1 (distribution of ZZ under H0H_0): since X1,…,XnX_1, \ldots, X_n are i.i.d. N(μ0,σ2)N(\mu_0, \sigma^2), the sample mean Xˉ\bar{X} is distributed as N(μ0,σ2/n)N(\mu_0, \sigma^2/n), so standardizing gives Z=Xˉ−μ0σ/n∼N(0,1)Z = \frac{\bar{X} - \mu_0}{\sigma/\sqrt{n}} \sim N(0, 1) exactly.

Step 2 (probability of the rejection event): the rejection event is ∣Z∣>zα/2|Z| > z_{\alpha/2}, which splits into the two disjoint events Z>zα/2Z > z_{\alpha/2} and Z<−zα/2Z < -z_{\alpha/2}, so P(∣Z∣>zα/2)=P(Z>zα/2)+P(Z<−zα/2)P(|Z| > z_{\alpha/2}) = P(Z > z_{\alpha/2}) + P(Z < -z_{\alpha/2}).

Step 3 (use symmetry of the standard normal): the standard normal density is symmetric about 00, so P(Z<−zα/2)=P(Z>zα/2)P(Z < -z_{\alpha/2}) = P(Z > z_{\alpha/2}); by the definition of zα/2z_{\alpha/2}, each of these two probabilities equals α/2\alpha/2.

Step 4 (add the two pieces): substituting into Step 2, P(∣Z∣>zα/2)=α/2+α/2=αP(|Z| > z_{\alpha/2}) = \alpha/2 + \alpha/2 = \alpha, and since this probability was computed exactly under H0H_0 (using Step 1), P(reject H0∣H0)=αP(\text{reject } H_0 \mid H_0) = \alpha exactly, as claimed.

UndergraduateReal-World Applications and Worked Examples

Hypothesis testing underlies clinical trials (deciding whether a new drug outperforms a placebo), A/B testing in software and marketing (deciding whether a new design increases conversions), quality control (deciding whether a production line has drifted out of specification), and the discovery thresholds used in physics (deciding whether an experimental signal is more than noise). The two examples below carry out a full z-test, including computing the test statistic, the p-value, and the final decision.

Example: Testing a bottling machine's mean fill volume

A bottling machine is supposed to fill bottles with a mean of μ0=500\mu_0 = 500 ml, and long records show the fill volume has standard deviation σ=10\sigma = 10 ml. A sample of n=25n = 25 bottles has sample mean xˉ=495\bar{x} = 495 ml. Test H0:μ=500H_0: \mu = 500 against H1:μ≠500H_1: \mu \neq 500 at significance level α=0.05\alpha = 0.05.

Solution

Step 1: compute the standard error, σ/n=10/25=10/5=2\sigma / \sqrt{n} = 10 / \sqrt{25} = 10 / 5 = 2.

Step 2: compute the test statistic, Z=xˉ−μ0σ/n=495−5002=−52=−2.5Z = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}} = \frac{495 - 500}{2} = \frac{-5}{2} = -2.5.

Step 3: compute the two-sided p-value, p=2 P(Z≥2.5)≈2×0.0062=0.0124p = 2 \, P(Z \geq 2.5) \approx 2 \times 0.0062 = 0.0124, using the standard normal distribution.

Step 4: compare p=0.0124p = 0.0124 to α=0.05\alpha = 0.05; since p<αp < \alpha, reject H0H_0. There is significant evidence that the machine's mean fill volume differs from 500500 ml.

Example: A one-sided test for a biased coin

A coin is flipped n=100n = 100 times and lands heads 6262 times. Test H0:p=0.5H_0: p = 0.5 against the one-sided alternative H1:p>0.5H_1: p > 0.5 at significance level α=0.05\alpha = 0.05, using the normal approximation with standard error p0(1−p0)/n\sqrt{p_0 (1 - p_0) / n} under H0H_0.

Solution

Step 1: the sample proportion is p^=62/100=0.62\hat{p} = 62/100 = 0.62, and under H0H_0 the standard error is p0(1−p0)/n=0.5×0.5/100=0.0025=0.05\sqrt{p_0 (1 - p_0) / n} = \sqrt{0.5 \times 0.5 / 100} = \sqrt{0.0025} = 0.05.

Step 2: compute the test statistic, Z=p^−p0p0(1−p0)/n=0.62−0.50.05=0.120.05=2.4Z = \frac{\hat{p} - p_0}{\sqrt{p_0 (1 - p_0) / n}} = \frac{0.62 - 0.5}{0.05} = \frac{0.12}{0.05} = 2.4.

Step 3: since the alternative is one-sided (H1:p>0.5H_1: p > 0.5), the p-value is p=P(Z≥2.4)≈0.0082p = P(Z \geq 2.4) \approx 0.0082, using only the upper tail of the standard normal distribution rather than both tails.

Step 4: since p≈0.0082<α=0.05p \approx 0.0082 < \alpha = 0.05, reject H0H_0. There is significant evidence that the coin is biased toward heads.

A sample of n=16n = 16 has mean xˉ=52\bar{x} = 52, and the population is known to have μ0=50\mu_0 = 50 and σ=8\sigma = 8. Compute the z-statistic Z=xˉ−μ0σ/nZ = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}.

In a hypothesis test, a Type I error occurs when:

A test is performed at significance level α=0.05\alpha = 0.05 and gives a p-value of 0.030.03. What is the correct conclusion?

According to the Neyman–Pearson lemma, for testing a simple H0:θ=θ0H_0: \theta = \theta_0 against a simple H1:θ=θ1H_1: \theta = \theta_1 at a fixed significance level α\alpha, which test is the most powerful?