Probability and statistics
Estimation
Using sample data to infer unknown values of a population, with point and interval estimates.
IntuitionFrom a Handful to the Whole
Imagine tasting one spoonful from a large pot of soup to judge how salty the whole pot is. You don't drink the entire pot -- a single representative spoonful (a sample) is enough to guess the salt level of the whole thing (the population). Estimation is the mathematics of turning that spoonful into a trustworthy guess, together with a sense of how far off the guess might be.
SchoolPoint Estimates: A Single Best Guess
Definition: Point estimator
A point estimator is a single number, computed from sample data, used as a best guess for an unknown population parameter. For estimating the population mean , the natural point estimator is the sample mean : the average of the observed values .
Here is the sample size, are the individual sample observations, and is their average. Because different samples give different values of , it is itself a random variable with its own distribution -- called the sampling distribution of the mean -- whose spread shrinks as grows.
This is the confidence interval for when the population standard deviation is known: is the standard-normal critical value that leaves probability in each tail, and is the standard error of , the typical distance between and .
| Confidence level | |
|---|---|
| 90% | |
| 95% | |
| 99% |
UndergraduateWhy the Confidence Interval Works
If are independent, identically distributed with for every , then the sample mean satisfies .
Why is it true?
Averaging removes systematic bias: on average, across many hypothetical samples, lands exactly on the true population mean , neither systematically too high nor too low.
Proof
By definition, . Expectation is a linear operator, so it distributes over the sum and the constant factor : .
Since every comes from the same population, for each of the terms, so the sum collapses to .
Therefore for every sample size : the sample mean is unbiased no matter how small or large the sample is. This is a statement about the center of its sampling distribution only -- the spread of that distribution, , still shrinks as grows, which is what the confidence interval below uses.
If are independent, identically distributed with mean and finite variance , then for large the confidence interval contains with probability approximately .
Why is it true?
The Central Limit Theorem says the standardized sample mean behaves like a standard normal variable once is large, so we can use normal quantiles to bracket with a known long-run success rate, regardless of the shape of the original population.
Proof
By the Central Limit Theorem, converges in distribution to as . So for large , holds with probability approximately , where is the value cutting probability from each tail of .
Multiply all three sides of the inequality by (this preserves the inequality direction): .
is subtracted from all sides, then all sides are multiplied by (flipping the inequalities): . This is exactly the confidence-interval formula: since the underlying probability statement held with approximate probability , so does this rearranged interval containing .
The margin of error shrinks only like , not like : to cut the margin of error in half you must quadruple the sample size , and to cut it to a third you need nine times as many observations. Precision is expensive: each extra digit of accuracy costs far more data than the one before it.
UndergraduateReal-World Applications and Worked Examples
Confidence intervals appear wherever a decision must be made from an incomplete sample: quality control on a factory line, opinion polling before an election, dosage trials in medicine, or calibrating a telescope from repeated measurements. In every case, the same formula turns a sample average into a range that is honest about its own uncertainty.
Example: Confidence interval for a bolt's mean diameter
A quality-control engineer samples bolts and finds a sample mean diameter mm, with known population standard deviation mm. Construct a 95% confidence interval for the true mean diameter , using .
Solution
The standard error is mm.
The margin of error is mm.
The 95% confidence interval is mm: we are 95% confident that the true mean bolt diameter lies in this range.
Example: Estimating the share of voters supporting a policy
A pollster surveys randomly chosen voters and finds a sample proportion in favor of a policy. Treating the sample proportion as a sample mean of 0/1 responses, construct a 95% confidence interval for the true population proportion in favor, using .
Solution
The estimated standard deviation of a single 0/1 response is , so the standard error of the sample proportion is .
The margin of error is .
The 95% confidence interval is : the poll is 95% confident the true support lies between about 48.1% and 57.9%, which is too wide to call the policy majority-supported with certainty -- a larger sample would be needed for a tighter call.
Using the margin-of-error formula , what is the margin of error when , , and ?
A 95% confidence interval for the mean daily temperature is computed from one sample. What is the correct interpretation of "95% confidence"?
The property for every sample size describes which characteristic of the estimator ?
If the sample size increases from 100 to 400, with and unchanged, the margin of error is multiplied by ...
References
- NIST/SEMATECH (2013). Confidence Limits for the Mean
- Diez, D.; Cetinkaya-Rundel, M.; Barr, C. (2019). OpenIntro Statistics