Probability and statistics
Descriptive statistics
Summarizing data with measures like mean, median and standard deviation, and displaying it in charts.
IntuitionIntuition: Telling the Story of Data with a Few Numbers
Imagine a teacher with 30 quiz scores, or a shop owner with a month of daily sales. Nobody wants to stare at 30 separate numbers to understand what happened. Descriptive statistics compresses a whole data set into a handful of numbers — a typical value, how spread out the values are — so we can compare, summarize and spot anything unusual at a glance.
SchoolMeasuring the Center: Mean, Median, Mode
Definition: Mean
The mean (or average) of data values adds them all up and divides by how many there are. It uses every value in the data set, which makes it sensitive to unusually large or small values.
Here is the mean, is the number of data values, and is the -th value being summed. The bar over is standard notation for "the mean of ".
Definition: Median and mode
The median is the middle value when the data is sorted from smallest to largest (or the average of the two middle values, if is even). The mode is the value that occurs most often. Unlike the mean, the median ignores how far away the extreme values are — only their rank matters — which makes it resistant to outliers.
So which measure should you use? If the data is roughly symmetric with no extreme values, the mean is efficient and uses all the information. If the data has outliers or is heavily skewed — like household incomes or house prices — the median gives a more representative "typical" value.
| Measure | Definition | Sensitive to outliers? |
|---|---|---|
| Mean | Sum of values divided by count | Yes — every value counts |
| Median | Middle value when sorted | No — only rank matters |
| Mode | Most frequent value | No |
SchoolMeasuring the Spread: Standard Deviation and Quartiles
Two data sets can have the same mean but look completely different — one tightly clustered, one wildly spread out. The standard deviation captures this by measuring the typical distance of a data value from the mean:
Each deviation is squared (so negative and positive deviations don't cancel out), averaged, and then the square root brings the units back to the original scale of the data. A small means the data clusters tightly around ; a large means it is spread out.
Quartiles split sorted data into four equal parts: (25% of the way through), the median (50%), and (75%). The interquartile range measures the spread of the middle half of the data and, like the median, is resistant to outliers — this is exactly what a boxplot draws: a box from to with a line at the median.
For any data set with mean , .
Why is it true?
The mean is defined precisely as the balance point of the data — think of a see-saw with weights at each — so the values above the mean must exactly offset the values below it, or else would not be the balance point.
Proof
Expand the sum by distributing: . The second sum adds the constant to itself times, so .
By definition, , so multiplying both sides by gives . This is exactly the same quantity as the first sum in the expansion above.
Substituting back, , which proves the claim.
If a new data set is formed by for constants , then and the new standard deviation is .
Why is it true?
Shifting every data point by the same amount shifts the mean by that same amount but does not change how spread out the points are relative to each other; scaling every point by scales both the mean and the spread by (using since spread cannot be negative).
Proof
For the mean: , using that and .
For the spread, first note that : the constant shift cancels out completely, leaving only the scaled deviation.
Squaring and averaging, . Taking the square root of both sides gives , since a square root is always non-negative.
SchoolReal-World Applications and Worked Examples
Descriptive statistics are everywhere: a sports analyst compares a batting average to a league average, a factory tracks the standard deviation of part sizes to catch quality problems, a weather service reports the median rainfall because a single flood year would distort the mean, and a company reports quarterly revenue with both a typical value and a measure of volatility.
Example: An outlier throws off the mean
A tutor records 6 quiz scores: . On the last quiz, a data entry mistake replaces the with . Compare the mean and median before and after the mistake.
Solution
Before the mistake: the mean is , and sorting the data gives a median of (average of the two middle values). Mean and median agree closely.
After the mistake, the data becomes . The mean is now — pulled far upward by the single bad value.
The sorted data is still , so the median is still , completely unaffected. This is exactly why the median is preferred when a data set might contain data entry errors or extreme outliers.
Example: Comparing consistency with standard deviation
Two players score points in 5 games. Player A: . Player B: . Both have the same mean, but which player is more consistent? Compute the standard deviation of each to compare.
Solution
Both data sets have mean : for A, ; for B, .
For player A, the deviations from are , with squares summing to , so the variance is and the standard deviation is .
For player B, the deviations are , with squares summing to , so the variance is and . Since , player A scores far more consistently even though both average points per game.
Compute the mean of the data set .
Which measure of central tendency is least affected by a single extreme outlier in the data?
For the data set , what is the standard deviation ?
A boxplot of delivery times (minutes) shows and . What is the interquartile range?
References
- David Freedman, Robert Pisani, Roger Purves (2007). Statistics
- David S. Moore, William I. Notz (2020). The Basic Practice of Statistics