MathLabs

Probability and statistics

Descriptive statistics

Summarizing data with measures like mean, median and standard deviation, and displaying it in charts.

IntuitionIntuition: Telling the Story of Data with a Few Numbers

Imagine a teacher with 30 quiz scores, or a shop owner with a month of daily sales. Nobody wants to stare at 30 separate numbers to understand what happened. Descriptive statistics compresses a whole data set into a handful of numbers — a typical value, how spread out the values are — so we can compare, summarize and spot anything unusual at a glance.

An asymmetric curve with a rise and a dip, illustrating a skewed data shape rather than a symmetric one.
A right-skewed count distribution (Poisson with rate λ=np\lambda = np): increase nn or pp to watch the center of mass shift right and the shape become more symmetric.

SchoolMeasuring the Center: Mean, Median, Mode

Definition: Mean

The mean (or average) of nn data values x1,x2,…,xnx_1, x_2, \ldots, x_n adds them all up and divides by how many there are. It uses every value in the data set, which makes it sensitive to unusually large or small values.

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i

Here xˉ\bar x is the mean, nn is the number of data values, and xix_i is the ii-th value being summed. The bar over xx is standard notation for "the mean of xx".

Definition: Median and mode

The median is the middle value when the data is sorted from smallest to largest (or the average of the two middle values, if nn is even). The mode is the value that occurs most often. Unlike the mean, the median ignores how far away the extreme values are — only their rank matters — which makes it resistant to outliers.

So which measure should you use? If the data is roughly symmetric with no extreme values, the mean is efficient and uses all the information. If the data has outliers or is heavily skewed — like household incomes or house prices — the median gives a more representative "typical" value.

Mean, median and mode compared
MeasureDefinitionSensitive to outliers?
MeanSum of values divided by countYes — every value counts
MedianMiddle value when sortedNo — only rank matters
ModeMost frequent valueNo

SchoolMeasuring the Spread: Standard Deviation and Quartiles

Two data sets can have the same mean but look completely different — one tightly clustered, one wildly spread out. The standard deviation ss captures this by measuring the typical distance of a data value from the mean:

s=1n∑i=1n(xi−xˉ)2s = \sqrt{\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar{x})^2}

Each deviation xi−xˉx_i-\bar x is squared (so negative and positive deviations don't cancel out), averaged, and then the square root brings the units back to the original scale of the data. A small ss means the data clusters tightly around xˉ\bar x; a large ss means it is spread out.

IQR=Q3−Q1IQR = Q_3 - Q_1

Quartiles split sorted data into four equal parts: Q1Q_1 (25% of the way through), the median (50%), and Q3Q_3 (75%). The interquartile range IQR=Q3−Q1IQR = Q_3-Q_1 measures the spread of the middle half of the data and, like the median, is resistant to outliers — this is exactly what a boxplot draws: a box from Q1Q_1 to Q3Q_3 with a line at the median.

For any data set x1,x2,…,xnx_1, x_2, \ldots, x_n with mean xˉ\bar x, ∑i=1n(xi−xˉ)=0\sum_{i=1}^{n}(x_i-\bar x) = 0.

Why is it true?

The mean is defined precisely as the balance point of the data — think of a see-saw with weights at each xix_i — so the values above the mean must exactly offset the values below it, or else xˉ\bar x would not be the balance point.

Proof

Expand the sum by distributing: ∑i=1n(xi−xˉ)=∑i=1nxi−∑i=1nxˉ\sum_{i=1}^{n}(x_i-\bar x) = \sum_{i=1}^{n}x_i - \sum_{i=1}^{n}\bar x. The second sum adds the constant xˉ\bar x to itself nn times, so ∑i=1nxˉ=nxˉ\sum_{i=1}^{n}\bar x = n\bar x.

By definition, xˉ=1n∑i=1nxi\bar x = \frac{1}{n}\sum_{i=1}^{n}x_i, so multiplying both sides by nn gives nxˉ=∑i=1nxin\bar x = \sum_{i=1}^{n}x_i. This is exactly the same quantity as the first sum in the expansion above.

Substituting back, ∑i=1n(xi−xˉ)=∑i=1nxi−nxˉ=∑i=1nxi−∑i=1nxi=0\sum_{i=1}^{n}(x_i-\bar x) = \sum_{i=1}^{n}x_i - n\bar x = \sum_{i=1}^{n}x_i - \sum_{i=1}^{n}x_i = 0, which proves the claim.

If a new data set is formed by yi=axi+by_i = ax_i+b for constants a,ba,b, then yˉ=axˉ+b\bar y = a\bar x+b and the new standard deviation is sy=∣a∣sxs_y = |a|s_x.

Why is it true?

Shifting every data point by the same amount bb shifts the mean by that same amount but does not change how spread out the points are relative to each other; scaling every point by aa scales both the mean and the spread by aa (using ∣a∣|a| since spread cannot be negative).

Proof

For the mean: yˉ=1n∑i=1nyi=1n∑i=1n(axi+b)=an∑i=1nxi+1n∑i=1nb=axˉ+b\bar y = \frac{1}{n}\sum_{i=1}^{n}y_i = \frac{1}{n}\sum_{i=1}^{n}(ax_i+b) = \frac{a}{n}\sum_{i=1}^{n}x_i + \frac{1}{n}\sum_{i=1}^{n}b = a\bar x + b, using that 1n∑xi=xˉ\frac{1}{n}\sum x_i = \bar x and 1n∑i=1nb=b\frac{1}{n}\sum_{i=1}^{n}b = b.

For the spread, first note that yi−yˉ=(axi+b)−(axˉ+b)=a(xi−xˉ)y_i-\bar y = (ax_i+b)-(a\bar x+b) = a(x_i-\bar x): the constant shift bb cancels out completely, leaving only the scaled deviation.

Squaring and averaging, sy2=1n∑i=1n(yi−yˉ)2=1n∑i=1na2(xi−xˉ)2=a2sx2s_y^2 = \frac{1}{n}\sum_{i=1}^{n}(y_i-\bar y)^2 = \frac{1}{n}\sum_{i=1}^{n}a^2(x_i-\bar x)^2 = a^2 s_x^2. Taking the square root of both sides gives sy=∣a∣sxs_y = |a|s_x, since a square root is always non-negative.

SchoolReal-World Applications and Worked Examples

Descriptive statistics are everywhere: a sports analyst compares a batting average to a league average, a factory tracks the standard deviation of part sizes to catch quality problems, a weather service reports the median rainfall because a single flood year would distort the mean, and a company reports quarterly revenue with both a typical value and a measure of volatility.

Example: An outlier throws off the mean

A tutor records 6 quiz scores: 5,6,7,7,8,95, 6, 7, 7, 8, 9. On the last quiz, a data entry mistake replaces the 99 with 5050. Compare the mean and median before and after the mistake.

Solution

Before the mistake: the mean is 5+6+7+7+8+96=426=7\frac{5+6+7+7+8+9}{6} = \frac{42}{6} = 7, and sorting the data 5,6,7,7,8,95,6,7,7,8,9 gives a median of 7+72=7\frac{7+7}{2}=7 (average of the two middle values). Mean and median agree closely.

After the mistake, the data becomes 5,6,7,7,8,505,6,7,7,8,50. The mean is now 5+6+7+7+8+506=836≈13.83\frac{5+6+7+7+8+50}{6} = \frac{83}{6} \approx 13.83 — pulled far upward by the single bad value.

The sorted data is still 5,6,7,7,8,505,6,7,7,8,50, so the median is still 7+72=7\frac{7+7}{2}=7, completely unaffected. This is exactly why the median is preferred when a data set might contain data entry errors or extreme outliers.

Example: Comparing consistency with standard deviation

Two players score points in 5 games. Player A: 9,10,11,12,139, 10, 11, 12, 13. Player B: 1,6,11,16,211, 6, 11, 16, 21. Both have the same mean, but which player is more consistent? Compute the standard deviation of each to compare.

Solution

Both data sets have mean 1111: for A, 9+10+11+12+135=555=11\frac{9+10+11+12+13}{5}=\frac{55}{5}=11; for B, 1+6+11+16+215=555=11\frac{1+6+11+16+21}{5}=\frac{55}{5}=11.

For player A, the deviations from 1111 are −2,−1,0,1,2-2,-1,0,1,2, with squares 4,1,0,1,44,1,0,1,4 summing to 1010, so the variance is 105=2\frac{10}{5}=2 and the standard deviation is sA=2≈1.41s_A=\sqrt{2}\approx 1.41.

For player B, the deviations are −10,−5,0,5,10-10,-5,0,5,10, with squares 100,25,0,25,100100,25,0,25,100 summing to 250250, so the variance is 2505=50\frac{250}{5}=50 and sB=50≈7.07s_B=\sqrt{50}\approx 7.07. Since sA<sBs_A < s_B, player A scores far more consistently even though both average 1111 points per game.

Compute the mean of the data set 2,4,4,6,92, 4, 4, 6, 9.

Which measure of central tendency is least affected by a single extreme outlier in the data?

For the data set 1,2,3,4,51, 2, 3, 4, 5, what is the standard deviation ss?

A boxplot of delivery times (minutes) shows Q1=20Q_1=20 and Q3=35Q_3=35. What is the interquartile range?

References

  1. David Freedman, Robert Pisani, Roger Purves (2007). Statistics
  2. David S. Moore, William I. Notz (2020). The Basic Practice of Statistics