MathLabs

Probability and statistics

Regression

Modeling the relationship between variables to predict one from the others.

IntuitionPredicting one variable from another

Suppose you plot the number of hours a group of students studied against the scores they got on an exam. The points do not fall on a perfect line, but they clearly drift upward: more hours tend to go with higher scores. Regression is the tool that turns this drifting cloud of points into a single best-fitting line (or curve), so you can summarize the trend and predict a score for a new number of hours.

A scatter of points drifting upward with a straight best-fit line drawn through them, showing the slope and intercept of a simple linear regression.
Ordinary least-squares linear regression: the amber line minimizes the sum of squared vertical residuals (dashed segments) to the observed data points.

UndergraduateThe simple linear regression model

Definition: Simple linear regression

Given paired data (x1,y1),…,(xn,yn)(x_1,y_1),\dots,(x_n,y_n), the simple linear regression model assumes that yy depends on xx through a straight line plus random noise: y=β0+β1x+ϵy=\beta_0+\beta_1x+\epsilon, where β0\beta_0 is the intercept, β1\beta_1 is the slope, and ϵ\epsilon is a random error term with mean 0 that absorbs everything the line does not capture.

y=β0+β1x+ϵy=\beta_0+\beta_1x+\epsilon

The intercept β0\beta_0 and slope β1\beta_1 are unknown population parameters: fixed numbers we never observe directly. What we do have is a sample of nn points, from which we compute estimates β^0\hat\beta_0 and β^1\hat\beta_1. The fitted line y^=β^0+β^1x\hat y=\hat\beta_0+\hat\beta_1x is our best guess at the true line, and for each data point the residual ei=yi−y^ie_i=y_i-\hat y_i measures how far the fitted line misses that particular observation.

β^1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2,β^0=yˉ−β^1xˉ\hat\beta_1=\dfrac{\sum_{i=1}^n (x_i-\bar x)(y_i-\bar y)}{\sum_{i=1}^n (x_i-\bar x)^2},\qquad \hat\beta_0=\bar y-\hat\beta_1\bar x

Among all lines y=b0+b1xy=b_0+b_1x, the choice that minimizes the sum of squared residuals ∑i=1n(yi−b0−b1xi)2\sum_{i=1}^n (y_i-b_0-b_1x_i)^2 is b1=β^1=SxySxxb_1=\hat\beta_1=\dfrac{S_{xy}}{S_{xx}} and b0=β^0=yˉ−β^1xˉb_0=\hat\beta_0=\bar y-\hat\beta_1\bar x, where Sxy=∑i=1n(xi−xˉ)(yi−yˉ)S_{xy}=\sum_{i=1}^n(x_i-\bar x)(y_i-\bar y) and Sxx=∑i=1n(xi−xˉ)2S_{xx}=\sum_{i=1}^n(x_i-\bar x)^2.

Why is it true?

The sum of squared residuals is a smooth (quadratic, convex) function of b0b_0 and b1b_1, so its minimum is found exactly where both partial derivatives vanish — the same idea as finding the bottom of a bowl by setting its slope to zero in every direction.

Proof

Write Q(b0,b1)=∑i=1n(yi−b0−b1xi)2Q(b_0,b_1)=\sum_{i=1}^n (y_i-b_0-b_1x_i)^2. Taking the partial derivative with respect to b0b_0 and setting it to zero: ∂Q∂b0=−2∑i=1n(yi−b0−b1xi)=0\dfrac{\partial Q}{\partial b_0}=-2\sum_{i=1}^n(y_i-b_0-b_1x_i)=0, which simplifies to the first normal equation ∑i=1nyi=nb0+b1∑i=1nxi\sum_{i=1}^n y_i = nb_0+b_1\sum_{i=1}^n x_i, i.e. b0=yˉ−b1xˉb_0=\bar y-b_1\bar x.

Taking the partial derivative with respect to b1b_1 and setting it to zero: ∂Q∂b1=−2∑i=1nxi(yi−b0−b1xi)=0\dfrac{\partial Q}{\partial b_1}=-2\sum_{i=1}^n x_i(y_i-b_0-b_1x_i)=0, the second normal equation ∑i=1nxiyi=b0∑i=1nxi+b1∑i=1nxi2\sum_{i=1}^n x_iy_i = b_0\sum_{i=1}^n x_i+b_1\sum_{i=1}^n x_i^2.

Substitute b0=yˉ−b1xˉb_0=\bar y-b_1\bar x from the first equation into the second: ∑i=1nxiyi=(yˉ−b1xˉ)nxˉ+b1∑i=1nxi2=nxˉyˉ+b1(∑i=1nxi2−nxˉ2)\sum_{i=1}^n x_iy_i = (\bar y-b_1\bar x)n\bar x+b_1\sum_{i=1}^n x_i^2 = n\bar x\bar y+b_1\left(\sum_{i=1}^n x_i^2-n\bar x^2\right). Rearranging isolates b1b_1: b1(∑i=1nxi2−nxˉ2)=∑i=1nxiyi−nxˉyˉb_1\left(\sum_{i=1}^n x_i^2-n\bar x^2\right)=\sum_{i=1}^n x_iy_i-n\bar x\bar y.

A direct algebraic expansion shows ∑i=1nxi2−nxˉ2=∑i=1n(xi−xˉ)2=Sxx\sum_{i=1}^n x_i^2-n\bar x^2=\sum_{i=1}^n(x_i-\bar x)^2=S_{xx} and ∑i=1nxiyi−nxˉyˉ=∑i=1n(xi−xˉ)(yi−yˉ)=Sxy\sum_{i=1}^n x_iy_i-n\bar x\bar y=\sum_{i=1}^n(x_i-\bar x)(y_i-\bar y)=S_{xy}, so b1=Sxy/Sxxb_1=S_{xy}/S_{xx}. Substituting back gives b0=yˉ−b1xˉb_0=\bar y-b_1\bar x. Because QQ is a sum of squares in a quadratic form that grows without bound as ∣b0∣,∣b1∣→∞|b_0|,|b_1|\to\infty, this unique stationary point must be the global minimum.

UndergraduateGoodness of fit: the coefficient of determination

R2=1−SSESSTR^2=1-\dfrac{SSE}{SST}

Definition: Sums of squares and R2R^2

Decompose the total variation in yy into three sums of squares: SST=∑i=1n(yi−yˉ)2SST=\sum_{i=1}^n(y_i-\bar y)^2 (total), SSR=∑i=1n(y^i−yˉ)2SSR=\sum_{i=1}^n(\hat y_i-\bar y)^2 (explained by the regression), and SSE=∑i=1n(yi−y^i)2SSE=\sum_{i=1}^n(y_i-\hat y_i)^2 (left over as residuals). The coefficient of determination R2=SSR/SST=1−SSE/SSTR^2=SSR/SST=1-SSE/SST is the fraction of the variation in yy that the line explains.

In simple linear regression, SST=SSR+SSESST=SSR+SSE, so 0≤R2≤10\le R^2\le 1; moreover R2R^2 equals the square of the sample correlation coefficient rr between xx and yy: R2=r2R^2=r^2.

Why is it true?

The residuals from a least-squares fit are always uncorrelated with the fitted values, because the normal equations that define β^0,β^1\hat\beta_0,\hat\beta_1 are precisely the conditions that force this. That orthogonality is exactly what makes the total variation split cleanly into an "explained" piece and a "leftover" piece with no cross-term.

Proof

The two normal equations from the least-squares theorem above say ∑i=1nei=0\sum_{i=1}^n e_i=0 and ∑i=1nxiei=0\sum_{i=1}^n x_ie_i=0, where ei=yi−y^ie_i=y_i-\hat y_i. Since y^i=β^0+β^1xi\hat y_i=\hat\beta_0+\hat\beta_1x_i is a linear combination of 11 and xix_i, both normal equations combine to give ∑i=1neiy^i=β^0∑i=1nei+β^1∑i=1nxiei=0\sum_{i=1}^n e_i\hat y_i=\hat\beta_0\sum_{i=1}^n e_i+\hat\beta_1\sum_{i=1}^n x_ie_i=0. Combined with ∑ei=0\sum e_i=0, this gives ∑i=1nei(y^i−yˉ)=∑eiy^i−yˉ∑ei=0\sum_{i=1}^n e_i(\hat y_i-\bar y)=\sum e_i\hat y_i-\bar y\sum e_i=0.

Now expand SST=∑i=1n(yi−yˉ)2=∑i=1n((yi−y^i)+(y^i−yˉ))2=∑ei2+2∑ei(y^i−yˉ)+∑(y^i−yˉ)2SST=\sum_{i=1}^n(y_i-\bar y)^2=\sum_{i=1}^n\big((y_i-\hat y_i)+(\hat y_i-\bar y)\big)^2=\sum e_i^2+2\sum e_i(\hat y_i-\bar y)+\sum(\hat y_i-\bar y)^2. The cross term vanishes by the orthogonality just shown, leaving SST=SSE+SSRSST=SSE+SSR.

Since SSE=∑ei2≥0SSE=\sum e_i^2\ge0 and SSR=∑(y^i−yˉ)2≥0SSR=\sum(\hat y_i-\bar y)^2\ge0, dividing SST=SSR+SSESST=SSR+SSE by SST>0SST>0 gives R2=SSR/SST∈[0,1]R^2=SSR/SST\in[0,1].

Finally, y^i−yˉ=β^1(xi−xˉ)\hat y_i-\bar y=\hat\beta_1(x_i-\bar x) (since y^i=β^0+β^1xi\hat y_i=\hat\beta_0+\hat\beta_1x_i and yˉ=β^0+β^1xˉ\bar y=\hat\beta_0+\hat\beta_1\bar x), so SSR=β^12SxxSSR=\hat\beta_1^2S_{xx}. Substituting β^1=Sxy/Sxx\hat\beta_1=S_{xy}/S_{xx} gives SSR=Sxy2/SxxSSR=S_{xy}^2/S_{xx}, and since SST=Syy=∑(yi−yˉ)2SST=S_{yy}=\sum(y_i-\bar y)^2, R2=Sxy2SxxSyy=(SxySxxSyy)2=r2R^2=\dfrac{S_{xy}^2}{S_{xx}S_{yy}}=\left(\dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}\right)^2=r^2, the square of the sample correlation coefficient.

Example: Fitting a line to five data points

A small dataset has x=(1,2,3,4,5)x=(1,2,3,4,5) and y=(2,4,5,4,5)y=(2,4,5,4,5). Find the least-squares line and compute R2R^2.

Solution

The means are xˉ=3\bar x=3 and yˉ=4\bar y=4. The deviations (xi−xˉ,yi−yˉ)(x_i-\bar x, y_i-\bar y) are (−2,−2),(−1,0),(0,1),(1,0),(2,1)(-2,-2),(-1,0),(0,1),(1,0),(2,1), so Sxy=4+0+0+0+2=6S_{xy}=4+0+0+0+2=6 and Sxx=4+1+0+1+4=10S_{xx}=4+1+0+1+4=10.

This gives β^1=Sxy/Sxx=6/10=0.6\hat\beta_1=S_{xy}/S_{xx}=6/10=0.6 and β^0=yˉ−β^1xˉ=4−0.6×3=2.2\hat\beta_0=\bar y-\hat\beta_1\bar x=4-0.6\times3=2.2, so the fitted line is y^=0.6x+2.2\hat y=0.6x+2.2.

For R2R^2: SST=∑(yi−yˉ)2=4+0+1+0+1=6SST=\sum(y_i-\bar y)^2=4+0+1+0+1=6, and SSR=β^12Sxx=0.36×10=3.6SSR=\hat\beta_1^2S_{xx}=0.36\times10=3.6, so R2=SSR/SST=3.6/6=0.6R^2=SSR/SST=3.6/6=0.6 — the line explains 60% of the variation in yy.

UndergraduateChecking the model: residual analysis

Fitting a line is easy; trusting it requires checking whether the model's assumptions actually hold. Plotting the residuals ei=yi−y^ie_i=y_i-\hat y_i against xix_i (or against y^i\hat y_i) is the standard diagnostic: if the linear model is appropriate, this residual plot should look like a random, structureless scatter around zero, with roughly constant spread.

Reading a residual plot
Pattern in the residual plotWhat it suggests
Random scatter around zero, constant spreadThe linear model and constant-variance assumption look reasonable
Funnel/fan shape (spread grows with xx)Heteroscedasticity: the constant-variance assumption is violated
Curved (U-shaped or arc) patternThe true relationship is nonlinear; a straight line is the wrong shape

UndergraduateReal-world applications and worked examples

Regression is one of the most widely used tools in applied science: economists use it to estimate how strongly a stock moves with the overall market, biologists use it to relate an animal's body size to its metabolic rate, engineers use it to calibrate instruments against a known reference, and epidemiologists use it to relate a drug's dose to a measured response. In every case, the same two questions arise: how good is the fit (R2R^2), and does the residual plot reveal a hidden problem?

Example: Forecasting ice-cream sales from temperature

A shop records daily temperature xx (°C) and ice-cream sales yy (units) on five days: x=(20,25,30,35,40)x=(20,25,30,35,40), y=(30,45,55,65,80)y=(30,45,55,65,80). Fit a regression line and use it to forecast sales at x=32x=32°C.

Solution

The means are xˉ=30\bar x=30, yˉ=55\bar y=55. Deviations give Sxy=(−10)(−25)+(−5)(−10)+0+5×10+10×25=250+50+0+50+250=600S_{xy}=(-10)(-25)+(-5)(-10)+0+5\times10+10\times25=250+50+0+50+250=600 and Sxx=100+25+0+25+100=250S_{xx}=100+25+0+25+100=250, so β^1=600/250=2.4\hat\beta_1=600/250=2.4 and β^0=55−2.4×30=−17\hat\beta_0=55-2.4\times30=-17: the fitted line is y^=2.4x−17\hat y=2.4x-17.

At x=32x=32: y^=2.4×32−17=76.8−17=59.8\hat y=2.4\times32-17=76.8-17=59.8, so the shop should expect roughly 60 units sold.

Checking fit: SST=625+100+0+100+625=1450SST=625+100+0+100+625=1450 and SSR=β^12Sxx=5.76×250=1440SSR=\hat\beta_1^2S_{xx}=5.76\times250=1440, so R2=1440/1450≈0.993R^2=1440/1450\approx0.993 — temperature explains about 99.3% of the variation in sales here, an unusually tight real-world fit.

Example: A high R2R^2 that still hides a problem

An engineer calibrates a new pressure sensor by fitting reading=β0+β1⋅true pressure\text{reading}=\beta_0+\beta_1\cdot\text{true pressure} against a reference instrument across many pressures, and gets R2=0.95R^2=0.95 — seemingly excellent. But plotting the residuals against true pressure reveals a clear U-shaped curve: residuals are positive at low and high pressures and negative in the middle. What does this mean, and what should the engineer do?

Solution

A high R2R^2 only says the line tracks the overall trend well; it does not certify that the relationship is linear. The systematic U-shaped residual pattern is exactly the "curved pattern" signature from the diagnostics table above: it means the true reading-versus-pressure relationship has curvature that the straight line is not capturing — the sensor likely has a genuine quadratic (or otherwise nonlinear) response.

Because the pattern is systematic rather than random, more data at the same pressures will not fix it — the model itself is misspecified, not just noisy. The residuals being simultaneously informative (clearly patterned) and numerically small (giving a high R2R^2) shows why relying on R2R^2 alone is a mistake for a calibration application where the whole point is precision.

The fix is to change the model, not just refit the line: add a quadratic term (fit reading=β0+β1⋅pressure+β2⋅pressure2\text{reading}=\beta_0+\beta_1\cdot\text{pressure}+\beta_2\cdot\text{pressure}^2) or apply a known physical transformation of the sensor's response, then re-check that the new residual plot looks like unstructured noise before trusting the calibration.

Given Sxy=50S_{xy}=50 and Sxx=25S_{xx}=25 for a dataset, what is β^1\hat\beta_1?

If R2=0.81R^2=0.81 for a regression, what fraction of the variation in yy is left unexplained by the linear model?

A residual plot shows a clear funnel shape, widening as xx increases. What does this indicate?

A real-estate model fits price=β^0+β^1⋅sqft\text{price}=\hat\beta_0+\hat\beta_1\cdot\text{sqft} (price in thousands of dollars) with β^0=20\hat\beta_0=20 and β^1=0.15\hat\beta_1=0.15. What price does it predict for a 1500-sqft house?