Modeling the relationship between variables to predict one from the others.
IntuitionPredicting one variable from another
Suppose you plot the number of hours a group of students studied against the scores they got on an exam. The points do not fall on a perfect line, but they clearly drift upward: more hours tend to go with higher scores. Regression is the tool that turns this drifting cloud of points into a single best-fitting line (or curve), so you can summarize the trend and predict a score for a new number of hours.
A scatter of points drifting upward with a straight best-fit line drawn through them, showing the slope and intercept of a simple linear regression.
Ordinary least-squares linear regression: the amber line minimizes the sum of squared vertical residuals (dashed segments) to the observed data points.
UndergraduateThe simple linear regression model
Definition: Simple linear regression
Given paired data (x1,y1),…,(xn,yn), the simple linear regression model assumes that y depends on x through a straight line plus random noise: y=β0+β1x+ϵ, where β0 is the intercept, β1 is the slope, and ϵ is a random error term with mean 0 that absorbs everything the line does not capture.
y=β0+β1x+ϵ
The intercept β0 and slope β1 are unknown population parameters: fixed numbers we never observe directly. What we do have is a sample of n points, from which we compute estimatesβ^0 and β^1. The fitted line y^=β^0+β^1x is our best guess at the true line, and for each data point the residualei=yi−y^i measures how far the fitted line misses that particular observation.
Among all lines y=b0+b1x, the choice that minimizes the sum of squared residuals ∑i=1n(yi−b0−b1xi)2 is b1=β^1=SxxSxy and b0=β^0=yˉ−β^1xˉ, where Sxy=∑i=1n(xi−xˉ)(yi−yˉ) and Sxx=∑i=1n(xi−xˉ)2.
Why is it true?
The sum of squared residuals is a smooth (quadratic, convex) function of b0 and b1, so its minimum is found exactly where both partial derivatives vanish — the same idea as finding the bottom of a bowl by setting its slope to zero in every direction.
Proof
Write Q(b0,b1)=∑i=1n(yi−b0−b1xi)2. Taking the partial derivative with respect to b0 and setting it to zero: ∂b0∂Q=−2∑i=1n(yi−b0−b1xi)=0, which simplifies to the first normal equation ∑i=1nyi=nb0+b1∑i=1nxi, i.e. b0=yˉ−b1xˉ.
Taking the partial derivative with respect to b1 and setting it to zero: ∂b1∂Q=−2∑i=1nxi(yi−b0−b1xi)=0, the second normal equation ∑i=1nxiyi=b0∑i=1nxi+b1∑i=1nxi2.
Substitute b0=yˉ−b1xˉ from the first equation into the second: ∑i=1nxiyi=(yˉ−b1xˉ)nxˉ+b1∑i=1nxi2=nxˉyˉ+b1(∑i=1nxi2−nxˉ2). Rearranging isolates b1: b1(∑i=1nxi2−nxˉ2)=∑i=1nxiyi−nxˉyˉ.
A direct algebraic expansion shows ∑i=1nxi2−nxˉ2=∑i=1n(xi−xˉ)2=Sxx and ∑i=1nxiyi−nxˉyˉ=∑i=1n(xi−xˉ)(yi−yˉ)=Sxy, so b1=Sxy/Sxx. Substituting back gives b0=yˉ−b1xˉ. Because Q is a sum of squares in a quadratic form that grows without bound as ∣b0∣,∣b1∣→∞, this unique stationary point must be the global minimum.
UndergraduateGoodness of fit: the coefficient of determination
R2=1−SSTSSE
Definition: Sums of squares and R2
Decompose the total variation in y into three sums of squares: SST=∑i=1n(yi−yˉ)2 (total), SSR=∑i=1n(y^i−yˉ)2 (explained by the regression), and SSE=∑i=1n(yi−y^i)2 (left over as residuals). The coefficient of determinationR2=SSR/SST=1−SSE/SST is the fraction of the variation in y that the line explains.
In simple linear regression, SST=SSR+SSE, so 0≤R2≤1; moreover R2 equals the square of the sample correlation coefficient r between x and y: R2=r2.
Why is it true?
The residuals from a least-squares fit are always uncorrelated with the fitted values, because the normal equations that define β^0,β^1 are precisely the conditions that force this. That orthogonality is exactly what makes the total variation split cleanly into an "explained" piece and a "leftover" piece with no cross-term.
Proof
The two normal equations from the least-squares theorem above say ∑i=1nei=0 and ∑i=1nxiei=0, where ei=yi−y^i. Since y^i=β^0+β^1xi is a linear combination of 1 and xi, both normal equations combine to give ∑i=1neiy^i=β^0∑i=1nei+β^1∑i=1nxiei=0. Combined with ∑ei=0, this gives ∑i=1nei(y^i−yˉ)=∑eiy^i−yˉ∑ei=0.
Now expand SST=∑i=1n(yi−yˉ)2=∑i=1n((yi−y^i)+(y^i−yˉ))2=∑ei2+2∑ei(y^i−yˉ)+∑(y^i−yˉ)2. The cross term vanishes by the orthogonality just shown, leaving SST=SSE+SSR.
Since SSE=∑ei2≥0 and SSR=∑(y^i−yˉ)2≥0, dividing SST=SSR+SSE by SST>0 gives R2=SSR/SST∈[0,1].
Finally, y^i−yˉ=β^1(xi−xˉ) (since y^i=β^0+β^1xi and yˉ=β^0+β^1xˉ), so SSR=β^12Sxx. Substituting β^1=Sxy/Sxx gives SSR=Sxy2/Sxx, and since SST=Syy=∑(yi−yˉ)2, R2=SxxSyySxy2=(SxxSyySxy)2=r2, the square of the sample correlation coefficient.
Example: Fitting a line to five data points
A small dataset has x=(1,2,3,4,5) and y=(2,4,5,4,5). Find the least-squares line and compute R2.
Solution
The means are xˉ=3 and yˉ=4. The deviations (xi−xˉ,yi−yˉ) are (−2,−2),(−1,0),(0,1),(1,0),(2,1), so Sxy=4+0+0+0+2=6 and Sxx=4+1+0+1+4=10.
This gives β^1=Sxy/Sxx=6/10=0.6 and β^0=yˉ−β^1xˉ=4−0.6×3=2.2, so the fitted line is y^=0.6x+2.2.
For R2: SST=∑(yi−yˉ)2=4+0+1+0+1=6, and SSR=β^12Sxx=0.36×10=3.6, so R2=SSR/SST=3.6/6=0.6 — the line explains 60% of the variation in y.
UndergraduateChecking the model: residual analysis
Fitting a line is easy; trusting it requires checking whether the model's assumptions actually hold. Plotting the residuals ei=yi−y^i against xi (or against y^i) is the standard diagnostic: if the linear model is appropriate, this residual plot should look like a random, structureless scatter around zero, with roughly constant spread.
Reading a residual plot
Pattern in the residual plot
What it suggests
Random scatter around zero, constant spread
The linear model and constant-variance assumption look reasonable
Funnel/fan shape (spread grows with x)
Heteroscedasticity: the constant-variance assumption is violated
Curved (U-shaped or arc) pattern
The true relationship is nonlinear; a straight line is the wrong shape
UndergraduateReal-world applications and worked examples
Regression is one of the most widely used tools in applied science: economists use it to estimate how strongly a stock moves with the overall market, biologists use it to relate an animal's body size to its metabolic rate, engineers use it to calibrate instruments against a known reference, and epidemiologists use it to relate a drug's dose to a measured response. In every case, the same two questions arise: how good is the fit (R2), and does the residual plot reveal a hidden problem?
Example: Forecasting ice-cream sales from temperature
A shop records daily temperature x (°C) and ice-cream sales y (units) on five days: x=(20,25,30,35,40), y=(30,45,55,65,80). Fit a regression line and use it to forecast sales at x=32°C.
Solution
The means are xˉ=30, yˉ=55. Deviations give Sxy=(−10)(−25)+(−5)(−10)+0+5×10+10×25=250+50+0+50+250=600 and Sxx=100+25+0+25+100=250, so β^1=600/250=2.4 and β^0=55−2.4×30=−17: the fitted line is y^=2.4x−17.
At x=32: y^=2.4×32−17=76.8−17=59.8, so the shop should expect roughly 60 units sold.
Checking fit: SST=625+100+0+100+625=1450 and SSR=β^12Sxx=5.76×250=1440, so R2=1440/1450≈0.993 — temperature explains about 99.3% of the variation in sales here, an unusually tight real-world fit.
Example: A high R2 that still hides a problem
An engineer calibrates a new pressure sensor by fitting reading=β0+β1⋅true pressure against a reference instrument across many pressures, and gets R2=0.95 — seemingly excellent. But plotting the residuals against true pressure reveals a clear U-shaped curve: residuals are positive at low and high pressures and negative in the middle. What does this mean, and what should the engineer do?
Solution
A high R2 only says the line tracks the overall trend well; it does not certify that the relationship is linear. The systematic U-shaped residual pattern is exactly the "curved pattern" signature from the diagnostics table above: it means the true reading-versus-pressure relationship has curvature that the straight line is not capturing — the sensor likely has a genuine quadratic (or otherwise nonlinear) response.
Because the pattern is systematic rather than random, more data at the same pressures will not fix it — the model itself is misspecified, not just noisy. The residuals being simultaneously informative (clearly patterned) and numerically small (giving a high R2) shows why relying on R2 alone is a mistake for a calibration application where the whole point is precision.
The fix is to change the model, not just refit the line: add a quadratic term (fit reading=β0+β1⋅pressure+β2⋅pressure2) or apply a known physical transformation of the sensor's response, then re-check that the new residual plot looks like unstructured noise before trusting the calibration.
Given Sxy=50 and Sxx=25 for a dataset, what is β^1?
If R2=0.81 for a regression, what fraction of the variation in y is left unexplained by the linear model?
A residual plot shows a clear funnel shape, widening as x increases. What does this indicate?
A real-estate model fits price=β^0+β^1⋅sqft (price in thousands of dollars) with β^0=20 and β^1=0.15. What price does it predict for a 1500-sqft house?