Linear Regression
Once a scatter plot shows a linear pattern, you can draw the line of best fit through it. That line is a model: it summarizes the relationship in an equation and lets you make predictions. Linear regression is how technology finds the best possible line, and residuals tell you how well it fits.
Key ideas
Section titled “Key ideas”The line of best fit
Section titled “The line of best fit”The line of best fit (or regression line) is written
where is the slope and is the -intercept. Technology finds it using the least squares method: of all possible lines, it picks the one that makes the sum of the squared vertical distances from the points to the line as small as possible.
- Spreadsheet:
=SLOPE(B2:B9, A2:A9)and=INTERCEPT(B2:B9, A2:A9)(note: -values first), or add a linear trendline to a chart. - Graphing calculator: STAT, CALC, LinReg(ax+b).
- Desmos:
y1 ~ a x1 + b.
The line always passes through the point , the means of the two variables.
Interpreting slope and intercept
Section titled “Interpreting slope and intercept”- Slope : for each increase of unit in , the model predicts changes by units.
- Intercept : the predicted when . This only makes sense if is reasonable and close to the data.
Always use the units and the context: “each extra degree predicts about more cones sold”, not just “the slope is ”.
Interpolation and extrapolation
Section titled “Interpolation and extrapolation”- Interpolation: predicting for an -value inside the range of the data. Usually reliable when the correlation is strong.
- Extrapolation: predicting outside the range. Risky, because the pattern may not continue.
Residuals
Section titled “Residuals”A residual is how far a point is above or below the line:
A positive residual means the point is above the line; negative means below.
A residual plot graphs the residuals against . If the line is a good model, the residuals look randomly scattered around . A clear pattern, like a U shape, means a straight line isn’t the right model.
Outliers
Section titled “Outliers”An outlier can pull the line toward itself, changing the slope and intercept, and usually lowering . Outliers far from the others in the -direction have the most pull.
Answers on this page are rounded; technology may differ slightly in the last digit.
On the SAT
Section titled “On the SAT”For raw data, use a Desmos table and type y_1 ~ mx_1 + b to get the slope, intercept, and . Most SAT regression questions give you the line of best fit already and ask what its slope or intercept means in context (“for each additional hour, the predicted score increases by…”), or ask for a predicted value; that’s interpretation, which Desmos can’t do. When a question asks how far an actual data point is from the line, find actual minus predicted (the residual), and remember that the line gives a predicted value, not a guaranteed one. See using Desmos on the SAT.
Worked examples
Section titled “Worked examples”Example 1: Finding and interpreting the line
Section titled “Example 1: Finding and interpreting the line”A student runs an ice-cream stand and records the afternoon temperature and the number of cones sold on eight days.
| Temperature (°C) | ||||||||
|---|---|---|---|---|---|---|---|---|
| Cones sold |
Find the line of best fit and . Interpret the slope and the intercept.
Solution. With technology, using for temperature and for cones sold:
The correlation is strong and positive, so a line is a good model.
- Slope: each °C rise in temperature predicts about more cones sold.
- Intercept: at °C the model predicts cones. That’s impossible! The data only go from °C to °C, so is far outside the data and the intercept has no real meaning here.
Example 2: Interpolation and extrapolation
Section titled “Example 2: Interpolation and extrapolation”Use the line from Example 1 to predict sales at (a) °C (b) °C. How reliable is each prediction?
Solution. Use the rounded equation from Example 1.
(a)
About cones. Since °C is inside the data range ( °C to °C), this is interpolation, and with it’s fairly reliable.
(b)
About cones. But °C is well outside the data, so this is extrapolation. On a very hot day, people might stay indoors, or the stand might run out of ice cream. Treat this prediction with caution.
Example 3: Residuals
Section titled “Example 3: Residuals”Find the residuals for the days at °C and °C.
Solution.
At °C: predicted . Observed: .
At °C: predicted . Observed: .
The °C day sold about cones fewer than predicted (below the line); the °C day sold about more (above the line).
Example 4: The effect of an outlier
Section titled “Example 4: The effect of an outlier”On one °C day, a summer camp group stopped by and the stand sold cones. Add this point to the data and find the new line and .
Solution. With all nine days:
The single outlier lifted the left side of the line, so the slope fell from about to about , and dropped from strong () to moderate (). Because the camp visit was a one-time event, it makes sense to leave it out when modelling ordinary sales, and to say that you did.
Common mistakes
Section titled “Common mistakes”Mixing up the variables. If you swap and , technology gives a different line. Put the independent variable in the list.
Subtracting the wrong way for a residual. It’s observed minus predicted. A point above the line has a positive residual.
Trusting extrapolations. The model is only supported by the data you have. Predicting far outside that range, or interpreting an intercept far from the data, can give nonsense (like cones).
Interpreting the slope without context or units. Say what changes, by how much, and for each unit of what.
Rounding the equation too early. Rounding and to one decimal before predicting can noticeably change your answers. Keep at least three significant digits.
Deleting outliers without a reason. Remove a point only if you have a reason (an error, or an unusual event), and always report that you did.
Practice
Section titled “Practice”1. (Warm-up) A line of best fit is . Predict when .
Solution
2. (Warm-up) A point has observed , and the line predicts . Find the residual. Is the point above or below the line?
Solution
Residual . It’s negative, so the point is below the line.
3. (Warm-up) The ice-cream data in Example 1 range from °C to °C. Is each prediction interpolation or extrapolation? (a) °C (b) °C (c) °C
Solution
(a) Interpolation. (b) Extrapolation. (c) Extrapolation.
4. (Core) A driver records distance driven and fuel used on six trips.
| Distance (km) | ||||||
|---|---|---|---|---|---|---|
| Fuel (L) |
- (a) Find the line of best fit and .
- (b) Interpret the slope.
- (c) Predict the fuel used on a km trip.
Solution
(a) , with , a very strong positive correlation.
(b) Each extra kilometre uses about L of fuel, or about L per km.
(c) L. This is interpolation, and with so close to , it’s very reliable.
5. (Core) For one house, the monthly heating bill (in dollars) and the average outdoor temperature (in °C) for the winter months are modelled by .
- (a) Interpret the slope and the intercept.
- (b) Predict the bill for a month averaging °C.
Solution
(a) Slope: each °C warmer average temperature predicts a bill about $4.20 lower. Intercept: a month averaging °C predicts a bill of about $135. Since °C is within the range of typical Ontario winter temperatures, this intercept is meaningful.
(b) , so about $168.60.
6. (Core) Using your line from Question 4, find the residuals for the km and km trips. Which of the two trips used more fuel than predicted?
Solution
km: predicted L, so the residual is L.
km: predicted L, so the residual is L. (Using unrounded coefficients gives about L.)
The km trip had a positive residual, so it used more fuel than predicted.
7. (Core) The hours-studied data from scatter plots and correlation are:
| Hours | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mark (%) |
The last student was sick on test day. Find the line of best fit with and without that student, and compare.
Solution
With: , .
Without: , .
The outlier flattens the line a lot (the slope drops by more than half) and weakens the correlation from strong to moderate. Since there’s a clear reason the point is unusual, it’s reasonable to leave it out and say so.
8. (Challenge) For the points , , , , , :
- (a) Find the line of best fit and .
- (b) Find all six residuals. What pattern do you see, and what does it tell you?
Solution
(a) , with .
(b) The predicted values are , , , , , , so the residuals are about
They go positive, negative, then positive again: a U shape. Even though is very high, the data curve upward, so a straight line is not the best model. (In fact fits exactly.)
9. (Challenge) Show that the line of best fit from Example 1 passes (very nearly) through the point .
Solution
At : . The tiny difference is from rounding the slope and intercept; with unrounded values the line passes through exactly.