Skip to content
Family Table Math

Linear Regression

Once a scatter plot shows a linear pattern, you can draw the line of best fit through it. That line is a model: it summarizes the relationship in an equation and lets you make predictions. Linear regression is how technology finds the best possible line, and residuals tell you how well it fits.

The line of best fit (or regression line) is written

y=ax+by = ax + b

where aa is the slope and bb is the yy-intercept. Technology finds it using the least squares method: of all possible lines, it picks the one that makes the sum of the squared vertical distances from the points to the line as small as possible.

  • Spreadsheet: =SLOPE(B2:B9, A2:A9) and =INTERCEPT(B2:B9, A2:A9) (note: yy-values first), or add a linear trendline to a chart.
  • Graphing calculator: STAT, CALC, LinReg(ax+b).
  • Desmos: y1 ~ a x1 + b.

The line always passes through the point (xˉ,yˉ)(\bar{x}, \bar{y}), the means of the two variables.

  • Slope aa: for each increase of 11 unit in xx, the model predicts yy changes by aa units.
  • Intercept bb: the predicted yy when x=0x = 0. This only makes sense if x=0x = 0 is reasonable and close to the data.

Always use the units and the context: “each extra degree predicts about 44 more cones sold”, not just “the slope is 44”.

  • Interpolation: predicting for an xx-value inside the range of the data. Usually reliable when the correlation is strong.
  • Extrapolation: predicting outside the range. Risky, because the pattern may not continue.

A residual is how far a point is above or below the line:

residual=observed y−predicted y\text{residual} = \text{observed } y - \text{predicted } y

A positive residual means the point is above the line; negative means below.

A residual plot graphs the residuals against xx. If the line is a good model, the residuals look randomly scattered around 00. A clear pattern, like a U shape, means a straight line isn’t the right model.

An outlier can pull the line toward itself, changing the slope and intercept, and usually lowering ∣r∣|r|. Outliers far from the others in the xx-direction have the most pull.

Answers on this page are rounded; technology may differ slightly in the last digit.

Example 1: Finding and interpreting the line

Section titled “Example 1: Finding and interpreting the line”

A student runs an ice-cream stand and records the afternoon temperature and the number of cones sold on eight days.

Temperature (°C)16161919212123232424262628283030
Cones sold5252707066668585808098989393112112

Find the line of best fit and rr. Interpret the slope and the intercept.

Solution. With technology, using xx for temperature and yy for cones sold:

y=3.957x−10.500,r≈0.958y = 3.957x - 10.500, \qquad r \approx 0.958

The correlation is strong and positive, so a line is a good model.

  • Slope: each 11 °C rise in temperature predicts about 44 more cones sold.
  • Intercept: at 00 °C the model predicts −10.5-10.5 cones. That’s impossible! The data only go from 1616 °C to 3030 °C, so x=0x = 0 is far outside the data and the intercept has no real meaning here.

Example 2: Interpolation and extrapolation

Section titled “Example 2: Interpolation and extrapolation”

Use the line from Example 1 to predict sales at (a) 2525 °C (b) 4040 °C. How reliable is each prediction?

Solution. Use the rounded equation y=3.957x−10.500y = 3.957x - 10.500 from Example 1.

(a)

y=3.957(25)−10.500=88.425y = 3.957(25) - 10.500 = 88.425

About 8888 cones. Since 2525 °C is inside the data range (1616 °C to 3030 °C), this is interpolation, and with r≈0.958r \approx 0.958 it’s fairly reliable.

(b)

y=3.957(40)−10.500=147.78y = 3.957(40) - 10.500 = 147.78

About 148148 cones. But 4040 °C is well outside the data, so this is extrapolation. On a very hot day, people might stay indoors, or the stand might run out of ice cream. Treat this prediction with caution.

Find the residuals for the days at 2121 °C and 2626 °C.

Solution.

At 2121 °C: predicted y=3.957(21)−10.500≈72.6y = 3.957(21) - 10.500 \approx 72.6. Observed: 6666.

residual=66−72.6=−6.6\text{residual} = 66 - 72.6 = -6.6

At 2626 °C: predicted y=3.957(26)−10.500≈92.4y = 3.957(26) - 10.500 \approx 92.4. Observed: 9898.

residual=98−92.4=5.6\text{residual} = 98 - 92.4 = 5.6

The 2121 °C day sold about 6.66.6 cones fewer than predicted (below the line); the 2626 °C day sold about 5.65.6 more (above the line).

Scatter plot of cones sold against temperature with the line of best fit y = 3.957x - 10.500. Two residuals are drawn as vertical dashed segments: at 21 °C the point (66 cones) is about 6.6 below the line, and at 26 °C the point (98 cones) is about 5.6 above it. 14 16 18 20 22 24 26 28 30 32 40 50 60 70 80 90 100 110 120 temperature (°C) cones sold residual ≈ −6.6 residual ≈ +5.6 y = 3.957x − 10.500
The line of best fit, with two residuals drawn as vertical segments.

On one 2020 °C day, a summer camp group stopped by and the stand sold 115115 cones. Add this point to the data and find the new line and rr.

Solution. With all nine days:

y=3.099x+14.395,r≈0.661y = 3.099x + 14.395, \qquad r \approx 0.661
The ice-cream data with one outlier added at 20 °C and 115 cones. The original line of best fit (dashed) is steep; the new line (solid) is pulled up on the left and is much flatter. 14 16 18 20 22 24 26 28 30 32 40 50 60 70 80 90 100 110 120 temperature (°C) cones sold outlier (20, 115) solid: with outlier dashed: without outlier
One outlier pulls the line toward it.

The single outlier lifted the left side of the line, so the slope fell from about 3.963.96 to about 3.103.10, and rr dropped from strong (0.9580.958) to moderate (0.6610.661). Because the camp visit was a one-time event, it makes sense to leave it out when modelling ordinary sales, and to say that you did.

Mixing up the variables. If you swap xx and yy, technology gives a different line. Put the independent variable in the xx list.

Subtracting the wrong way for a residual. It’s observed minus predicted. A point above the line has a positive residual.

Trusting extrapolations. The model is only supported by the data you have. Predicting far outside that range, or interpreting an intercept far from the data, can give nonsense (like −10.5-10.5 cones).

Interpreting the slope without context or units. Say what changes, by how much, and for each unit of what.

Rounding the equation too early. Rounding aa and bb to one decimal before predicting can noticeably change your answers. Keep at least three significant digits.

Deleting outliers without a reason. Remove a point only if you have a reason (an error, or an unusual event), and always report that you did.

1. (Warm-up) A line of best fit is y=2.5x+12y = 2.5x + 12. Predict yy when x=8x = 8.

Solution

y=2.5(8)+12=32y = 2.5(8) + 12 = 32

2. (Warm-up) A point has observed y=47y = 47, and the line predicts y=50.3y = 50.3. Find the residual. Is the point above or below the line?

Solution

Residual =47−50.3=−3.3= 47 - 50.3 = -3.3. It’s negative, so the point is below the line.

3. (Warm-up) The ice-cream data in Example 1 range from 1616 °C to 3030 °C. Is each prediction interpolation or extrapolation? (a) 1818 °C (b) 3535 °C (c) 1010 °C

Solution

(a) Interpolation. (b) Extrapolation. (c) Extrapolation.

4. (Core) A driver records distance driven and fuel used on six trips.

Distance (km)5050120120180180250250300300410410
Fuel (L)4.34.38.68.613.913.918.118.123.023.029.829.8
  • (a) Find the line of best fit and rr.
  • (b) Interpret the slope.
  • (c) Predict the fuel used on a 200200 km trip.
Solution

(a) y=0.0723x+0.508y = 0.0723x + 0.508, with r≈0.998r \approx 0.998, a very strong positive correlation.

(b) Each extra kilometre uses about 0.07230.0723 L of fuel, or about 7.27.2 L per 100100 km.

(c) y=0.0723(200)+0.508≈15.0y = 0.0723(200) + 0.508 \approx 15.0 L. This is interpolation, and with rr so close to 11, it’s very reliable.

5. (Core) For one house, the monthly heating bill yy (in dollars) and the average outdoor temperature xx (in °C) for the winter months are modelled by y=−4.2x+135y = -4.2x + 135.

  • (a) Interpret the slope and the intercept.
  • (b) Predict the bill for a month averaging −8-8 °C.
Solution

(a) Slope: each 11 °C warmer average temperature predicts a bill about $4.20 lower. Intercept: a month averaging 00 °C predicts a bill of about $135. Since 00 °C is within the range of typical Ontario winter temperatures, this intercept is meaningful.

(b) y=−4.2(−8)+135=33.6+135=168.6y = -4.2(-8) + 135 = 33.6 + 135 = 168.6, so about $168.60.

6. (Core) Using your line from Question 4, find the residuals for the 120120 km and 300300 km trips. Which of the two trips used more fuel than predicted?

Solution

120120 km: predicted 0.0723(120)+0.508≈9.180.0723(120) + 0.508 \approx 9.18 L, so the residual is 8.6−9.18≈−0.588.6 - 9.18 \approx -0.58 L.

300300 km: predicted 0.0723(300)+0.508≈22.200.0723(300) + 0.508 \approx 22.20 L, so the residual is 23.0−22.20≈0.823.0 - 22.20 \approx 0.8 L. (Using unrounded coefficients gives about 0.820.82 L.)

The 300300 km trip had a positive residual, so it used more fuel than predicted.

7. (Core) The hours-studied data from scatter plots and correlation are:

Hours1122223344445566778888
Mark (%)55555252646460607171636374747070858580804545

The last student was sick on test day. Find the line of best fit with and without that student, and compare.

Solution

With: y=1.973x+56.395y = 1.973x + 56.395, r≈0.402r \approx 0.402.

Without: y=4.143x+50.000y = 4.143x + 50.000, r≈0.900r \approx 0.900.

The outlier flattens the line a lot (the slope drops by more than half) and weakens the correlation from strong to moderate. Since there’s a clear reason the point is unusual, it’s reasonable to leave it out and say so.

8. (Challenge) For the points (1,2)(1, 2), (2,5)(2, 5), (3,10)(3, 10), (4,17)(4, 17), (5,26)(5, 26), (6,37)(6, 37):

  • (a) Find the line of best fit and rr.
  • (b) Find all six residuals. What pattern do you see, and what does it tell you?
Solution

(a) y=7x−8.333y = 7x - 8.333, with r≈0.979r \approx 0.979.

(b) The predicted values are −1.333-1.333, 5.6675.667, 12.66712.667, 19.66719.667, 26.66726.667, 33.66733.667, so the residuals are about

3.33, −0.67, −2.67, −2.67, −0.67, 3.333.33,\ -0.67,\ -2.67,\ -2.67,\ -0.67,\ 3.33

They go positive, negative, then positive again: a U shape. Even though rr is very high, the data curve upward, so a straight line is not the best model. (In fact y=x2+1y = x^2 + 1 fits exactly.)

9. (Challenge) Show that the line of best fit from Example 1 passes (very nearly) through the point (xˉ,yˉ)(\bar{x}, \bar{y}).

Solutionxˉ=16+19+21+23+24+26+28+308=1878=23.375\bar{x} = \frac{16 + 19 + 21 + 23 + 24 + 26 + 28 + 30}{8} = \frac{187}{8} = 23.375yˉ=52+70+66+85+80+98+93+1128=6568=82\bar{y} = \frac{52 + 70 + 66 + 85 + 80 + 98 + 93 + 112}{8} = \frac{656}{8} = 82

At x=23.375x = 23.375: y=3.957(23.375)−10.500≈81.995≈82y = 3.957(23.375) - 10.500 \approx 81.995 \approx 82. The tiny difference is from rounding the slope and intercept; with unrounded values the line passes through (23.375,82)(23.375, 82) exactly.