Skip to content
Family Table Math
Auto

Regression Line of x on y

The usual line of best fit is built to predict yy from xx. But sometimes you know yy and want to estimate xx: you know a student’s test score and want to estimate how long they revised. For that job there’s a different line, the regression line of xx on yy. This page shows why it’s different, how to find it, and how to pick the right line for each prediction.

The regression line of yy on xx, y=ax+by = ax + b, is found by making the sum of the squared vertical distances from the points to the line as small as possible. That makes it the best line for predicting yy when you know xx.

The regression line of xx on yy is written with xx as the subject:

x=cy+dx = cy + d

It makes the sum of the squared horizontal distances as small as possible, so it’s the best line for predicting xx when you know yy.

Because they minimize different distances, the two lines are different (unless the points lie exactly on a straight line, r=±1r = \pm 1, when they coincide). The weaker the correlation, the further apart they are.

Both regression lines pass through the mean point (xˉ,yˉ)(\bar{x}, \bar{y}). So if you’re given both equations, solving them simultaneously gives xˉ\bar{x} and yˉ\bar{y}.

Scatter plot of test score against hours of revision with the two regression lines, which cross at the mean point (6.6, 66). 0 1 2 3 4 5 6 7 8 9 10 11 12 13 40 50 60 70 80 90 (6.6, 66) y on x x on y hours of revision, x test score (%), y
The line of yy on xx (blue) and the line of xx on yy (orange) are different lines, but both pass through the mean point (xˉ,yˉ)=(6.6,66)(\bar{x}, \bar{y}) = (6.6, 66).
You knowYou wantUse
xxyythe line of yy on xx:  y=ax+b\ y = ax + b
yyxxthe line of xx on yy:  x=cy+d\ x = cy + d

You cannot get the line of xx on yy by rearranging y=ax+by = ax + b to make xx the subject. That gives the same line written differently, which is still the best line for predicting yy, not xx. Likewise, you can’t reliably predict yy from xx using the line of xx on yy.

On a GDC, the line of xx on yy is found by running the usual linear regression with the lists swapped: put the yy-values in the list the calculator treats as "xx" and the xx-values in the list it treats as "yy". The calculator reports something like "y=ax+by = ax + b", but remember its "xx" is really your yy, so rewrite the result as x=cy+dx = cy + d.

The correlation coefficient rr is the same either way round.

The same cautions as for yy on xx apply:

  • Interpolation (the known value is inside the range of the data) is reasonably reliable when the correlation is strong.
  • Extrapolation (the known value is outside the range of the data) is unreliable: the pattern may not continue, and it can give impossible answers.
  • If the correlation is weak, no prediction is very reliable.

For a line of xx on yy, “inside the range” means inside the range of the yy-values, since yy is what you’re given.

A café records the daily maximum temperature xx (∘C^\circ\text{C}) and the number of iced drinks sold, yy, for 3030 summer days. Which regression line should be used to:

  • (a) predict the number of iced drinks sold on a day forecast to reach 27 ∘C27\,^\circ\text{C};
  • (b) estimate the maximum temperature on a day when 140140 iced drinks were sold?

Solution.

(a) The temperature xx is known and the sales yy are wanted: use the line of yy on xx.

(b) The sales yy are known and the temperature xx is wanted: use the line of xx on yy.

Example 2: Finding both lines and predicting

Section titled “Example 2: Finding both lines and predicting”

The table shows the hours of revision, xx, and the test score (%), yy, for 1010 students.

xx223344556677889910101212
yy4848606052526666585874746565808072728585
  • (a) Find the regression line of yy on xx and use it to predict the score of a student who revised for 6.56.5 hours.
  • (b) Find the regression line of xx on yy and use it to estimate how long a student who scored 70%70\% revised.

Solution.

(a) Enter xx in list 1 and yy in list 2 and run linear regression:

y=3.29x+44.3(r=0.885, to 3 s.f.)y = 3.29x + 44.3 \qquad (r = 0.885, \text{ to 3 s.f.})

For x=6.5x = 6.5: y≈3.29(6.5)+44.3≈65.7y \approx 3.29(6.5) + 44.3 \approx 65.7. The predicted score is about 65.7%65.7\%.

(b) Run the regression again with the lists swapped (yy as the input list, xx as the output list):

x=0.238y−9.10x = 0.238y - 9.10

For y=70y = 70: x≈0.238(70)−9.10≈7.55x \approx 0.238(70) - 9.10 \approx 7.55 hours (using the unrounded coefficients). Both predictions are interpolations (6.56.5 is inside 22 to 1212, and 7070 is inside 4848 to 8585) and r=0.885r = 0.885 shows strong positive correlation, so both are reasonably reliable.

Check: the mean point is (xˉ,yˉ)=(6.6,66)(\bar{x}, \bar{y}) = (6.6, 66). Both lines pass through it: 3.29(6.6)+44.3≈66.03.29(6.6) + 44.3 \approx 66.0 and 0.238(66)−9.10≈6.610.238(66) - 9.10 \approx 6.61. ✓

For a set of data, the regression line of yy on xx is y=0.5x+8y = 0.5x + 8, and the regression line of xx on yy is x=1.5y−7x = 1.5y - 7.

  • (a) Find xˉ\bar{x} and yˉ\bar{y}.
  • (b) Estimate xx when y=22y = 22, and estimate yy when x=24x = 24.

Solution.

(a) Both lines pass through (xˉ,yˉ)(\bar{x}, \bar{y}), so solve them simultaneously. Substitute the first into the second:

x=1.5(0.5x+8)−7x=0.75x+50.25x=5x=20\begin{aligned} x &= 1.5(0.5x + 8) - 7 \\ x &= 0.75x + 5 \\ 0.25x &= 5 \\ x &= 20 \end{aligned}

Then y=0.5(20)+8=18y = 0.5(20) + 8 = 18. So xˉ=20\bar{x} = 20 and yˉ=18\bar{y} = 18.

(b) Given yy, use the line of xx on yy: x=1.5(22)−7=26x = 1.5(22) - 7 = 26.

Given xx, use the line of yy on xx: y=0.5(24)+8=20y = 0.5(24) + 8 = 20.

Using the revision data from Example 2, a student uses x=0.238y−9.10x = 0.238y - 9.10 to estimate the revision time for scores of (a) 95%95\% and (b) 30%30\%. Comment on each answer.

Solution.

(a) x≈0.238(95)−9.10≈13.5x \approx 0.238(95) - 9.10 \approx 13.5 hours. The score 9595 is outside the observed scores (4848 to 8585), so this is extrapolation and unreliable. The pattern might not continue: scores can’t rise forever as revision time increases.

(b) x≈0.238(30)−9.10≈−1.96x \approx 0.238(30) - 9.10 \approx -1.96 hours. A negative revision time is impossible. Again, 3030 is far outside the observed scores, and the model shouldn’t be used there.

Rearranging the y-on-x line. Making xx the subject of y=3.29x+44.3y = 3.29x + 44.3 gives x=y−44.33.29x = \dfrac{y - 44.3}{3.29}, which predicts x≈7.82x \approx 7.82 for y=70y = 70. That isn’t the line of xx on yy, and its answer differs from the correct 7.557.55. Always run the regression with the lists swapped.

Using the x-on-y line to predict y. The guide is clear that predicting yy from xx with an xx-on-yy line isn’t reliable. Match the line to what you know.

Reading the GDC output literally. After swapping the lists, the calculator still writes "y=ax+by = ax + b". Rewrite it as x=ay+bx = ay + b so you don’t plug a value into the wrong letter.

Mixing up which equation is which. A line written as x=…x = \ldots is the line of xx on yy. If both equations are given as "y=…y = \ldots", read the question carefully to see which is which.

Checking the wrong range for extrapolation. For an xx-on-yy prediction, compare the given yy-value with the range of the observed yy-values, not the xx-values.

Forgetting the mean point. Questions often give both lines and ask for xˉ\bar{x} and yˉ\bar{y}. Solve the equations simultaneously; you don’t need the original data.

1. (Warm-up) At an outdoor pool, xx is the daily maximum temperature (∘C^\circ\text{C}) and yy is the number of swimmers. Which regression line should be used to:

  • (a) predict the number of swimmers when the forecast is 28 ∘C28\,^\circ\text{C};
  • (b) estimate the temperature on a day when 350350 swimmers came?
Solution

(a) xx is known, yy is wanted: the line of yy on xx.

(b) yy is known, xx is wanted: the line of xx on yy.

2. (Warm-up) The regression line of yy on xx is y=2x+3y = 2x + 3, and the regression line of xx on yy is x=0.4y+1x = 0.4y + 1. Find xˉ\bar{x} and yˉ\bar{y}.

Solution

Both lines pass through (xˉ,yˉ)(\bar{x}, \bar{y}). Substitute:

x=0.4(2x+3)+1=0.8x+2.2⇒0.2x=2.2⇒x=11x = 0.4(2x + 3) + 1 = 0.8x + 2.2 \quad\Rightarrow\quad 0.2x = 2.2 \quad\Rightarrow\quad x = 11

y=2(11)+3=25y = 2(11) + 3 = 25. So xˉ=11\bar{x} = 11 and yˉ=25\bar{y} = 25.

3. (Core) The age xx (years) and price yy (thousands of dollars) of 77 used cars of the same model are shown.

xx11223344556688
yy2828252523231919181814141010

Use your GDC to find the regression line of xx on yy. Use it to estimate the age of a car of this model priced at $21 000.

Solution

Run linear regression with the prices as the input list and the ages as the output list:

x=−0.381y+11.6(r=−0.995, to 3 s.f.)x = -0.381y + 11.6 \qquad (r = -0.995, \text{ to 3 s.f.})

For a price of $21 000, y=21y = 21:

x≈−0.381(21)+11.6≈3.60 yearsx \approx -0.381(21) + 11.6 \approx 3.60 \text{ years}

(using the unrounded coefficients). This is interpolation (2121 is between 1010 and 2828) and the correlation is very strong, so it’s reliable.

4. (Core) For the cars in question 3, a buyer wants to predict the price of a 77-year-old car. Explain why the line from question 3 shouldn’t be used, find the correct line, and make the prediction.

Solution

The age xx is known and the price yy is wanted, so the line of yy on xx is needed. The line of xx on yy isn’t designed to predict yy from xx.

GDC (ages as input, prices as output):

y=−2.60x+30.3y = -2.60x + 30.3

For x=7x = 7: y≈−2.60(7)+30.3≈12.1y \approx -2.60(7) + 30.3 \approx 12.1, so about $12 100 (using unrounded coefficients). Since 77 is inside the ages 11 to 88, this is interpolation.

5. (Core) Using the line from question 3, estimate the age of a car priced at $5000, and comment on the reliability of your answer.

Solution

y=5y = 5: x≈−0.381(5)+11.6≈9.69x \approx -0.381(5) + 11.6 \approx 9.69 years.

The prices in the data run from 1010 to 2828 thousand dollars, so 55 is outside the range: this is extrapolation. Even though rr is very close to −1-1, the straight-line pattern may not continue (car prices usually level off rather than keep dropping at the same rate), so the estimate is unreliable.

6. (Core) For a data set, xˉ=20\bar{x} = 20. The regression line of xx on yy is x=0.8y−4x = 0.8y - 4, and the regression line of yy on xx has gradient 1.051.05.

  • (a) Find yˉ\bar{y}.
  • (b) Find the equation of the regression line of yy on xx.
Solution

(a) The line of xx on yy passes through (xˉ,yˉ)(\bar{x}, \bar{y}): 20=0.8yˉ−420 = 0.8\bar{y} - 4, so 0.8yˉ=240.8\bar{y} = 24 and yˉ=30\bar{y} = 30.

(b) The line of yy on xx also passes through (20,30)(20, 30): 30=1.05(20)+b30 = 1.05(20) + b, so b=9b = 9.

y=1.05x+9y = 1.05x + 9

7. (Core) For the revision data in Example 2, a student wants to estimate the revision time for a score of 70%70\%. She rearranges y=3.29x+44.3y = 3.29x + 44.3 to get x=y−44.33.29x = \dfrac{y - 44.3}{3.29}.

  • (a) Find her estimate.
  • (b) Explain why it differs from the estimate in Example 2, and which estimate is better.
Solution

(a) x=70−44.33.29≈7.81x = \dfrac{70 - 44.3}{3.29} \approx 7.81 hours. (With the unrounded GDC coefficients, 7.827.82 hours.)

(b) Rearranging doesn’t change the line: it’s still the line of yy on xx, which minimizes vertical distances and is designed for predicting yy. The line of xx on yy minimizes horizontal distances, so it’s the best line for predicting xx. The estimate of 7.557.55 hours from Example 2 is the better one.

8. (Challenge) The gradients of the two regression lines can be written as a=rsysxa = r\dfrac{s_y}{s_x} (for yy on xx) and c=rsxsyc = r\dfrac{s_x}{s_y} (for xx on yy), where sxs_x, sys_y are the standard deviations of xx and yy.

  • (a) Show that ac=r2ac = r^2.
  • (b) For Example 3, the lines are y=0.5x+8y = 0.5x + 8 and x=1.5y−7x = 1.5y - 7. Find rr.
  • (c) Use (a) to explain why the two lines are the same line when r=±1r = \pm 1.
Solution

(a)

ac=rsysx×rsxsy=r2ac = r\frac{s_y}{s_x} \times r\frac{s_x}{s_y} = r^2

(b) r2=0.5×1.5=0.75r^2 = 0.5 \times 1.5 = 0.75. Both gradients are positive, so rr is positive: r=0.75≈0.866r = \sqrt{0.75} \approx 0.866 (3 s.f.).

(c) If r=±1r = \pm 1, then ac=1ac = 1, so c=1ac = \dfrac{1}{a}. The line of xx on yy is x=1ay+dx = \dfrac{1}{a}y + d, which rearranges to y=ax−ady = ax - ad: a line with the same gradient as y=ax+by = ax + b. Both lines also pass through (xˉ,yˉ)(\bar{x}, \bar{y}), and two lines with the same gradient through the same point are the same line.

9. (Challenge) The two regression lines for a data set are y=−0.5x+20y = -0.5x + 20 and x=−1.6y+34x = -1.6y + 34.

  • (a) State which line is the line of xx on yy.
  • (b) Find the mean point.
  • (c) Using the result of question 8, find rr.
  • (d) Estimate yy when x=12x = 12.
Solution

(a) x=−1.6y+34x = -1.6y + 34 has xx as the subject, so it’s the line of xx on yy.

(b) Substitute y=−0.5x+20y = -0.5x + 20 into the other line:

x=−1.6(−0.5x+20)+34=0.8x+2⇒0.2x=2⇒x=10x = -1.6(-0.5x + 20) + 34 = 0.8x + 2 \quad\Rightarrow\quad 0.2x = 2 \quad\Rightarrow\quad x = 10

y=−0.5(10)+20=15y = -0.5(10) + 20 = 15. The mean point is (10,15)(10, 15).

(c) r2=(−0.5)(−1.6)=0.8r^2 = (-0.5)(-1.6) = 0.8. Both gradients are negative, so r=−0.8≈−0.894r = -\sqrt{0.8} \approx -0.894 (3 s.f.).

(d) xx is known, so use the line of yy on xx: y=−0.5(12)+20=14y = -0.5(12) + 20 = 14.