Regression Line of x on y
The usual line of best fit is built to predict from . But sometimes you know and want to estimate : you know a student’s test score and want to estimate how long they revised. For that job there’s a different line, the regression line of on . This page shows why it’s different, how to find it, and how to pick the right line for each prediction.
Key ideas
Section titled “Key ideas”Two different “best” lines
Section titled “Two different “best” lines”The regression line of on , , is found by making the sum of the squared vertical distances from the points to the line as small as possible. That makes it the best line for predicting when you know .
The regression line of on is written with as the subject:
It makes the sum of the squared horizontal distances as small as possible, so it’s the best line for predicting when you know .
Because they minimize different distances, the two lines are different (unless the points lie exactly on a straight line, , when they coincide). The weaker the correlation, the further apart they are.
Both lines pass through the mean point
Section titled “Both lines pass through the mean point”Both regression lines pass through the mean point . So if you’re given both equations, solving them simultaneously gives and .
Which line to use
Section titled “Which line to use”| You know | You want | Use |
|---|---|---|
| the line of on : | ||
| the line of on : |
You cannot get the line of on by rearranging to make the subject. That gives the same line written differently, which is still the best line for predicting , not . Likewise, you can’t reliably predict from using the line of on .
Finding the line with technology
Section titled “Finding the line with technology”On a GDC, the line of on is found by running the usual linear regression with the lists swapped: put the -values in the list the calculator treats as "" and the -values in the list it treats as "". The calculator reports something like "", but remember its "" is really your , so rewrite the result as .
The correlation coefficient is the same either way round.
Making predictions
Section titled “Making predictions”The same cautions as for on apply:
- Interpolation (the known value is inside the range of the data) is reasonably reliable when the correlation is strong.
- Extrapolation (the known value is outside the range of the data) is unreliable: the pattern may not continue, and it can give impossible answers.
- If the correlation is weak, no prediction is very reliable.
For a line of on , “inside the range” means inside the range of the -values, since is what you’re given.
Worked examples
Section titled “Worked examples”Example 1: Choosing the right line
Section titled “Example 1: Choosing the right line”A café records the daily maximum temperature () and the number of iced drinks sold, , for summer days. Which regression line should be used to:
- (a) predict the number of iced drinks sold on a day forecast to reach ;
- (b) estimate the maximum temperature on a day when iced drinks were sold?
Solution.
(a) The temperature is known and the sales are wanted: use the line of on .
(b) The sales are known and the temperature is wanted: use the line of on .
Example 2: Finding both lines and predicting
Section titled “Example 2: Finding both lines and predicting”The table shows the hours of revision, , and the test score (%), , for students.
- (a) Find the regression line of on and use it to predict the score of a student who revised for hours.
- (b) Find the regression line of on and use it to estimate how long a student who scored revised.
Solution.
(a) Enter in list 1 and in list 2 and run linear regression:
For : . The predicted score is about .
(b) Run the regression again with the lists swapped ( as the input list, as the output list):
For : hours (using the unrounded coefficients). Both predictions are interpolations ( is inside to , and is inside to ) and shows strong positive correlation, so both are reasonably reliable.
Check: the mean point is . Both lines pass through it: and . ✓
Example 3: Using the mean point
Section titled “Example 3: Using the mean point”For a set of data, the regression line of on is , and the regression line of on is .
- (a) Find and .
- (b) Estimate when , and estimate when .
Solution.
(a) Both lines pass through , so solve them simultaneously. Substitute the first into the second:
Then . So and .
(b) Given , use the line of on : .
Given , use the line of on : .
Example 4: When predictions go wrong
Section titled “Example 4: When predictions go wrong”Using the revision data from Example 2, a student uses to estimate the revision time for scores of (a) and (b) . Comment on each answer.
Solution.
(a) hours. The score is outside the observed scores ( to ), so this is extrapolation and unreliable. The pattern might not continue: scores can’t rise forever as revision time increases.
(b) hours. A negative revision time is impossible. Again, is far outside the observed scores, and the model shouldn’t be used there.
Common mistakes
Section titled “Common mistakes”Rearranging the y-on-x line. Making the subject of gives , which predicts for . That isn’t the line of on , and its answer differs from the correct . Always run the regression with the lists swapped.
Using the x-on-y line to predict y. The guide is clear that predicting from with an -on- line isn’t reliable. Match the line to what you know.
Reading the GDC output literally. After swapping the lists, the calculator still writes "". Rewrite it as so you don’t plug a value into the wrong letter.
Mixing up which equation is which. A line written as is the line of on . If both equations are given as "", read the question carefully to see which is which.
Checking the wrong range for extrapolation. For an -on- prediction, compare the given -value with the range of the observed -values, not the -values.
Forgetting the mean point. Questions often give both lines and ask for and . Solve the equations simultaneously; you don’t need the original data.
Practice
Section titled “Practice”1. (Warm-up) At an outdoor pool, is the daily maximum temperature () and is the number of swimmers. Which regression line should be used to:
- (a) predict the number of swimmers when the forecast is ;
- (b) estimate the temperature on a day when swimmers came?
Solution
(a) is known, is wanted: the line of on .
(b) is known, is wanted: the line of on .
2. (Warm-up) The regression line of on is , and the regression line of on is . Find and .
Solution
Both lines pass through . Substitute:
. So and .
3. (Core) The age (years) and price (thousands of dollars) of used cars of the same model are shown.
Use your GDC to find the regression line of on . Use it to estimate the age of a car of this model priced at $21 000.
Solution
Run linear regression with the prices as the input list and the ages as the output list:
For a price of $21 000, :
(using the unrounded coefficients). This is interpolation ( is between and ) and the correlation is very strong, so it’s reliable.
4. (Core) For the cars in question 3, a buyer wants to predict the price of a -year-old car. Explain why the line from question 3 shouldn’t be used, find the correct line, and make the prediction.
Solution
The age is known and the price is wanted, so the line of on is needed. The line of on isn’t designed to predict from .
GDC (ages as input, prices as output):
For : , so about $12 100 (using unrounded coefficients). Since is inside the ages to , this is interpolation.
5. (Core) Using the line from question 3, estimate the age of a car priced at $5000, and comment on the reliability of your answer.
Solution
: years.
The prices in the data run from to thousand dollars, so is outside the range: this is extrapolation. Even though is very close to , the straight-line pattern may not continue (car prices usually level off rather than keep dropping at the same rate), so the estimate is unreliable.
6. (Core) For a data set, . The regression line of on is , and the regression line of on has gradient .
- (a) Find .
- (b) Find the equation of the regression line of on .
Solution
(a) The line of on passes through : , so and .
(b) The line of on also passes through : , so .
7. (Core) For the revision data in Example 2, a student wants to estimate the revision time for a score of . She rearranges to get .
- (a) Find her estimate.
- (b) Explain why it differs from the estimate in Example 2, and which estimate is better.
Solution
(a) hours. (With the unrounded GDC coefficients, hours.)
(b) Rearranging doesn’t change the line: it’s still the line of on , which minimizes vertical distances and is designed for predicting . The line of on minimizes horizontal distances, so it’s the best line for predicting . The estimate of hours from Example 2 is the better one.
8. (Challenge) The gradients of the two regression lines can be written as (for on ) and (for on ), where , are the standard deviations of and .
- (a) Show that .
- (b) For Example 3, the lines are and . Find .
- (c) Use (a) to explain why the two lines are the same line when .
Solution
(a)
(b) . Both gradients are positive, so is positive: (3 s.f.).
(c) If , then , so . The line of on is , which rearranges to : a line with the same gradient as . Both lines also pass through , and two lines with the same gradient through the same point are the same line.
9. (Challenge) The two regression lines for a data set are and .
- (a) State which line is the line of on .
- (b) Find the mean point.
- (c) Using the result of question 8, find .
- (d) Estimate when .
Solution
(a) has as the subject, so it’s the line of on .
(b) Substitute into the other line:
. The mean point is .
(c) . Both gradients are negative, so (3 s.f.).
(d) is known, so use the line of on : .