Skip to content
Family Table Math

Scatter Plots and Correlation

Do students who study longer get higher marks? Do taller people have longer arms? Questions like these are about how two attributes of the same person or thing are related. A scatter plot lets you see the relationship, and the correlation coefficient rr puts a number on how strong it is.

In two-variable (bivariate) data, you measure two attributes for each individual, like hours studied and test mark for each student. The goal is to study the relationship between the two.

  • The independent variable (xx) is the one you think might explain or influence the other. It goes on the horizontal axis.
  • The dependent variable (yy) is the one that might respond. It goes on the vertical axis.

For hours studied and test mark, hours studied is independent and test mark is dependent.

A scatter plot shows each individual as one point (x,y)(x, y). Don’t join the dots: each point is a separate individual.

Describe four things:

FeatureWhat to look for
DirectionPositive (yy tends to rise as xx rises) or negative (yy tends to fall)
FormLinear (points follow a line) or non-linear (a curve)
StrengthStrong (points close to a line or curve), moderate, or weak (widely scattered)
OutliersPoints that sit far away from the overall pattern

The correlation coefficient rr measures how closely the points fit a straight line:

−1≤r≤1-1 \le r \le 1
  • r=1r = 1: all points lie exactly on a rising line. r=−1r = -1: exactly on a falling line.
  • rr near 00: no linear relationship.
  • The sign gives the direction; the size ∣r∣|r| gives the strength.
Four scatter plots: r = 0.95 (points tightly along a rising line), r = 0.50 (rising but spread out), r about 0 (no pattern), and r = -0.80 (falling, fairly tight). r = 0.95 strong positive r = 0.50 moderate positive r ≈ 0 no linear correlation r = −0.80 strong negative
Four sets of points and their correlation coefficients.

Many Ontario textbooks use these cut-offs (others draw the lines a little differently):

Value of rDescription
0.67≤r≤10.67 \le r \le 1strong positive
0.33≤r<0.670.33 \le r \lt 0.67moderate positive
0<r<0.330 \lt r \lt 0.33weak positive
r=0r = 0no linear correlation
−0.33<r<0-0.33 \lt r \lt 0weak negative
−0.67<r≤−0.33-0.67 \lt r \le -0.33moderate negative
−1≤r≤−0.67-1 \le r \le -0.67strong negative

You’ll almost always find rr with technology:

  • Spreadsheet: =CORREL(A2:A11, B2:B11), with the xx-values in column A and the yy-values in column B.
  • Graphing calculator (TI-83/84): enter the data in lists L1 and L2, turn on DiagnosticOn, then choose STAT, CALC, LinReg(ax+b).
  • Desmos: enter a table, then type y1 ~ m x1 + b. Desmos shows rr.

For the curious, the formula is:

r=n∑xy−(∑x)(∑y)[n∑x2−(∑x)2][n∑y2−(∑y)2]r = \frac{n\sum xy - \left(\sum x\right)\left(\sum y\right)}{\sqrt{\left[n\sum x^2 - \left(\sum x\right)^2\right]\left[n\sum y^2 - \left(\sum y\right)^2\right]}}

You aren’t expected to use it by hand for real data sets. Answers on this page are rounded to three decimals; your technology may differ in the last digit if you round along the way.

Comparing groups with side-by-side boxplots

Section titled “Comparing groups with side-by-side boxplots”

When one variable is categorical (like grade) and the other is numerical (like hours of sleep), a scatter plot doesn’t work. Instead, draw a boxplot for each category on the same scale and compare their medians, spreads (IQR and range), and outliers. You met boxplots in quartiles and percentiles.

For each pair, identify the independent and dependent variables.

  • (a) The outdoor temperature and the number of hot chocolates a café sells.
  • (b) A car’s age and its resale value.

Solution.

(a) Temperature might affect sales, not the other way around. Independent: temperature. Dependent: hot chocolates sold.

(b) The age of the car influences its value. Independent: age. Dependent: resale value.

Ten students recorded how long they studied for a test and their mark.

Hours studied11222233444455667788
Mark (%)5555525264646060717163637474707085858080

Draw a scatter plot, describe it, and find rr.

Solution.

Scatter plot of test mark against hours studied for 10 students. The points rise from about 55% at 1 hour to about 80-85% at 7-8 hours, roughly along a line. 0 1 2 3 4 5 6 7 8 9 40 50 60 70 80 90 hours studied test mark (%)
Test mark against hours studied for ten students.
  • Direction: positive; students who studied longer tended to score higher.
  • Form: roughly linear.
  • Strength: fairly strong; the points stay close to a line, with some scatter.
  • Outliers: none.

With =CORREL or a calculator:

r≈0.900r \approx 0.900

That’s a strong positive linear correlation, which matches the picture.

Describe each correlation using the table above: (a) r=−0.92r = -0.92 (b) r=0.45r = 0.45 (c) r=−0.18r = -0.18.

Solution.

(a) ∣r∣=0.92≥0.67|r| = 0.92 \ge 0.67 and rr is negative: strong negative.

(b) 0.33≤0.45<0.670.33 \le 0.45 \lt 0.67: moderate positive.

(c) ∣r∣=0.18<0.33|r| = 0.18 \lt 0.33 and rr is negative: weak negative.

A survey asked Grade 9 and Grade 12 students how many hours they sleep on a school night. Compare the groups.

Side-by-side boxplots of hours of sleep. Grade 9: minimum 6.5, Q1 7.5, median 8, Q3 8.5, maximum 9.5. Grade 12: minimum 5, Q1 6, median 7, Q3 7.5, maximum 9. Grade 9 Grade 12 4 5 6 7 8 9 10 hours of sleep on a school night
Hours of sleep on a school night for two grades.

Solution.

  • Centre: the Grade 9 median is 88 hours; the Grade 12 median is 77 hours. Grade 9 students typically sleep about an hour more.
  • Spread: Grade 9 IQR =8.5−7.5=1= 8.5 - 7.5 = 1 hour; Grade 12 IQR =7.5−6=1.5= 7.5 - 6 = 1.5 hours. Ranges are 33 and 44 hours. Grade 12 sleep times are more spread out.
  • Overall: in fact, 75%75\% of Grade 12 students sleep 7.57.5 hours or less, while 75%75\% of Grade 9 students sleep 7.57.5 hours or more.

So there does seem to be a relationship between grade and sleep in this survey.

Swapping the axes. The independent variable goes on the xx-axis. Ask: “Which one might influence the other?”

Thinking a negative r means a weak relationship. The sign is only the direction. r=−0.9r = -0.9 is just as strong as r=0.9r = 0.9.

Using r for a curved pattern. rr measures linear fit only. Points that follow a clear curve can have rr near 00 (see Practice Question 9). Always look at the scatter plot first.

Ignoring outliers. A single outlier can change rr a lot (see Practice Question 6). Check whether it’s a data-entry error or a genuinely unusual individual before deciding what to do.

Reading r as a percentage. r=0.5r = 0.5 does not mean “50% related”. It’s just a number on the scale from −1-1 to 11.

Concluding that x causes y. A strong correlation alone doesn’t prove cause and effect. See correlation and causation.

1. (Warm-up) Identify the independent and dependent variables.

  • (a) The number of hours a phone is used and its remaining battery percentage.
  • (b) A person’s height and their shoe size.
Solution

(a) Independent: hours of use. Dependent: battery percentage.

(b) Height is usually treated as independent and shoe size as dependent, since shoe size is thought of as following from overall body size. (Either choice can be defended here, as long as you explain it.)

2. (Warm-up) Describe each correlation: (a) r=0.82r = 0.82 (b) r=−0.45r = -0.45 (c) r=0.12r = 0.12 (d) r=−0.93r = -0.93

Solution

(a) Strong positive. (b) Moderate negative. (c) Weak positive. (d) Strong negative.

3. (Warm-up) Which shows a stronger linear relationship, r=−0.85r = -0.85 or r=0.70r = 0.70? Explain.

Solution

r=−0.85r = -0.85, because ∣−0.85∣=0.85|-0.85| = 0.85 is larger than 0.700.70. The sign only tells you the direction.

4. (Core) Eight people measured their arm span and height in centimetres.

Arm span (cm)152152160160165165170170174174178178183183190190
Height (cm)155155158158167167168168177177175175186186188188

Use technology to find rr, and describe the correlation.

Solution

r≈0.976r \approx 0.976. This is a strong positive linear correlation: people with longer arm spans tend to be taller, and the points lie very close to a line.

5. (Core) A used-car website lists the asking prices of eight cars of the same model.

Age (years)1122334455667788
Price (thousands of $)27.527.524.024.022.522.518.018.017.517.513.013.012.512.59.09.0

Find rr and describe the scatter plot (direction, form, strength, outliers).

Solution

r≈−0.991r \approx -0.991. The scatter plot shows a strong, negative, linear relationship with no outliers: older cars have lower prices, and the price falls steadily with age.

6. (Core) An eleventh student is added to the data in Example 2: they studied 88 hours but scored 45%45\%. Find the new value of rr. What does this show?

Solution

With the extra point, r≈0.402r \approx 0.402 (down from 0.9000.900). One outlier changed a strong correlation into a moderate one. Outliers can have a big effect on rr, especially in small data sets.

7. (Core) A class recorded their travel times to school. Five-number summaries, in minutes:

MinQ1MedianQ3Max
Bus10101818252530304545
Walk5588121215152020

Describe how side-by-side boxplots of these would compare the two groups.

Solution

The bus box would sit well to the right of the walking box. The median bus trip (2525 min) is more than twice the median walk (1212 min). Bus times are also more spread out: IQR 30−18=1230 - 18 = 12 min versus 15−8=715 - 8 = 7 min, and range 3535 min versus 1515 min. Even the slowest walker (2020 min) is faster than the median bus rider, so method of travel and travel time seem to be related.

8. (Challenge) Use the formula to calculate rr by hand for the points (1,2)(1, 2), (2,4)(2, 4), (3,5)(3, 5), (4,4)(4, 4), (5,6)(5, 6). Check with technology.

Solution

n=5n = 5, ∑x=15\sum x = 15, ∑y=21\sum y = 21, ∑xy=2+8+15+16+30=71\sum xy = 2 + 8 + 15 + 16 + 30 = 71, ∑x2=55\sum x^2 = 55, ∑y2=4+16+25+16+36=97\sum y^2 = 4 + 16 + 25 + 16 + 36 = 97.

n∑xy−(∑x)(∑y)=5(71)−15(21)=40n∑x2−(∑x)2=5(55)−152=50n∑y2−(∑y)2=5(97)−212=44\begin{aligned} n\sum xy - \left(\sum x\right)\left(\sum y\right) &= 5(71) - 15(21) = 40 \\ n\sum x^2 - \left(\sum x\right)^2 &= 5(55) - 15^2 = 50 \\ n\sum y^2 - \left(\sum y\right)^2 &= 5(97) - 21^2 = 44 \end{aligned}r=4050×44=402200≈0.853r = \frac{40}{\sqrt{50 \times 44}} = \frac{40}{\sqrt{2200}} \approx 0.853

Technology gives the same value: a strong positive correlation.

9. (Challenge) Find rr for the points (−2,4)(-2, 4), (−1,1)(-1, 1), (0,0)(0, 0), (1,1)(1, 1), (2,4)(2, 4). Does this mean xx and yy are unrelated?

Solution

r=0r = 0. In the formula, ∑x=0\sum x = 0 and ∑xy=−8−1+0+1+8=0\sum xy = -8 - 1 + 0 + 1 + 8 = 0, so the numerator is 5(0)−0(10)=05(0) - 0(10) = 0.

But xx and yy are perfectly related: every point is on y=x2y = x^2. The relationship is just not linear, so rr can’t detect it. This is why you should always look at the scatter plot, not just rr.