Skip to content
Family Table Math

Correlation and Causation

When two variables are strongly correlated, it’s tempting to say one causes the other. Sometimes that’s true, but often it isn’t. Learning to ask “what else could explain this?” is one of the most useful skills in statistics, and it will help you spot misleading claims in the news and in ads.

A strong correlation tells you that two variables tend to change together. It does not tell you why. To show that xx causes yy, you usually need a well-designed experiment, where you change xx on purpose and control everything else (see survey and experiment design).

TypeWhat’s going onExample
Cause and effectA change in xx directly causes a change in yy.The more kilometres you drive, the more fuel you use.
Reverse cause and effectThe cause runs the other way: yy causes xx.Students who get extra help often have lower marks, but low marks lead to extra help, not the reverse.
Common causeA third variable, a lurking (hidden) variable, causes both xx and yy.Ice-cream sales and sunburns rise together because both increase on hot, sunny days.
AccidentalThe correlation is pure coincidence; there’s no logical connection.The number of library cards issued in a town and the goals scored by a hockey team in the same seasons.
PresumedThere seems to be a logical connection, but no clear cause and effect and no obvious common cause.People who exercise more tend to watch fewer hours of TV.

A lurking variable is a variable that isn’t in your data but affects the variables you measured. When you see a correlation, ask: “Is there some third factor that could push both of these up or down?” Common lurking variables include time, temperature, age, population size, and income.

Watch for these when you read a claim:

  • Weak correlations presented as strong. A study finds r=0.25r = 0.25 and reports a “clear link”. By the usual cut-offs, that’s weak.
  • Assuming causation. “Students who eat breakfast get better marks, so breakfast raises marks.” Maybe, but family routines, sleep, or income could explain both.
  • Small samples. A correlation from 66 data points can be coincidence.
  • Extrapolating. Using a trend far outside the data range (see linear regression).
  • Hiding outliers or cherry-picking data to make a correlation look stronger.

Classify each relationship and explain.

  • (a) The mass of apples in a bag and the price of the bag.
  • (b) Across Ontario towns, the number of fire stations and the number of fires per year.
  • (c) Over several years, the number of people who watched a cooking show and a country’s honey production.

Solution.

(a) Cause and effect. Apples sold by the kilogram: more mass directly means a higher price.

(b) Common cause. Bigger towns have more people, more buildings, and so more fires, and they also need more fire stations. Population is the lurking variable. Fire stations don’t cause fires. (There’s some reverse cause too, since towns with many fires may build more stations, but the main driver is population.)

(c) Accidental. There’s no sensible connection; any correlation is a coincidence.

A beach clinic records weekly ice-cream cone sales at the nearby stand and the number of sunburn visits for eight summer weeks.

Cones sold42042051051064064070070082082091091098098010501050
Sunburn visits33554488998812121313

Find rr. Does eating ice cream cause sunburn?

Solution. With technology, r≈0.932r \approx 0.932: a strong positive correlation.

But ice cream doesn’t cause sunburn. Hot, sunny weeks bring more people to the beach, who buy more ice cream and spend more time in the sun. Weather (and the number of beach visitors) is the lurking variable. This is a common-cause relationship.

A survey finds that people who own more pairs of running shoes tend to run more kilometres per week. A shoe store’s ad says, “Buy more shoes, and you’ll run more!” What’s wrong?

Solution. The ad assumes owning shoes causes running. It’s much more likely to be reverse cause and effect: people who run a lot wear out shoes faster, so they own more pairs. Running causes shoe buying, not the other way around.

A blog post says: “Our survey of 1515 students found a correlation of r=0.28r = 0.28 between hours of music listened to per day and math marks. Music makes you better at math!” Give two problems with this claim.

Solution.

  1. The correlation is weak. r=0.28r = 0.28 is below 0.330.33, so it’s a weak positive correlation. The points would be widely scattered; music time is a poor predictor of marks.
  2. It assumes causation. Even a strong correlation wouldn’t prove music causes better marks. A lurking variable (for example, students who study with music may also study longer) could explain it.

Also, 1515 students is a small sample, so even this weak correlation could be due to chance.

Saying “x causes y” from a scatter plot. A scatter plot and rr show association only. Use words like “is associated with” or “tends to increase with” unless an experiment supports cause and effect.

Thinking a strong correlation is more likely to be causal. Strength doesn’t tell you the type. Common-cause and accidental correlations can be very strong, like r≈0.932r \approx 0.932 in Example 2.

Forgetting about time. Two quantities that both grow over the years (phone ownership and average house prices, say) will be correlated even if they’re unrelated. Time is a common lurking variable.

Mixing up common cause and reverse cause. In common cause, a third variable drives both. In reverse cause, the two variables are linked directly, but in the opposite direction from what was claimed.

Overlooking weak correlations dressed up as strong. Check the actual value of rr before believing words like “strong link”.

1. (Warm-up) Classify each relationship as cause and effect, common cause, reverse cause and effect, or accidental.

  • (a) Hours a kettle is switched on and electricity used.
  • (b) The number of umbrellas sold and the number of car accidents on the same days.
  • (c) A city’s number of dentists and its number of pizza restaurants.
Solution

(a) Cause and effect.

(b) Common cause: rainy days bring both more umbrella sales and more accidents.

(c) Common cause: both depend on the city’s population.

2. (Warm-up) In your own words, what is a lurking variable?

Solution

A lurking variable is a third variable, not included in the study, that affects both of the variables being compared. It can create a correlation between them even if neither causes the other.

3. (Warm-up) A study reports r=−0.21r = -0.21 between two variables and calls it “a strong negative relationship”. Is that fair?

Solution

No. The direction is negative, but ∣r∣=0.21<0.33|r| = 0.21 \lt 0.33, so it’s a weak correlation.

4. (Core) Across Canadian cities, the number of coffee shops is strongly correlated with the number of traffic lights. Identify a lurking variable and explain.

Solution

City size (population or area). Bigger cities have more people buying coffee and more roads and intersections needing traffic lights. Neither one causes the other.

5. (Core) Over six seasons, a town’s library recorded new library cards issued, and the local hockey team’s goals:

Library cards121012101350135011801180142014201390139015051505
Goals221221240240215215236236251251262262

Find rr. What type of relationship is this most likely to be?

Solution

r≈0.933r \approx 0.933, a strong positive correlation. There’s no sensible reason library cards would affect goals (or vice versa), and no obvious common cause, so it’s most likely accidental. With only six data points, a coincidence like this is not surprising.

6. (Core) “Hospitals with more doctors have more deaths per year, so doctors are dangerous.” Explain the flaw.

Solution

Larger hospitals serve more patients, including the most seriously ill, so they have both more doctors and more deaths. Hospital size (number of patients) is a common cause. A fair comparison would look at death rates for similar patients, not total deaths.

7. (Core) Students who sleep more tend to report less stress. Give one cause-and-effect explanation, one reverse cause-and-effect explanation, and one possible lurking variable.

Solution

Cause and effect: more sleep helps the body and mind recover, reducing stress.

Reverse: stressed students lie awake worrying, so stress causes less sleep.

Lurking variable: workload. A heavy load of homework or a part-time job could cause both less sleep and more stress.

(Other reasonable answers are fine.)

8. (Challenge) A company says, “Customers who use our app have 20%20\% higher grades on average, so our app improves grades.” Describe how you could design a study to test whether the app actually causes higher grades.

Solution

Use a controlled experiment. Randomly assign a large group of students to two groups: one uses the app, and one doesn’t (or uses a different study method for the same amount of time). Make the groups as alike as possible in every other way, and measure grades before and after. Random assignment spreads lurking variables (like motivation or family support) evenly across the groups, so if the app group improves noticeably more, you have good evidence that the app caused it.

9. (Challenge) Over 2020 years, the number of smartphones in Canada and the average price of a concert ticket both rose steadily, giving r=0.96r = 0.96. A student says, “Phones are making concerts more expensive.” Explain what’s really going on, and suggest how to remove its effect.

Solution

Both variables increased over time: technology spread and prices rose with inflation and demand. Time is the lurking variable, so the correlation doesn’t show that phones affect ticket prices. To reduce its effect, you could compare the year-to-year changes in each variable instead of the totals, or adjust prices for inflation, and see whether a relationship remains.