Skip to content
Family Table Math

Bias in Sampling

A sample is supposed to be a small, fair picture of a population. Bias is anything in the way data is collected that pushes the results consistently in one direction, so the picture is distorted. A biased study can collect thousands of responses and still give the wrong answer, so spotting and reducing bias is one of the most useful skills in statistics.

Results from any sample will differ a little from the true population value just by chance. That’s sampling variability, and it goes both ways: sometimes a bit high, sometimes a bit low. Larger samples reduce it.

Bias is different: it’s a systematic error that tends to push results the same way every time. Making the sample bigger doesn’t fix it. You have to fix the method.

TypeWhat happensExample
Sampling biasthe sampling method favours some parts of the population, or leaves some outsurveying students about bus service by asking only students in the parking lot
Non-response biassome chosen individuals don’t respond, and they differ from those who doonly 20%20\% of mailed surveys come back, mostly from people with strong opinions
Response biaspeople give false or misleading answers, on purpose or notstudents under-report how much time they spend on their phones when the teacher is watching
Measurement biasthe way data is measured consistently overestimates or underestimatesa bathroom scale that always reads 0.50.5 kg too high; a question worded to push one answer

Even a random sample can be biased if:

  • the list you sample from (the sampling frame) leaves people out, like a phone survey that only calls landlines and misses households with only cell phones
  • many chosen people don’t respond
  • the questions or measurements are biased

Randomness protects against the person choosing the sample playing favourites. It doesn’t protect against a bad list, a low response rate, or a leading question.

Convenience and voluntary response samples are almost always biased:

  • A convenience sample only reaches people who are easy to reach, and they’re often similar to each other (and to the person asking).
  • A voluntary response sample over-represents people with strong feelings, especially negative ones. Online polls and call-in votes are classic examples.
BiasHow to reduce it
Sampling biasuse a random method; make sure the sampling frame includes the whole population; stratify if some groups might be missed
Non-response biasfollow up with non-responders; keep surveys short; offer several ways to respond; report the response rate
Response biasmake responses anonymous; use neutral interviewers; avoid questions that embarrass people
Measurement biascalibrate instruments; use neutral question wording; test the survey on a small group first

You’ll learn to write fair questions in survey and experiment design.

Identify the type of bias in each situation.

  • (a) To estimate how many students play sports, a student surveys people in the school gym after practice.
  • (b) A teacher asks her class, out loud, “Who did all of last night’s homework?” and counts the hands.
  • (c) A town mails a survey to 400400 random households, and only 9292 are returned.
  • (d) Students measure their heights against a wall chart that starts 22 cm above the floor.

Solution.

(a) Sampling bias. Students in the gym after practice are far more likely to play sports than the average student.

(b) Response bias. Students who didn’t do the homework may raise their hands anyway rather than admit it in front of the teacher and class.

(c) Non-response bias. The response rate is 92400=0.23\tfrac{92}{400} = 0.23, or 23%23\%. The households that replied may care more about the issue than the 77%77\% that didn’t.

(d) Measurement bias. Every height is read from a chart that is 22 cm off, so every measurement is 22 cm too low. (A height of 165165 cm reads as 163163 cm.)

A school has 12001200 students, and 300300 of them are in Grade 12. A student surveys 5050 people in the senior lounge about whether the school should start later in the morning. 3838 of the 5050 are in Grade 12. Explain why the results may be biased.

Solution. In the whole school, Grade 12 students make up 3001200=25%\tfrac{300}{1200} = 25\%. In the sample, they make up 3850=76%\tfrac{38}{50} = 76\%. Grade 12 is heavily over-represented, and other grades are under-represented.

If Grade 12 students feel differently about start times (for example, because more of them have jobs or drive), the survey’s result won’t reflect the whole school. This is a convenience sample with sampling bias.

A fix: use a stratified random sample from the full student list, with about 25%25\% of the sample from Grade 12.

Example 3: A random sample that’s still biased

Section titled “Example 3: A random sample that’s still biased”

A polling firm picks phone numbers randomly from an old directory of landline numbers to survey a city’s residents about public transit. Why might the results be biased, even though the sample is random?

Solution. The sampling frame is incomplete. Many residents, especially younger people and renters, don’t have a listed landline, so they have no chance of being chosen. Younger people may also use public transit differently, so leaving them out can push the results one way. This is sampling bias.

A fix: use a frame that includes cell phones (random digit dialling) or combine phone, online, and mail methods, and check that the sample’s ages match the city’s.

A website asks, “Should our city ban gas-powered leaf blowers?” and gets 35003500 votes: 81%81\% say yes. A councillor says, “With 35003500 votes, this must reflect the city.” Is she right?

Solution. No. This is a voluntary response sample: only people who visit the site and care enough to vote are counted. People who are annoyed by leaf blowers are probably more motivated to vote than people who don’t mind them. One person may also be able to vote many times.

A large sample doesn’t remove bias. A random sample of a few hundred residents would give a more trustworthy picture.

Thinking a larger sample removes bias. Size reduces random variability, but a biased method stays biased no matter how many people respond. Fix the method first.

Assuming a random sample is automatically unbiased. Random selection from an incomplete list, low response rates, and leading questions can all bias a random sample.

Mixing up response bias and non-response bias. Response bias is about the answers people give (they’re inaccurate). Non-response bias is about people who don’t answer at all.

Mixing up response bias and measurement bias. If the tool is off (a broken scale, a loaded question), it’s measurement bias. If the person gives an inaccurate answer (embarrassment, wanting to please), it’s response bias. Some situations involve both.

Naming the bias without explaining its direction. Good answers say which way the results are pushed: “This survey will probably overestimate the number of students who play sports, because…”

1. (Warm-up) Match each situation with sampling, non-response, response, or measurement bias.

  • (a) A thermometer always reads 1 ∘C1\,^\circ\text{C} too high.
  • (b) Only half the students given a survey hand it back.
  • (c) A survey about reading habits is done only at the public library.
  • (d) Students exaggerate how much they exercise because their gym teacher is asking.
Solution

(a) Measurement bias. (b) Non-response bias. (c) Sampling bias. (d) Response bias.

2. (Warm-up) Explain in one or two sentences the difference between bias and sampling variability.

Solution

Sampling variability is the random difference between a sample result and the population value, which can go either way and shrinks with larger samples. Bias is a systematic error from the method that pushes results the same way and doesn’t go away with a larger sample.

3. (Warm-up) A hockey league emails a satisfaction survey to all 640640 families, and 160160 reply. Find the response rate, and name the type of bias this could cause.

Solution

160640=0.25\tfrac{160}{640} = 0.25, so the response rate is 25%25\%. This could cause non-response bias: families who reply may be especially happy or especially unhappy with the league.

4. (Core) To find out how often residents use a town’s new bike lanes, a council member surveys people at a bike repair shop. Identify the bias, say whether the results will probably overestimate or underestimate bike lane use, and suggest a better method.

Solution

Sampling bias (a convenience sample). Customers at a bike shop are much more likely to cycle than typical residents, so the survey will probably overestimate bike lane use.

Better: choose a simple random sample of households from the town’s address list and survey them by mail and online, following up with non-responders.

5. (Core) A student weighs 2020 backpacks on a scale that reads 0.30.3 kg too heavy. She finds a mean mass of 6.86.8 kg.

  • (a) What type of bias is this?
  • (b) What is the true mean mass?
  • (c) Would weighing 200200 backpacks on the same scale fix the problem?
Solution

(a) Measurement bias.

(b) Every reading is 0.30.3 kg too high, so the mean is too: 6.8−0.3=6.56.8 - 0.3 = 6.5 kg.

(c) No. Every reading would still be 0.30.3 kg too high. She needs to recalibrate (zero) the scale or correct each reading.

6. (Core) An interviewer wearing a team jersey asks fans outside an arena, “How much do you love our team?” Identify two sources of bias and suggest how to reduce each.

Solution
  • Sampling bias: fans outside the arena are already likely to like the team. Survey a random sample of people from the whole city instead.
  • Response bias (and measurement bias from the wording): the jersey and the question “How much do you love…” push people toward positive answers. Use a neutral interviewer and a neutral question, such as “How would you rate the team’s season, from 1 (poor) to 5 (excellent)?”

7. (Core) A school of 900900 students has 450450 girls and 450450 boys. A random sample of 6060 students is chosen, but only 4040 return the survey: 2828 girls and 1212 boys. What percentage of the responses came from girls? Explain how this could bias a question about which school club should get more funding, and suggest a fix.

Solution

2840=0.7\tfrac{28}{40} = 0.7, so 70%70\% of responses came from girls, compared with 50%50\% of the school.

If girls and boys tend to prefer different clubs, the results will lean toward the girls’ preferences. This is non-response bias. Fixes: follow up with the students who didn’t respond (especially the boys), or use a stratified sample by gender: choose extra boys at random from the list and follow up, so each group’s responses match its share of the school. Even then, the boys who answer may differ from those who don’t.

8. (Challenge) A city wants to estimate the average household income of its residents. It sends a survey to a random sample of households asking, “What was your household income last year?” List three different sources of bias that could affect the result, and say for each whether it would probably push the estimate up or down.

Solution

Sample answer:

  • Non-response bias: households at the very high and very low ends may be less willing to share income, and busy, lower-income households may not have time to respond. The direction depends on who doesn’t respond, so the city should check the response rate in different neighbourhoods.
  • Response bias: some people round up to seem better off, or round down to seem modest, or don’t remember exactly. Rounding up would push the estimate up.
  • Sampling bias: if the sampling frame is a list of property owners, renters (who often have lower incomes) are left out, pushing the estimate up.

Using a complete address list, keeping answers anonymous, and offering income ranges instead of exact amounts all reduce these biases.

9. (Challenge) Two students estimate how many hours per week students at their school spend on homework. Student A asks every person in her Grade 12 calculus class (3030 students). Student B randomly chooses 3030 students from the whole school list, but only 1818 respond. Discuss the bias in each method. Which estimate would you trust more, and why?

Solution

Student A: a convenience sample with sampling bias. Students in Grade 12 calculus likely do more homework than average, so her estimate will probably be too high. She also didn’t sample any other grades.

Student B: a random sample, so no one was favoured in the selection. But the response rate is 1830=60%\tfrac{18}{30} = 60\%, so there could be non-response bias: students who do a lot of homework (or very little) might be less likely to respond.

Student B’s estimate is more trustworthy, since its bias is smaller and can be reduced by following up with the 1212 non-responders. Student A’s method can’t represent the school no matter what.