Bias in Sampling
A sample is supposed to be a small, fair picture of a population. Bias is anything in the way data is collected that pushes the results consistently in one direction, so the picture is distorted. A biased study can collect thousands of responses and still give the wrong answer, so spotting and reducing bias is one of the most useful skills in statistics.
Key ideas
Section titled “Key ideas”What bias is (and isn’t)
Section titled “What bias is (and isn’t)”Results from any sample will differ a little from the true population value just by chance. That’s sampling variability, and it goes both ways: sometimes a bit high, sometimes a bit low. Larger samples reduce it.
Bias is different: it’s a systematic error that tends to push results the same way every time. Making the sample bigger doesn’t fix it. You have to fix the method.
Four kinds of bias
Section titled “Four kinds of bias”| Type | What happens | Example |
|---|---|---|
| Sampling bias | the sampling method favours some parts of the population, or leaves some out | surveying students about bus service by asking only students in the parking lot |
| Non-response bias | some chosen individuals don’t respond, and they differ from those who do | only of mailed surveys come back, mostly from people with strong opinions |
| Response bias | people give false or misleading answers, on purpose or not | students under-report how much time they spend on their phones when the teacher is watching |
| Measurement bias | the way data is measured consistently overestimates or underestimates | a bathroom scale that always reads kg too high; a question worded to push one answer |
Biased random samples
Section titled “Biased random samples”Even a random sample can be biased if:
- the list you sample from (the sampling frame) leaves people out, like a phone survey that only calls landlines and misses households with only cell phones
- many chosen people don’t respond
- the questions or measurements are biased
Randomness protects against the person choosing the sample playing favourites. It doesn’t protect against a bad list, a low response rate, or a leading question.
Non-random samples
Section titled “Non-random samples”Convenience and voluntary response samples are almost always biased:
- A convenience sample only reaches people who are easy to reach, and they’re often similar to each other (and to the person asking).
- A voluntary response sample over-represents people with strong feelings, especially negative ones. Online polls and call-in votes are classic examples.
Reducing bias
Section titled “Reducing bias”| Bias | How to reduce it |
|---|---|
| Sampling bias | use a random method; make sure the sampling frame includes the whole population; stratify if some groups might be missed |
| Non-response bias | follow up with non-responders; keep surveys short; offer several ways to respond; report the response rate |
| Response bias | make responses anonymous; use neutral interviewers; avoid questions that embarrass people |
| Measurement bias | calibrate instruments; use neutral question wording; test the survey on a small group first |
You’ll learn to write fair questions in survey and experiment design.
Worked examples
Section titled “Worked examples”Example 1: Naming the bias
Section titled “Example 1: Naming the bias”Identify the type of bias in each situation.
- (a) To estimate how many students play sports, a student surveys people in the school gym after practice.
- (b) A teacher asks her class, out loud, “Who did all of last night’s homework?” and counts the hands.
- (c) A town mails a survey to random households, and only are returned.
- (d) Students measure their heights against a wall chart that starts cm above the floor.
Solution.
(a) Sampling bias. Students in the gym after practice are far more likely to play sports than the average student.
(b) Response bias. Students who didn’t do the homework may raise their hands anyway rather than admit it in front of the teacher and class.
(c) Non-response bias. The response rate is , or . The households that replied may care more about the issue than the that didn’t.
(d) Measurement bias. Every height is read from a chart that is cm off, so every measurement is cm too low. (A height of cm reads as cm.)
Example 2: How bias changes the result
Section titled “Example 2: How bias changes the result”A school has students, and of them are in Grade 12. A student surveys people in the senior lounge about whether the school should start later in the morning. of the are in Grade 12. Explain why the results may be biased.
Solution. In the whole school, Grade 12 students make up . In the sample, they make up . Grade 12 is heavily over-represented, and other grades are under-represented.
If Grade 12 students feel differently about start times (for example, because more of them have jobs or drive), the survey’s result won’t reflect the whole school. This is a convenience sample with sampling bias.
A fix: use a stratified random sample from the full student list, with about of the sample from Grade 12.
Example 3: A random sample that’s still biased
Section titled “Example 3: A random sample that’s still biased”A polling firm picks phone numbers randomly from an old directory of landline numbers to survey a city’s residents about public transit. Why might the results be biased, even though the sample is random?
Solution. The sampling frame is incomplete. Many residents, especially younger people and renters, don’t have a listed landline, so they have no chance of being chosen. Younger people may also use public transit differently, so leaving them out can push the results one way. This is sampling bias.
A fix: use a frame that includes cell phones (random digit dialling) or combine phone, online, and mail methods, and check that the sample’s ages match the city’s.
Example 4: Spotting the non-random sample
Section titled “Example 4: Spotting the non-random sample”A website asks, “Should our city ban gas-powered leaf blowers?” and gets votes: say yes. A councillor says, “With votes, this must reflect the city.” Is she right?
Solution. No. This is a voluntary response sample: only people who visit the site and care enough to vote are counted. People who are annoyed by leaf blowers are probably more motivated to vote than people who don’t mind them. One person may also be able to vote many times.
A large sample doesn’t remove bias. A random sample of a few hundred residents would give a more trustworthy picture.
Common mistakes
Section titled “Common mistakes”Thinking a larger sample removes bias. Size reduces random variability, but a biased method stays biased no matter how many people respond. Fix the method first.
Assuming a random sample is automatically unbiased. Random selection from an incomplete list, low response rates, and leading questions can all bias a random sample.
Mixing up response bias and non-response bias. Response bias is about the answers people give (they’re inaccurate). Non-response bias is about people who don’t answer at all.
Mixing up response bias and measurement bias. If the tool is off (a broken scale, a loaded question), it’s measurement bias. If the person gives an inaccurate answer (embarrassment, wanting to please), it’s response bias. Some situations involve both.
Naming the bias without explaining its direction. Good answers say which way the results are pushed: “This survey will probably overestimate the number of students who play sports, because…”
Practice
Section titled “Practice”1. (Warm-up) Match each situation with sampling, non-response, response, or measurement bias.
- (a) A thermometer always reads too high.
- (b) Only half the students given a survey hand it back.
- (c) A survey about reading habits is done only at the public library.
- (d) Students exaggerate how much they exercise because their gym teacher is asking.
Solution
(a) Measurement bias. (b) Non-response bias. (c) Sampling bias. (d) Response bias.
2. (Warm-up) Explain in one or two sentences the difference between bias and sampling variability.
Solution
Sampling variability is the random difference between a sample result and the population value, which can go either way and shrinks with larger samples. Bias is a systematic error from the method that pushes results the same way and doesn’t go away with a larger sample.
3. (Warm-up) A hockey league emails a satisfaction survey to all families, and reply. Find the response rate, and name the type of bias this could cause.
Solution
, so the response rate is . This could cause non-response bias: families who reply may be especially happy or especially unhappy with the league.
4. (Core) To find out how often residents use a town’s new bike lanes, a council member surveys people at a bike repair shop. Identify the bias, say whether the results will probably overestimate or underestimate bike lane use, and suggest a better method.
Solution
Sampling bias (a convenience sample). Customers at a bike shop are much more likely to cycle than typical residents, so the survey will probably overestimate bike lane use.
Better: choose a simple random sample of households from the town’s address list and survey them by mail and online, following up with non-responders.
5. (Core) A student weighs backpacks on a scale that reads kg too heavy. She finds a mean mass of kg.
- (a) What type of bias is this?
- (b) What is the true mean mass?
- (c) Would weighing backpacks on the same scale fix the problem?
Solution
(a) Measurement bias.
(b) Every reading is kg too high, so the mean is too: kg.
(c) No. Every reading would still be kg too high. She needs to recalibrate (zero) the scale or correct each reading.
6. (Core) An interviewer wearing a team jersey asks fans outside an arena, “How much do you love our team?” Identify two sources of bias and suggest how to reduce each.
Solution
- Sampling bias: fans outside the arena are already likely to like the team. Survey a random sample of people from the whole city instead.
- Response bias (and measurement bias from the wording): the jersey and the question “How much do you love…” push people toward positive answers. Use a neutral interviewer and a neutral question, such as “How would you rate the team’s season, from 1 (poor) to 5 (excellent)?”
7. (Core) A school of students has girls and boys. A random sample of students is chosen, but only return the survey: girls and boys. What percentage of the responses came from girls? Explain how this could bias a question about which school club should get more funding, and suggest a fix.
Solution
, so of responses came from girls, compared with of the school.
If girls and boys tend to prefer different clubs, the results will lean toward the girls’ preferences. This is non-response bias. Fixes: follow up with the students who didn’t respond (especially the boys), or use a stratified sample by gender: choose extra boys at random from the list and follow up, so each group’s responses match its share of the school. Even then, the boys who answer may differ from those who don’t.
8. (Challenge) A city wants to estimate the average household income of its residents. It sends a survey to a random sample of households asking, “What was your household income last year?” List three different sources of bias that could affect the result, and say for each whether it would probably push the estimate up or down.
Solution
Sample answer:
- Non-response bias: households at the very high and very low ends may be less willing to share income, and busy, lower-income households may not have time to respond. The direction depends on who doesn’t respond, so the city should check the response rate in different neighbourhoods.
- Response bias: some people round up to seem better off, or round down to seem modest, or don’t remember exactly. Rounding up would push the estimate up.
- Sampling bias: if the sampling frame is a list of property owners, renters (who often have lower incomes) are left out, pushing the estimate up.
Using a complete address list, keeping answers anonymous, and offering income ranges instead of exact amounts all reduce these biases.
9. (Challenge) Two students estimate how many hours per week students at their school spend on homework. Student A asks every person in her Grade 12 calculus class ( students). Student B randomly chooses students from the whole school list, but only respond. Discuss the bias in each method. Which estimate would you trust more, and why?
Solution
Student A: a convenience sample with sampling bias. Students in Grade 12 calculus likely do more homework than average, so her estimate will probably be too high. She also didn’t sample any other grades.
Student B: a random sample, so no one was favoured in the selection. But the response rate is , so there could be non-response bias: students who do a lot of homework (or very little) might be less likely to respond.
Student B’s estimate is more trustworthy, since its bias is smaller and can be reduced by following up with the non-responders. Student A’s method can’t represent the school no matter what.