Sampling Methods
You usually can’t ask everyone. A school can’t easily survey every student about the cafeteria menu, and a polling company can’t phone every voter in Canada. Instead, you collect data from a smaller group, a sample, and use it to draw conclusions about everyone. How you choose that sample decides whether your conclusions can be trusted.
Key ideas
Section titled “Key ideas”Population and sample
Section titled “Population and sample”- The population is the whole group you want to learn about: all students at a school, every household in a town.
- A sample is the part of the population you actually collect data from.
- A census collects data from the entire population.
Why sample?
Section titled “Why sample?”A census is often impossible or impractical. Sampling is:
- cheaper and faster, especially for large populations
- sometimes the only option, for example when testing destroys the item (you can’t crash-test every car, or taste-test every cookie in a batch)
What makes a good sample?
Section titled “What makes a good sample?”A good sample is representative: it looks like a small version of the population, so its results are close to what a census would give. That needs:
- randomness, so the person choosing can’t (even accidentally) favour some individuals
- enough individuals, because larger samples vary less from the truth (more on this in margin of error)
Random sampling methods
Section titled “Random sampling methods”In each of these, chance decides who is chosen.
- Simple random sample: every individual, and every group of the same size, has an equal chance of being chosen. Example: put all names in a spreadsheet and pick with a random number generator.
- Systematic sample: list the population, find the sampling interval (population size divided by sample size), choose a random start between and , then take every th individual.
- Stratified sample: split the population into groups (strata) that share a characteristic, like grade level, then take a simple random sample from each stratum, in proportion to its size:
- Cluster sample: split the population into groups (clusters), like homerooms, randomly choose some whole clusters, and include everyone in those clusters.
- Multistage sample: sample in stages, for example randomly choose schools in a board, then classes in those schools, then students in those classes.
Non-random sampling methods
Section titled “Non-random sampling methods”These are easy, but usually not representative:
- Convenience sample: choose whoever is easiest to reach, like your friends or the people in your class.
- Voluntary response sample: people choose themselves, like an online poll or a call-in vote. People with strong opinions are much more likely to respond.
You’ll see exactly how these go wrong in bias in sampling.
Collecting and organizing data in a spreadsheet
Section titled “Collecting and organizing data in a spreadsheet”Whether you collect primary data (your own survey) or download secondary data (for example a table from Statistics Canada), a spreadsheet keeps it organized:
- Use one row per individual and one column per variable, with a clear heading in the first row.
- Record units in the heading, like “Commute time (min)”, and use consistent categories (“Bus”, not sometimes “bus” and sometimes “school bus”).
- To draw a simple random sample from a list, add a column with
=RAND(), sort by that column, and take the first rows. Or use=RANDBETWEEN(1,N)to pick ID numbers, skipping repeats. - For secondary data, note the source and the date, and read the table’s notes so you know exactly what each column measures.
Worked examples
Section titled “Worked examples”Example 1: Population, sample, and method
Section titled “Example 1: Population, sample, and method”A town council mails a survey about a new skate park to randomly chosen households from its list of households. Identify the population, the sample, and the sampling method.
Solution.
- Population: all households in the town.
- Sample: the households that were mailed the survey.
- Method: simple random sample, since the households were chosen at random from the full list, so every possible group of households was equally likely.
Example 2: A systematic sample
Section titled “Example 2: A systematic sample”A school has students on an alphabetical list. The principal wants a systematic sample of students. Describe how to choose them.
Solution. Find the sampling interval:
Pick a random number from to , say . Then choose every th student starting there:
Check: the th student chosen is number , which is on the list. ✓
Example 3: A stratified sample
Section titled “Example 3: A stratified sample”A youth soccer league has players: in U10, in U12, in U14, and in U16. The league wants a stratified sample of players. How many should come from each division?
Solution. The sample is of the league, so take of each division:
| Division | Players | Sample |
|---|---|---|
| U10 | ||
| U12 | ||
| U14 | ||
| U16 |
Check: . ✓ Then choose a simple random sample of that size within each division.
Example 4: Choosing a method
Section titled “Example 4: Choosing a method”A school board wants to survey Grade 10 students across its high schools about online learning. Visiting every school is expensive. Suggest a sampling method and explain how to carry it out.
Solution. A multistage (or cluster) sample works well:
- Randomly choose some of the schools, say .
- In each chosen school, randomly choose two Grade 10 classes.
- Survey every student in those classes (or a random sample of them).
This saves travel, and randomness is used at every stage. One risk: if the chosen schools happen to be similar (all rural, say), the sample may not represent the whole board. Choosing more schools, or stratifying schools by type first, reduces that risk.
Common mistakes
Section titled “Common mistakes”Thinking “random” means “haphazard”. Standing in the hallway and picking “whoever” isn’t random: you’ll favour people who look approachable. Random means chance decides, using a random number generator or something equivalent.
Taking equal numbers from each stratum. A stratified sample is usually proportional. If Grade 9 is twice as big as Grade 12, it should get twice as many people in the sample.
Forgetting the random start in a systematic sample. If you always start at the first name, the sample isn’t random. Choose the start randomly from to .
Mixing up stratified and cluster samples. In a stratified sample, you take some individuals from every group. In a cluster sample, you take every individual from some groups.
Believing a big sample fixes a bad method. An online poll with votes is still a voluntary response sample. A smaller random sample is usually more trustworthy.
Practice
Section titled “Practice”1. (Warm-up) Identify the population and the sample.
- (a) A farmer tests apples from a shipment of for bruising.
- (b) A student asks randomly chosen students in her school of about their favourite lunch.
Solution
(a) Population: the apples in the shipment. Sample: the tested apples.
(b) Population: the students at the school. Sample: the students asked.
2. (Warm-up) Name the sampling method.
- (a) Every th customer leaving a grocery store is asked to complete a survey.
- (b) A radio station asks listeners to text in their vote.
- (c) A teacher surveys the students in her own two classes.
- (d) Student numbers are entered into a random number generator, which picks of them.
Solution
(a) Systematic. (b) Voluntary response. (c) Convenience. (d) Simple random.
3. (Warm-up) Give two reasons a company that makes light bulbs would test a sample rather than every bulb.
Solution
Testing how long a bulb lasts destroys it (it burns out), so testing every bulb would leave nothing to sell. Testing a sample is also much cheaper and faster.
4. (Core) A sports club has members on a numbered list. Describe how to choose a systematic sample of members. If the random start is , list the first four members chosen and the last one.
Solution
. Choose a random start from to , then take every th member.
With a start of : , and the last is .
5. (Core) A high school has students in Grade 9, in Grade 10, in Grade 11, and in Grade 12. Find the number from each grade in a stratified sample of students.
Solution
The school has students, so the sample is of each grade:
Grade 9: . Grade 10: . Grade 11: . Grade 12: .
Check: . ✓
6. (Core) Explain the difference between a stratified sample and a cluster sample, using homerooms in a school as the example.
Solution
Stratified: treat each homeroom as a stratum and randomly choose a few students from every homeroom (in proportion to its size).
Cluster: randomly choose a few homerooms and survey every student in those homerooms.
Stratified samples guarantee every group is represented. Cluster samples are easier to carry out, but if the chosen homerooms aren’t typical, the results can be off.
7. (Core) A student wants to know how many hours per week students at his school of spend on part-time jobs. He has a spreadsheet of all student names. Describe how to use the spreadsheet to choose a simple random sample of , and how he should organize the data once it’s collected.
Solution
Sample answer: add a column with =RAND() beside the names, sort the whole list by that column, and take the first names. (Re-sorting gives a different sample, so he should pick once and save it.)
For the data, use one row per student and columns such as “Student ID”, “Grade”, “Has a job (Y/N)”, and “Hours worked last week (h)”. Keep categories consistent and leave a clear note for anyone who didn’t respond, rather than entering .
8. (Challenge) A school has students in Grade 9, in Grade 10, in Grade 11, and in Grade 12. The student council wants a stratified sample of .
- (a) Find the exact (unrounded) number for each grade.
- (b) Rounding each to the nearest whole number gives the wrong total. Show this, and suggest a fair fix.
Solution
(a) The school has students. Each grade gets :
Grade 9: . Grade 10: about . Grade 11: about . Grade 12: .
(b) Rounding gives , one short of .
A fair fix is to give the extra person to the grade whose number was rounded down the most. Grade 12’s lost , the most of any grade, so use . Check: . ✓
9. (Challenge) A list has names and you want a systematic sample of .
- (a) Why is a problem here?
- (b) If you use and a random start from to , which names can never be chosen?
- (c) Suggest a way to fix this.
Solution
(a) isn’t a whole number, so you can’t take “every th name”.
(b) With start , the th name is number , which is at most . So names to can never be chosen, and those people have no chance of being in the sample.
(c) One fix: choose the random start from to , and when you pass the end of the list, wrap around to the beginning (treat the list as a circle). Now every name has the same chance of being chosen.