Skip to content
Family Table Math
Auto

Sampling Methods

You usually can’t ask everyone. A school can’t easily survey every student about the cafeteria menu, and a polling company can’t phone every voter in Canada. Instead, you collect data from a smaller group, a sample, and use it to draw conclusions about everyone. How you choose that sample decides whether your conclusions can be trusted.

  • The population is the whole group you want to learn about: all 12001200 students at a school, every household in a town.
  • A sample is the part of the population you actually collect data from.
  • A census collects data from the entire population.

A census is often impossible or impractical. Sampling is:

  • cheaper and faster, especially for large populations
  • sometimes the only option, for example when testing destroys the item (you can’t crash-test every car, or taste-test every cookie in a batch)

A good sample is representative: it looks like a small version of the population, so its results are close to what a census would give. That needs:

  • randomness, so the person choosing can’t (even accidentally) favour some individuals
  • enough individuals, because larger samples vary less from the truth (more on this in margin of error)

In each of these, chance decides who is chosen.

  • Simple random sample: every individual, and every group of the same size, has an equal chance of being chosen. Example: put all names in a spreadsheet and pick 5050 with a random number generator.
  • Systematic sample: list the population, find the sampling interval k=Nnk = \dfrac{N}{n} (population size divided by sample size), choose a random start between 11 and kk, then take every kkth individual.
  • Stratified sample: split the population into groups (strata) that share a characteristic, like grade level, then take a simple random sample from each stratum, in proportion to its size:
sample size from a stratum=stratum sizepopulation size×n\text{sample size from a stratum} = \frac{\text{stratum size}}{\text{population size}} \times n
  • Cluster sample: split the population into groups (clusters), like homerooms, randomly choose some whole clusters, and include everyone in those clusters.
  • Multistage sample: sample in stages, for example randomly choose schools in a board, then classes in those schools, then students in those classes.
Three ways to choose a sample of 8 from a population of 40 dots in 4 rows of 10. Simple random: 8 dots scattered with no pattern. Systematic: every 5th dot, starting at dot 3. Stratified: the top row is group A (10 people) and the other rows are group B (30 people); 2 dots are chosen at random from A and 6 from B. Simple random any 8 of the 40 Systematic every 5th, starting at 3 Stratified 2 from group A, 6 from group B A B in the sample not chosen
Choosing 88 from a population of 4040 three ways. For the systematic sample, k=408=5k = \tfrac{40}{8} = 5. For the stratified sample, 1040×8=2\tfrac{10}{40} \times 8 = 2 from group A and 3040×8=6\tfrac{30}{40} \times 8 = 6 from group B.

These are easy, but usually not representative:

  • Convenience sample: choose whoever is easiest to reach, like your friends or the people in your class.
  • Voluntary response sample: people choose themselves, like an online poll or a call-in vote. People with strong opinions are much more likely to respond.

You’ll see exactly how these go wrong in bias in sampling.

Collecting and organizing data in a spreadsheet

Section titled “Collecting and organizing data in a spreadsheet”

Whether you collect primary data (your own survey) or download secondary data (for example a table from Statistics Canada), a spreadsheet keeps it organized:

  • Use one row per individual and one column per variable, with a clear heading in the first row.
  • Record units in the heading, like “Commute time (min)”, and use consistent categories (“Bus”, not sometimes “bus” and sometimes “school bus”).
  • To draw a simple random sample from a list, add a column with =RAND(), sort by that column, and take the first nn rows. Or use =RANDBETWEEN(1,N) to pick ID numbers, skipping repeats.
  • For secondary data, note the source and the date, and read the table’s notes so you know exactly what each column measures.

There’s nothing to calculate here, so Desmos doesn’t help. The SAT tests reasoning: results from a random sample can be generalized only to the population that was sampled. For example, a random sample of students at one high school tells you about that school, not about all teenagers in the province. A sample of volunteers, or of people who are easy to reach, may be biased however large it is. See using Desmos on the SAT.

A town council mails a survey about a new skate park to 300300 randomly chosen households from its list of 85008500 households. Identify the population, the sample, and the sampling method.

Solution.

  • Population: all 85008500 households in the town.
  • Sample: the 300300 households that were mailed the survey.
  • Method: simple random sample, since the 300300 households were chosen at random from the full list, so every possible group of 300300 households was equally likely.

A school has 12001200 students on an alphabetical list. The principal wants a systematic sample of 6060 students. Describe how to choose them.

Solution. Find the sampling interval:

k=Nn=120060=20k = \frac{N}{n} = \frac{1200}{60} = 20

Pick a random number from 11 to 2020, say 77. Then choose every 2020th student starting there:

7, 27, 47, 67, …, 11877, \ 27, \ 47, \ 67, \ \dots, \ 1187

Check: the 6060th student chosen is number 7+59×20=11877 + 59 \times 20 = 1187, which is on the list. ✓

A youth soccer league has 480480 players: 120120 in U10, 160160 in U12, 128128 in U14, and 7272 in U16. The league wants a stratified sample of 6060 players. How many should come from each division?

Solution. The sample is 60480=18\tfrac{60}{480} = \tfrac{1}{8} of the league, so take 18\tfrac{1}{8} of each division:

DivisionPlayersSample
U10120120120480×60=15\tfrac{120}{480} \times 60 = 15
U12160160160480×60=20\tfrac{160}{480} \times 60 = 20
U14128128128480×60=16\tfrac{128}{480} \times 60 = 16
U16727272480×60=9\tfrac{72}{480} \times 60 = 9

Check: 15+20+16+9=6015 + 20 + 16 + 9 = 60. ✓ Then choose a simple random sample of that size within each division.

A school board wants to survey Grade 10 students across its 2525 high schools about online learning. Visiting every school is expensive. Suggest a sampling method and explain how to carry it out.

Solution. A multistage (or cluster) sample works well:

  1. Randomly choose some of the 2525 schools, say 55.
  2. In each chosen school, randomly choose two Grade 10 classes.
  3. Survey every student in those classes (or a random sample of them).

This saves travel, and randomness is used at every stage. One risk: if the chosen schools happen to be similar (all rural, say), the sample may not represent the whole board. Choosing more schools, or stratifying schools by type first, reduces that risk.

Thinking “random” means “haphazard”. Standing in the hallway and picking “whoever” isn’t random: you’ll favour people who look approachable. Random means chance decides, using a random number generator or something equivalent.

Taking equal numbers from each stratum. A stratified sample is usually proportional. If Grade 9 is twice as big as Grade 12, it should get twice as many people in the sample.

Forgetting the random start in a systematic sample. If you always start at the first name, the sample isn’t random. Choose the start randomly from 11 to kk.

Mixing up stratified and cluster samples. In a stratified sample, you take some individuals from every group. In a cluster sample, you take every individual from some groups.

Believing a big sample fixes a bad method. An online poll with 10 00010\,000 votes is still a voluntary response sample. A smaller random sample is usually more trustworthy.

1. (Warm-up) Identify the population and the sample.

  • (a) A farmer tests 2020 apples from a shipment of 50005000 for bruising.
  • (b) A student asks 3030 randomly chosen students in her school of 900900 about their favourite lunch.
Solution

(a) Population: the 50005000 apples in the shipment. Sample: the 2020 tested apples.

(b) Population: the 900900 students at the school. Sample: the 3030 students asked.

2. (Warm-up) Name the sampling method.

  • (a) Every 1010th customer leaving a grocery store is asked to complete a survey.
  • (b) A radio station asks listeners to text in their vote.
  • (c) A teacher surveys the students in her own two classes.
  • (d) Student numbers are entered into a random number generator, which picks 4040 of them.
Solution

(a) Systematic. (b) Voluntary response. (c) Convenience. (d) Simple random.

3. (Warm-up) Give two reasons a company that makes light bulbs would test a sample rather than every bulb.

Solution

Testing how long a bulb lasts destroys it (it burns out), so testing every bulb would leave nothing to sell. Testing a sample is also much cheaper and faster.

4. (Core) A sports club has 850850 members on a numbered list. Describe how to choose a systematic sample of 5050 members. If the random start is 44, list the first four members chosen and the last one.

Solution

k=85050=17k = \tfrac{850}{50} = 17. Choose a random start from 11 to 1717, then take every 1717th member.

With a start of 44: 4,21,38,554, 21, 38, 55, and the last is 4+49×17=8374 + 49 \times 17 = 837.

5. (Core) A high school has 240240 students in Grade 9, 220220 in Grade 10, 180180 in Grade 11, and 160160 in Grade 12. Find the number from each grade in a stratified sample of 4040 students.

Solution

The school has 800800 students, so the sample is 40800=120\tfrac{40}{800} = \tfrac{1}{20} of each grade:

Grade 9: 1212. Grade 10: 1111. Grade 11: 99. Grade 12: 88.

Check: 12+11+9+8=4012 + 11 + 9 + 8 = 40. ✓

6. (Core) Explain the difference between a stratified sample and a cluster sample, using homerooms in a school as the example.

Solution

Stratified: treat each homeroom as a stratum and randomly choose a few students from every homeroom (in proportion to its size).

Cluster: randomly choose a few homerooms and survey every student in those homerooms.

Stratified samples guarantee every group is represented. Cluster samples are easier to carry out, but if the chosen homerooms aren’t typical, the results can be off.

7. (Core) A student wants to know how many hours per week students at his school of 11001100 spend on part-time jobs. He has a spreadsheet of all student names. Describe how to use the spreadsheet to choose a simple random sample of 5050, and how he should organize the data once it’s collected.

Solution

Sample answer: add a column with =RAND() beside the names, sort the whole list by that column, and take the first 5050 names. (Re-sorting gives a different sample, so he should pick once and save it.)

For the data, use one row per student and columns such as “Student ID”, “Grade”, “Has a job (Y/N)”, and “Hours worked last week (h)”. Keep categories consistent and leave a clear note for anyone who didn’t respond, rather than entering 00.

8. (Challenge) A school has 318318 students in Grade 9, 301301 in Grade 10, 290290 in Grade 11, and 291291 in Grade 12. The student council wants a stratified sample of 8080.

  • (a) Find the exact (unrounded) number for each grade.
  • (b) Rounding each to the nearest whole number gives the wrong total. Show this, and suggest a fair fix.
Solution

(a) The school has 12001200 students. Each grade gets grade size1200×80\tfrac{\text{grade size}}{1200} \times 80:

Grade 9: 21.221.2. Grade 10: about 20.0720.07. Grade 11: about 19.3319.33. Grade 12: 19.419.4.

(b) Rounding gives 21+20+19+19=7921 + 20 + 19 + 19 = 79, one short of 8080.

A fair fix is to give the extra person to the grade whose number was rounded down the most. Grade 12’s 19.419.4 lost 0.40.4, the most of any grade, so use 21,20,19,2021, 20, 19, 20. Check: 21+20+19+20=8021 + 20 + 19 + 20 = 80. ✓

9. (Challenge) A list has 10001000 names and you want a systematic sample of 3030.

  • (a) Why is k=Nnk = \tfrac{N}{n} a problem here?
  • (b) If you use k=33k = 33 and a random start from 11 to 3333, which names can never be chosen?
  • (c) Suggest a way to fix this.
Solution

(a) 100030≈33.3\tfrac{1000}{30} \approx 33.3 isn’t a whole number, so you can’t take “every 33.333.3th name”.

(b) With start ss, the 3030th name is number s+29×33=s+957s + 29 \times 33 = s + 957, which is at most 33+957=99033 + 957 = 990. So names 991991 to 10001000 can never be chosen, and those people have no chance of being in the sample.

(c) One fix: choose the random start from 11 to 10001000, and when you pass the end of the list, wrap around to the beginning (treat the list as a circle). Now every name has the same chance of being chosen.