Skip to content
Family Table Math

Types of Data

Every time you hear “a new study shows…”, someone collected data, organized it, and drew a conclusion from it. Before you can analyze data, you need to know what kind it is, because the type of data decides which graphs, averages, and methods make sense. This page is your vocabulary toolkit for the rest of the statistics units.

A statistical study collects data to answer a question. A few examples of the kinds of questions studies try to answer:

  • Medical research: Does a new treatment help patients recover faster than the old one?
  • Political polling: What percentage of voters in a riding support each candidate?
  • Market research: Would people in a town buy a new flavour of sparkling water, and at what price?

Two studies of the same question can reach different conclusions. That doesn’t always mean someone made a mistake. They may have studied different groups of people, asked the question differently, measured things in different ways, or simply been affected by chance. That’s why it’s important to ask how data was collected before trusting a conclusion.

If you measure the same thing many times, you rarely get identical results. Data varies because of:

  • Measurement variability: instruments and people aren’t perfectly precise. Two people timing the same 100100 m sprint with stopwatches will get slightly different times.
  • Experimental conditions: temperature, time of day, or equipment can change results. A runner might be faster on a dry track than a wet one.
  • Sample differences: different groups of people (or objects) give different results. Two classes in the same school won’t have exactly the same average height.

Some variation is natural and expected. Statistics is all about separating real patterns from this background variation.

Variables, one-variable and two-variable data

Section titled “Variables, one-variable and two-variable data”

A variable is a characteristic that can change from one individual to the next, like height, favourite sport, or number of siblings.

  • One-variable data looks at a single characteristic: “What are the heights of the students in this class?”
  • Two-variable data looks at how two characteristics relate: “Do taller students have longer arm spans?” Two-variable data comes in pairs, one pair per individual.
TypeMeaningExamples
Quantitative (numerical)values are numbers you can do arithmetic withheight, number of siblings, test score
Qualitative (categorical)values are categories or labelsfavourite sport, eye colour, postal code
Discrete (quantitative)countable values, often whole numbersnumber of goals, number of pets
Continuous (quantitative)measured values that can be anything in a rangemass, time, temperature
Nominal (categorical)categories with no natural orderprovince, favourite colour
Ordinal (categorical)categories with a natural orderT-shirt size (S, M, L), “agree / neutral / disagree”

“Categorical” and “qualitative” mean essentially the same thing, as do “numerical” and “quantitative”. Discrete and continuous are the same idea you met for sample spaces.

A number isn’t automatically quantitative. A jersey number or a phone number is a label: averaging jersey numbers makes no sense, so they’re categorical.

  • Primary data is collected by you (or the person doing the study) for this purpose: your own survey or experiment.
  • Secondary data was collected by someone else and you’re reusing it: Statistics Canada tables, a sports league’s published standings, a newspaper’s poll.
  • Experimental data comes from an experiment, where the researcher deliberately changes something (like giving one group a new study app) and measures the effect.
  • Observational data comes from an observational study, where the researcher only watches or records what already happens, without changing anything.
  • Microdata gives the information for each individual: one row per person in a survey.
  • Aggregate data is summarized or grouped: totals, averages, or percentages for a whole group, like “average household size by province”.

A hockey coach records the following for each player on a team. Classify each variable.

  • (a) position (forward, defence, goalie)
  • (b) number of goals this season
  • (c) time on ice in the last game
  • (d) jersey number
  • (e) skill level (beginner, intermediate, advanced)

Solution.

(a) Categorical, nominal. The positions are labels with no natural order.

(b) Quantitative, discrete. Goals are counted in whole numbers.

(c) Quantitative, continuous. Time is measured and could be 1414 min 32.632.6 s, or any value in a range.

(d) Categorical, nominal. It’s a number, but only as a label. Adding or averaging jersey numbers means nothing.

(e) Categorical, ordinal. The levels have a clear order, but the “gap” between beginner and intermediate isn’t a measured amount.

Example 2: Primary or secondary, experimental or observational

Section titled “Example 2: Primary or secondary, experimental or observational”

Classify the data in each situation.

  • (a) A student gives half her volunteers a sugary drink and half water, then times how long each takes to solve a puzzle.
  • (b) A town council uses Statistics Canada census tables to compare the age of its residents with the provincial average.
  • (c) A student stands at the school entrance for a week and records how each student arrives (walk, bike, bus, car).

Solution.

(a) Primary, experimental. She collected it herself, and she deliberately changed something (the drink) to see its effect.

(b) Secondary, observational, aggregate. Someone else collected it, nobody changed anything, and census tables report totals and averages for groups rather than individual people.

(c) Primary, observational. The student collected it, but only watched what students already do.

Two student groups each survey 4040 students about how many hours of sleep they get on school nights. Group A surveys a Grade 9 homeroom at 8:30 a.m. and gets a mean of 8.18.1 hours. Group B surveys students in the library during a Grade 12 exam week and gets a mean of 6.46.4 hours. Give two reasons the results might differ.

Solution.

  • Sample differences: Group A asked only Grade 9 students, and Group B asked mostly Grade 12 students. Younger and older students may have different sleep habits.
  • Conditions: Group B surveyed during exam week, when students may be sleeping less than usual. Students in the library during exams are also probably not typical of the whole school.

Neither group’s result describes “all students at the school”. To do that, they’d need a sample from the whole school at a typical time.

A sports league publishes two files. File 1 lists each player’s age, team, and goals. File 2 lists each team’s total goals and average player age. Which is microdata, and what’s one thing you can do with File 1 that you can’t do with File 2?

Solution. File 1 is microdata (one row per player). File 2 is aggregate data (one row per team).

With File 1 you can find the top scorer, or check whether older players score more goals. File 2 has already combined the players, so the individual values are gone.

Treating every number as quantitative. Jersey numbers, postal codes, student numbers, and phone numbers are labels. Ask: “Does it make sense to average this?” If not, it’s categorical.

Calling ordinal data quantitative. A rating like “1 = poor, 2 = fair, 3 = good” uses numbers, but the categories are ordered labels. The jump from “poor” to “fair” isn’t necessarily the same size as from “fair” to “good”.

Calling something discrete just because it’s rounded. Height recorded to the nearest centimetre is still continuous: the true height can be any value. Discrete means the values themselves are countable, like number of siblings.

Mixing up experimental and observational. The key question is: did the researcher deliberately change something? Surveys and simply recording what happens are observational, even if the researcher is very careful.

Assuming that different results mean someone cheated. Natural variation, different samples, and different conditions all lead to different results. Look at how the data was collected before deciding which study to trust.

1. (Warm-up) Classify each variable as quantitative or categorical.

  • (a) favourite subject
  • (b) number of text messages sent yesterday
  • (c) mass of a backpack
  • (d) area code
Solution

(a) Categorical. (b) Quantitative. (c) Quantitative. (d) Categorical (it’s a label, not an amount).

2. (Warm-up) Classify each quantitative variable as discrete or continuous.

  • (a) the number of cars in a school parking lot
  • (b) the time to download a file
  • (c) the number of pages in a book
  • (d) the volume of water in a bottle
Solution

(a) Discrete. (b) Continuous. (c) Discrete. (d) Continuous.

3. (Warm-up) Classify each categorical variable as nominal or ordinal.

  • (a) blood type
  • (b) coffee size (small, medium, large)
  • (c) level of agreement (strongly disagree to strongly agree)
  • (d) favourite music genre
Solution

(a) Nominal. (b) Ordinal. (c) Ordinal. (d) Nominal.

4. (Core) Is each study one-variable or two-variable? Explain.

  • (a) A student records the commute time of each person in her class.
  • (b) A student records each classmate’s commute time and the distance they live from school.
Solution

(a) One-variable: only commute time is studied.

(b) Two-variable: each student gives a pair (distance, commute time), and the question is how the two are related.

5. (Core) Classify the data in each situation as primary or secondary, and experimental or observational.

  • (a) A gardener grows tomato plants with two different fertilizers and measures the mass of fruit from each.
  • (b) A student downloads a table of Canadian provinces’ populations from a government website.
  • (c) A student surveys classmates about their favourite streaming service.
Solution

(a) Primary, experimental: the gardener chose which plants got which fertilizer.

(b) Secondary, observational: someone else collected it, and nothing was changed.

(c) Primary, observational: a survey records opinions without changing anything.

6. (Core) A school’s student council has a spreadsheet with one row for each student: grade, homeroom, and whether they bought a yearbook. The principal receives a summary showing the percentage of each grade that bought a yearbook. Which is microdata and which is aggregate data? Give one question only the spreadsheet can answer.

Solution

The spreadsheet is microdata; the principal’s summary is aggregate data.

Only the spreadsheet can answer questions about individuals or smaller groups, such as “Which homerooms had the lowest yearbook sales?” or “Did a particular student buy one?”

7. (Core) Give one example of each source of variability for a class experiment measuring how high a tennis ball bounces when dropped from 11 m.

  • (a) measurement variability
  • (b) experimental conditions
  • (c) sample differences
Solution

Sample answers:

(a) Different students read the bounce height off the metre stick slightly differently, especially since the ball is moving.

(b) Dropping the ball onto carpet versus a hard floor, or on a cold day versus a warm one, changes the bounce.

(c) Different tennis balls (old ones versus new ones) bounce to different heights.

8. (Challenge) Two market research firms study whether people in a town would buy a new energy bar. Firm A interviews 200200 shoppers at a fitness centre and finds strong interest. Firm B mails a survey to 200200 randomly chosen households and finds weak interest. Give two reasons the conclusions differ, and say which result you’d trust more for the whole town.

Solution
  • Different samples: fitness-centre shoppers are probably more interested in energy bars than the average resident. Firm B’s random households better represent the town.
  • Different methods: a face-to-face interview may make people say “yes” to be polite, while a mailed survey is anonymous. Also, people who bother to return a mailed survey may differ from those who don’t.

Firm B’s result is more trustworthy for the whole town because its sample was chosen randomly from all households, though it would be worth checking how many surveys were returned.

9. (Challenge) A variable can sometimes be recorded in different ways. For each pair, classify both versions, and say what information is lost in the second.

  • (a) exact age in years versus age group (under 18, 18–64, 65 and over)
  • (b) exact time to finish a 55 km race versus finishing place (1st, 2nd, 3rd, …)
Solution

(a) Exact age is quantitative, continuous (even though we usually round it to whole years). Age group is categorical, ordinal. You lose the detail within each group: a 1919-year-old and a 6060-year-old look the same.

(b) Race time is quantitative, continuous. Finishing place is categorical, ordinal. You lose the gaps between runners: 1st and 2nd might be 0.40.4 s apart or 33 min apart, and the places don’t show which.