Types of Data
Every time you hear “a new study shows…”, someone collected data, organized it, and drew a conclusion from it. Before you can analyze data, you need to know what kind it is, because the type of data decides which graphs, averages, and methods make sense. This page is your vocabulary toolkit for the rest of the statistics units.
Key ideas
Section titled “Key ideas”The role of data in statistical studies
Section titled “The role of data in statistical studies”A statistical study collects data to answer a question. A few examples of the kinds of questions studies try to answer:
- Medical research: Does a new treatment help patients recover faster than the old one?
- Political polling: What percentage of voters in a riding support each candidate?
- Market research: Would people in a town buy a new flavour of sparkling water, and at what price?
Two studies of the same question can reach different conclusions. That doesn’t always mean someone made a mistake. They may have studied different groups of people, asked the question differently, measured things in different ways, or simply been affected by chance. That’s why it’s important to ask how data was collected before trusting a conclusion.
Why data varies
Section titled “Why data varies”If you measure the same thing many times, you rarely get identical results. Data varies because of:
- Measurement variability: instruments and people aren’t perfectly precise. Two people timing the same m sprint with stopwatches will get slightly different times.
- Experimental conditions: temperature, time of day, or equipment can change results. A runner might be faster on a dry track than a wet one.
- Sample differences: different groups of people (or objects) give different results. Two classes in the same school won’t have exactly the same average height.
Some variation is natural and expected. Statistics is all about separating real patterns from this background variation.
Variables, one-variable and two-variable data
Section titled “Variables, one-variable and two-variable data”A variable is a characteristic that can change from one individual to the next, like height, favourite sport, or number of siblings.
- One-variable data looks at a single characteristic: “What are the heights of the students in this class?”
- Two-variable data looks at how two characteristics relate: “Do taller students have longer arm spans?” Two-variable data comes in pairs, one pair per individual.
Classifying variables
Section titled “Classifying variables”| Type | Meaning | Examples |
|---|---|---|
| Quantitative (numerical) | values are numbers you can do arithmetic with | height, number of siblings, test score |
| Qualitative (categorical) | values are categories or labels | favourite sport, eye colour, postal code |
| Discrete (quantitative) | countable values, often whole numbers | number of goals, number of pets |
| Continuous (quantitative) | measured values that can be anything in a range | mass, time, temperature |
| Nominal (categorical) | categories with no natural order | province, favourite colour |
| Ordinal (categorical) | categories with a natural order | T-shirt size (S, M, L), “agree / neutral / disagree” |
“Categorical” and “qualitative” mean essentially the same thing, as do “numerical” and “quantitative”. Discrete and continuous are the same idea you met for sample spaces.
A number isn’t automatically quantitative. A jersey number or a phone number is a label: averaging jersey numbers makes no sense, so they’re categorical.
Classifying data sets
Section titled “Classifying data sets”- Primary data is collected by you (or the person doing the study) for this purpose: your own survey or experiment.
- Secondary data was collected by someone else and you’re reusing it: Statistics Canada tables, a sports league’s published standings, a newspaper’s poll.
- Experimental data comes from an experiment, where the researcher deliberately changes something (like giving one group a new study app) and measures the effect.
- Observational data comes from an observational study, where the researcher only watches or records what already happens, without changing anything.
- Microdata gives the information for each individual: one row per person in a survey.
- Aggregate data is summarized or grouped: totals, averages, or percentages for a whole group, like “average household size by province”.
Worked examples
Section titled “Worked examples”Example 1: Classifying variables
Section titled “Example 1: Classifying variables”A hockey coach records the following for each player on a team. Classify each variable.
- (a) position (forward, defence, goalie)
- (b) number of goals this season
- (c) time on ice in the last game
- (d) jersey number
- (e) skill level (beginner, intermediate, advanced)
Solution.
(a) Categorical, nominal. The positions are labels with no natural order.
(b) Quantitative, discrete. Goals are counted in whole numbers.
(c) Quantitative, continuous. Time is measured and could be min s, or any value in a range.
(d) Categorical, nominal. It’s a number, but only as a label. Adding or averaging jersey numbers means nothing.
(e) Categorical, ordinal. The levels have a clear order, but the “gap” between beginner and intermediate isn’t a measured amount.
Example 2: Primary or secondary, experimental or observational
Section titled “Example 2: Primary or secondary, experimental or observational”Classify the data in each situation.
- (a) A student gives half her volunteers a sugary drink and half water, then times how long each takes to solve a puzzle.
- (b) A town council uses Statistics Canada census tables to compare the age of its residents with the provincial average.
- (c) A student stands at the school entrance for a week and records how each student arrives (walk, bike, bus, car).
Solution.
(a) Primary, experimental. She collected it herself, and she deliberately changed something (the drink) to see its effect.
(b) Secondary, observational, aggregate. Someone else collected it, nobody changed anything, and census tables report totals and averages for groups rather than individual people.
(c) Primary, observational. The student collected it, but only watched what students already do.
Example 3: Why studies disagree
Section titled “Example 3: Why studies disagree”Two student groups each survey students about how many hours of sleep they get on school nights. Group A surveys a Grade 9 homeroom at 8:30 a.m. and gets a mean of hours. Group B surveys students in the library during a Grade 12 exam week and gets a mean of hours. Give two reasons the results might differ.
Solution.
- Sample differences: Group A asked only Grade 9 students, and Group B asked mostly Grade 12 students. Younger and older students may have different sleep habits.
- Conditions: Group B surveyed during exam week, when students may be sleeping less than usual. Students in the library during exams are also probably not typical of the whole school.
Neither group’s result describes “all students at the school”. To do that, they’d need a sample from the whole school at a typical time.
Example 4: Microdata or aggregate data
Section titled “Example 4: Microdata or aggregate data”A sports league publishes two files. File 1 lists each player’s age, team, and goals. File 2 lists each team’s total goals and average player age. Which is microdata, and what’s one thing you can do with File 1 that you can’t do with File 2?
Solution. File 1 is microdata (one row per player). File 2 is aggregate data (one row per team).
With File 1 you can find the top scorer, or check whether older players score more goals. File 2 has already combined the players, so the individual values are gone.
Common mistakes
Section titled “Common mistakes”Treating every number as quantitative. Jersey numbers, postal codes, student numbers, and phone numbers are labels. Ask: “Does it make sense to average this?” If not, it’s categorical.
Calling ordinal data quantitative. A rating like “1 = poor, 2 = fair, 3 = good” uses numbers, but the categories are ordered labels. The jump from “poor” to “fair” isn’t necessarily the same size as from “fair” to “good”.
Calling something discrete just because it’s rounded. Height recorded to the nearest centimetre is still continuous: the true height can be any value. Discrete means the values themselves are countable, like number of siblings.
Mixing up experimental and observational. The key question is: did the researcher deliberately change something? Surveys and simply recording what happens are observational, even if the researcher is very careful.
Assuming that different results mean someone cheated. Natural variation, different samples, and different conditions all lead to different results. Look at how the data was collected before deciding which study to trust.
Practice
Section titled “Practice”1. (Warm-up) Classify each variable as quantitative or categorical.
- (a) favourite subject
- (b) number of text messages sent yesterday
- (c) mass of a backpack
- (d) area code
Solution
(a) Categorical. (b) Quantitative. (c) Quantitative. (d) Categorical (it’s a label, not an amount).
2. (Warm-up) Classify each quantitative variable as discrete or continuous.
- (a) the number of cars in a school parking lot
- (b) the time to download a file
- (c) the number of pages in a book
- (d) the volume of water in a bottle
Solution
(a) Discrete. (b) Continuous. (c) Discrete. (d) Continuous.
3. (Warm-up) Classify each categorical variable as nominal or ordinal.
- (a) blood type
- (b) coffee size (small, medium, large)
- (c) level of agreement (strongly disagree to strongly agree)
- (d) favourite music genre
Solution
(a) Nominal. (b) Ordinal. (c) Ordinal. (d) Nominal.
4. (Core) Is each study one-variable or two-variable? Explain.
- (a) A student records the commute time of each person in her class.
- (b) A student records each classmate’s commute time and the distance they live from school.
Solution
(a) One-variable: only commute time is studied.
(b) Two-variable: each student gives a pair (distance, commute time), and the question is how the two are related.
5. (Core) Classify the data in each situation as primary or secondary, and experimental or observational.
- (a) A gardener grows tomato plants with two different fertilizers and measures the mass of fruit from each.
- (b) A student downloads a table of Canadian provinces’ populations from a government website.
- (c) A student surveys classmates about their favourite streaming service.
Solution
(a) Primary, experimental: the gardener chose which plants got which fertilizer.
(b) Secondary, observational: someone else collected it, and nothing was changed.
(c) Primary, observational: a survey records opinions without changing anything.
6. (Core) A school’s student council has a spreadsheet with one row for each student: grade, homeroom, and whether they bought a yearbook. The principal receives a summary showing the percentage of each grade that bought a yearbook. Which is microdata and which is aggregate data? Give one question only the spreadsheet can answer.
Solution
The spreadsheet is microdata; the principal’s summary is aggregate data.
Only the spreadsheet can answer questions about individuals or smaller groups, such as “Which homerooms had the lowest yearbook sales?” or “Did a particular student buy one?”
7. (Core) Give one example of each source of variability for a class experiment measuring how high a tennis ball bounces when dropped from m.
- (a) measurement variability
- (b) experimental conditions
- (c) sample differences
Solution
Sample answers:
(a) Different students read the bounce height off the metre stick slightly differently, especially since the ball is moving.
(b) Dropping the ball onto carpet versus a hard floor, or on a cold day versus a warm one, changes the bounce.
(c) Different tennis balls (old ones versus new ones) bounce to different heights.
8. (Challenge) Two market research firms study whether people in a town would buy a new energy bar. Firm A interviews shoppers at a fitness centre and finds strong interest. Firm B mails a survey to randomly chosen households and finds weak interest. Give two reasons the conclusions differ, and say which result you’d trust more for the whole town.
Solution
- Different samples: fitness-centre shoppers are probably more interested in energy bars than the average resident. Firm B’s random households better represent the town.
- Different methods: a face-to-face interview may make people say “yes” to be polite, while a mailed survey is anonymous. Also, people who bother to return a mailed survey may differ from those who don’t.
Firm B’s result is more trustworthy for the whole town because its sample was chosen randomly from all households, though it would be worth checking how many surveys were returned.
9. (Challenge) A variable can sometimes be recorded in different ways. For each pair, classify both versions, and say what information is lost in the second.
- (a) exact age in years versus age group (under 18, 18–64, 65 and over)
- (b) exact time to finish a km race versus finishing place (1st, 2nd, 3rd, …)
Solution
(a) Exact age is quantitative, continuous (even though we usually round it to whole years). Age group is categorical, ordinal. You lose the detail within each group: a -year-old and a -year-old look the same.
(b) Race time is quantitative, continuous. Finishing place is categorical, ordinal. You lose the gaps between runners: 1st and 2nd might be s apart or min apart, and the places don’t show which.