Continuous Random Variables
Some quantities are counted (the number of goals in a game), but many are measured: heights, masses, times, temperatures. A measured quantity can take any value in a range, so you can’t list its outcomes one by one the way you did for discrete random variables. Instead, you group the data into intervals, draw a histogram, and think of probability as area. This idea is the foundation for the normal distribution.
Key ideas
Section titled “Key ideas”Discrete vs continuous random variables
Section titled “Discrete vs continuous random variables”A random variable gives a number to each outcome of an experiment.
| Discrete | Continuous | |
|---|---|---|
| Values | separate values you can count | any value in an interval |
| How you get them | counting | measuring |
| Examples | number of heads, number of absent students | height in cm, mass in g, time in minutes |
| Graph | bar graph (gaps between bars) | histogram (bars touch) |
A good test: could there be a value between any two possible values? Between cm and cm there’s cm, cm, and so on. That makes height continuous.
Why continuous distributions need models
Section titled “Why continuous distributions need models”You can never pin down a continuous distribution exactly:
- Sampling. You can’t measure every apple in Ontario, so you only ever see a sample, and a different sample gives a slightly different picture.
- Measurement uncertainty. Every measurement is rounded. A mass recorded as g really means “somewhere from g to g”. No ruler or scale gives the exact value.
So statisticians use a smooth model (a curve, like the normal curve) that describes the data well, rather than trying to know the true distribution perfectly.
Frequency tables, histograms, and polygons
Section titled “Frequency tables, histograms, and polygons”To organize continuous data:
- Split the range into intervals (also called classes or bins) of equal width, such as . Each value goes in exactly one interval.
- Count the frequency in each interval.
- Draw a frequency histogram: bars with no gaps, one per interval, with height equal to the frequency.
- For a frequency polygon, plot a point at the midpoint of the top of each bar and join the points with straight lines. Add a point with frequency at the midpoint of an empty interval at each end, so the polygon starts and ends on the axis.
Here are the masses of apples from one orchard:
| Mass (g) | – | – | – | – | – | – |
|---|---|---|---|---|---|---|
| Frequency |
(Each interval includes its left end but not its right end, so – means .)
Choosing the interval width
Section titled “Choosing the interval width”The interval width changes what you see:
- Too wide (say, or intervals): the details of the shape disappear.
- Too narrow (say, g intervals for apples): most intervals hold , , or values, and the graph looks jagged and random.
- A useful middle ground is often about to intervals.
As you collect more and more data and make the intervals narrower, the frequency polygon starts to look like a smooth curve. That curve is the model for the distribution.
Probability is area
Section titled “Probability is area”If you divide each frequency by the total, you get relative frequencies, which estimate probabilities. For example, .
For a smooth model curve (a probability density curve), the total area under the curve is , and
A single value has no width, so its area is :
That’s why, for continuous variables, and are the same. It also matches real life: the chance an apple has a mass of exactly g is zero. Only ranges make sense.
Worked examples
Section titled “Worked examples”Example 1: Discrete or continuous?
Section titled “Example 1: Discrete or continuous?”Is each random variable discrete or continuous?
- (a) the time it takes to run m
- (b) the number of cars in a parking lot
- (c) the volume of juice in a bottle
- (d) a shoe size
Solution. (a) Continuous: time is measured, and any value in a range is possible. (b) Discrete: cars are counted. (c) Continuous: volume is measured. (d) Discrete: shoe sizes only come in set values like , , , with nothing in between. (The foot length that a shoe size is based on is continuous, though.)
Example 2: Building a frequency table
Section titled “Example 2: Building a frequency table”Twenty students recorded their commute to school, in minutes:
Make a frequency table with intervals of width minutes, starting at . Describe the shape.
Solution. The smallest value is and the largest is , so intervals from to cover everything. Tally each value (remember goes in –, not –):
| Time (min) | – | – | – | – | – |
|---|---|---|---|---|---|
| Frequency |
Check: . ✓
The distribution is symmetric and mound-shaped: most commutes are to minutes, with fewer very short or very long ones.
Example 3: Probability from a frequency table
Section titled “Example 3: Probability from a frequency table”Use the apple table above. Estimate the probability that a randomly chosen apple from the orchard has a mass of at least g, and the probability that it has a mass of exactly g.
Solution. The intervals – and – hold apples:
Mass is continuous, so the probability of any single exact value is : . (An apple recorded as ” g” really has a mass somewhere from g to g, which is a range, not a single value.)
Example 4: Probability as the area of a rectangle
Section titled “Example 4: Probability as the area of a rectangle”A subway train arrives every minutes. If you show up at a random time, your waiting time is equally likely to be anywhere from to minutes. This is a continuous uniform distribution. Find .
Solution. The density “curve” is a flat line. The total area must be , and the base is , so the height is .
The probability is the area of the shaded rectangle:
Common mistakes
Section titled “Common mistakes”Calling something discrete because the data were rounded. Heights recorded to the nearest centimetre look like whole numbers, but height itself is continuous. Ask what’s being measured, not how it was written down.
Leaving gaps between histogram bars. For continuous data, the intervals touch (–, –, …), so the bars touch too. Gaps are for bar graphs of discrete or categorical data.
Putting a boundary value in two intervals. Decide which end each interval includes (usually the left) and stick to it. With , a mass of exactly g goes in the next interval.
Forgetting the zero points on a frequency polygon. The polygon should start and end on the horizontal axis, at the midpoints of the empty intervals just outside the data.
Giving a non-zero probability for a single value. For a continuous variable, . Probabilities only come from ranges, as areas.
Practice
Section titled “Practice”1. (Warm-up) Is each random variable discrete or continuous?
- (a) the mass of a newborn baby
- (b) the number of texts you get in a day
- (c) the temperature of a cup of tea
- (d) the number of red cars that pass in an hour
Solution
(a) Continuous. (b) Discrete. (c) Continuous. (d) Discrete.
2. (Warm-up) A student’s height is recorded as cm, to the nearest centimetre. What range of actual heights could this be? Why is ?
Solution
Any height from cm up to (but not including) cm rounds to cm.
Height is continuous, so a single exact value has no width and its area under the density curve is . Only a range like has a non-zero probability.
3. (Warm-up) Using the apple table, estimate the probability that a randomly chosen apple has a mass less than g.
Solution
apples are under g:
4. (Core) Sixteen reaction times, in seconds, were measured in a science class:
- (a) Make a frequency table with intervals of width s, starting at s.
- (b) Which interval(s) have the highest frequency?
- (c) Estimate the probability that a reaction time is under s.
Solution
(a)
| Time (s) | – | – | – | – | – |
|---|---|---|---|---|---|
| Frequency |
Check: . ✓ (Remember goes in – and goes in –.)
(b) – and – are tied, with each.
(c) times are under s: .
5. (Core) List the points you would plot for the frequency polygon of the commute times in Example 2.
Solution
Plot (midpoint, frequency), including a at each end:
6. (Core) A ferry leaves every minutes, and you arrive at a random time. Your waiting time is uniform from to minutes. Find:
- (a)
- (b)
- (c)
Solution
The height of the rectangle is , so each probability is (width) .
(a)
(b)
(c) : a single value has no width.
7. (Core) A biologist weighs trout from one lake and finds an average mass of g. A classmate weighs different trout from the same lake and gets g. Give two reasons the results differ, and explain why statisticians describe trout masses with a model instead of an exact distribution.
Solution
Sampling: they weighed different fish, and each sample of is only part of the population, so random variation between samples is expected.
Measurement uncertainty: each scale rounds, may be calibrated slightly differently, and a wet, wriggling fish is hard to weigh precisely.
Because you can never weigh every fish exactly, the true distribution can’t be known perfectly. A model (a smooth curve with a few parameters, like a mean and standard deviation) describes the data well and can be used to make predictions.
8. (Challenge) Regroup the apple data into intervals of width g: –, –, –. What is lost compared with the g intervals? What might happen if you used g intervals instead?
Solution
The new frequencies are , , and .
With only three bars, you can still see that the middle is most common, but you lose detail. For example, you can no longer see that – is slightly more common than –, or how quickly the frequencies drop off at the ends.
With g intervals, there would be intervals for only apples, so most would hold , , or apples. The histogram would look jagged, and random ups and downs would hide the overall shape.
9. (Challenge) A continuous random variable takes values from to . Its density curve is a triangle: it rises in a straight line from height at to a peak at , then falls back to at .
- (a) Find the height of the peak.
- (b) Find .
Solution
(a) The total area must be . The triangle has base , so , giving .
(b) From to the height rises from to , so at the height is . The region from to is a small triangle: