The Normal Distribution
Measure the heights of a few thousand teenagers, the masses of eggs from a farm, or the times of runners in a km race, and draw a histogram. Again and again, you’ll see the same shape: a symmetric hump in the middle, tapering off on both sides. This bell curve is the normal distribution, the most important model in statistics. Once you know a population is roughly normal, its mean and standard deviation tell you almost everything about it.
Key ideas
Section titled “Key ideas”A model for continuous data
Section titled “A model for continuous data”In continuous random variables, you saw that a frequency polygon starts to look like a smooth curve as you collect more data and narrow the intervals. For many measurements, that smooth curve is the normal curve. Like any probability density curve, the total area under it is , and probabilities are areas.
Notation
Section titled “Notation”If is normally distributed with mean and standard deviation , write
Read it as ” is normal with mean and variance .” Watch out: the second number is the variance, not the standard deviation. For example, has . It’s often written as to make easy to see.
Properties of the normal curve
Section titled “Properties of the normal curve”- It’s bell-shaped and symmetric about the mean.
- The mean, median, and mode are all equal, at the centre (the peak).
- It never touches the horizontal axis, but it gets very close beyond about standard deviations from the mean.
- The mean sets the centre (where the curve sits). The standard deviation sets the spread: a larger gives a wider, flatter curve; a smaller gives a narrower, taller one. (The area is always , so a wider curve must be lower.)
- The total area under the curve is , and half of it () is on each side of the mean.
The 68–95–99.7 rule
Section titled “The 68–95–99.7 rule”For any normal distribution:
- about of the data are within standard deviation of the mean, between and
- about are within standard deviations, between and
- about are within standard deviations, between and
This is also called the empirical rule. Splitting each band in half (using symmetry) gives the percentages in the diagram.
(The rule is rounded. More precise values are , , and . For areas at other values, use z-scores.)
When is the normal model a good fit?
Section titled “When is the normal model a good fit?”The normal distribution models data where many small, independent factors add up, so values near the average are common and extreme values in either direction are rare:
- physical measurements of a large group: heights, arm spans, foot lengths
- masses or volumes of products filled by a machine (bags of chips, bottles of juice)
- measurement errors, and scores on large standardized tests
- times for a repeated task, like running a race many times
It’s not a good model for data that are clearly skewed (lopsided) or that can’t go below when the mean is close to : incomes, house prices, waiting times, or the number of goals in a game. A quick check: for normal data, the mean and median should be close, and a histogram should look roughly symmetric and mound-shaped.
Worked examples
Section titled “Worked examples”Example 1: Using the rule directly
Section titled “Example 1: Using the rule directly”The heights of Grade 12 students at a large school are normally distributed with , in centimetres. Between which two heights are the middle of students? The middle ?
Solution. Here and .
Middle : cm to cm.
Middle : cm to cm.
Example 2: Adding up regions
Section titled “Example 2: Adding up regions”For the same heights, estimate:
- (a)
- (b)
Solution. First mark the key values: is , is , and is .
(a) are between and , so are outside, split equally between the two tails:
(b) From to is , and from to is :
Example 3: How many?
Section titled “Example 3: How many?”The school has Grade 12 students. About how many are taller than cm?
Solution. . Above is (or: outside, half above).
Example 4: Finding the mean and standard deviation
Section titled “Example 4: Finding the mean and standard deviation”The masses of bags of carrots are normally distributed. About of bags have masses between g and g. Find and .
Solution. The middle is centred on the mean, so is halfway:
The interval runs from to , which is wide:
So .
Common mistakes
Section titled “Common mistakes”Reading the variance as the standard deviation. In , the standard deviation is , not . Always check which one you’re given.
Using 68% for a one-sided region. is the area on both sides of the mean, from to . From to alone is .
Forgetting to split the tails. If is in the middle, the left over is split: in each tail, not in one.
Using the rule for values that aren’t whole standard deviations away. The rule only works at , , and . For a value like , use z-scores and a table or technology.
Assuming all data are normal. Incomes, wait times, and many counts are skewed. Look at a histogram, or compare the mean and median, before using the normal model.
Practice
Section titled “Practice”1. (Warm-up) For each distribution, state the mean and the standard deviation.
- (a)
- (b)
- (c)
Solution
(a) , .
(b) , .
(c) , .
2. (Warm-up) True or false? For a normal distribution:
- (a) the median is greater than the mean
- (b) about half the data are above the mean
- (c) about of the data are within standard deviations of the mean
Solution
(a) False: the mean, median, and mode are equal. (b) True: the curve is symmetric. (c) True.
3. (Warm-up) The masses of eggs from a farm are normally distributed with a mean of g and a standard deviation of g. What percentage of eggs have masses between g and g?
Solution
and , so this is within standard deviation of the mean: about .
4. (Core) For the eggs in Question 3, estimate:
- (a) the percentage of eggs heavier than g
- (b) the percentage of eggs between g and g
Solution
(a) . Above that is .
(b) and : .
5. (Core) The lifetimes of a type of LED bulb are normally distributed with hours and hours. A store sells of these bulbs. About how many will last:
- (a) more than hours?
- (b) fewer than hours?
Solution
(a) , so about last longer: bulbs.
(b) , so about last less: bulbs.
6. (Core) Which of these would you expect to be approximately normally distributed? Explain briefly.
- (a) the arm spans of all Grade 10 students in Ontario
- (b) the annual incomes of Canadian households
- (c) the volume of pop in cans filled by a machine
- (d) the number of siblings students have
Solution
(a) Yes: a physical measurement of a large group.
(b) No: incomes are skewed to the right, with a few very large incomes pulling the mean above the median.
(c) Yes: machine fills vary a little above and below the target, symmetrically.
(d) No: it’s a small count, can’t go below , and is skewed to the right (most students have , , or siblings, a few have many).
7. (Core) The times for a school’s cross-country runners to complete a km course are normally distributed. About of runners finish between and minutes. Find and , and the time that only about of runners beat.
Solution
minutes, and , so minutes.
Beating a time means running faster (a lower time). About of runners are below minutes.
8. (Challenge) Commute times for a city’s workers are approximately normal. About of commutes are longer than minutes, and about are shorter than minutes. Find and .
Solution
above means (since ). below means .
Subtract the equations: , so minutes. Then minutes.
Check: . ✓
9. (Challenge) A machine fills bottles of water with volumes that are normally distributed, with mL and mL. A quality inspector finds a bottle holding mL. Is this unusual? What might it suggest?
Solution
mL is mL below the mean, which is standard deviations. Only about of bottles are even below mL (), so mL is very unusual if the machine is working properly.
It suggests something may be wrong: the machine could need adjusting (its mean may have drifted lower, or its spread increased). The inspector should check more bottles.