Standard Deviation
Two bakeries can both sell loaves with a mean mass of g, but if one bakery’s loaves range from g to g and the other’s from g to g, they’re not the same at all. The mean tells you the centre; the standard deviation tells you how spread out the data are around it. It’s the most important measure of spread in statistics, and you need it for the normal distribution and z-scores.
Key ideas
Section titled “Key ideas”The range is the simplest measure of spread:
It’s quick, but it depends on only two values, so one unusual value (an outlier) can make it huge. The interquartile range is less affected by outliers.
Deviations, variance, and standard deviation
Section titled “Deviations, variance, and standard deviation”The deviation of a value is how far it is from the mean: . Some deviations are positive and some negative, and they always add to , so you can’t just average them. Instead, you square them first.
- The variance is the mean of the squared deviations.
- The standard deviation is the square root of the variance. Taking the square root brings the units back to the units of the data (cm, g, minutes).
Population vs sample
Section titled “Population vs sample”Use the population formulas when your data are the whole group you care about. Use the sample formulas when your data are a sample used to estimate a larger population.
| Population | Sample | |
|---|---|---|
| Mean | ||
| Variance | ||
| Standard deviation |
Why for a sample? A sample’s values tend to be a little closer to their own mean than to the true population mean, so dividing by would underestimate the spread. Dividing by the slightly smaller corrects for this. That’s why is a bit bigger than for the same numbers (unless all the values are equal, when both are ).
Using technology
Section titled “Using technology”| Tool | Population | Sample |
|---|---|---|
| Spreadsheet (Excel, Google Sheets) | =STDEV.P(A1:A20) | =STDEV.S(A1:A20) |
| Variance in a spreadsheet | =VAR.P(A1:A20) | =VAR.S(A1:A20) |
| TI-83/84 (STAT → CALC → 1-Var Stats) |
Grouped data
Section titled “Grouped data”When data are in a frequency table with intervals, use the midpoint of each interval to stand for every value in it, and weight by the frequency :
The answers are estimates, because the actual values aren’t all at the midpoints.
Interpreting the standard deviation
Section titled “Interpreting the standard deviation”- A small standard deviation means the values are clustered close to the mean (consistent).
- A large standard deviation means the values are spread out (variable).
- A standard deviation of means every value is the same.
- For mound-shaped data, most values are within about standard deviations of the mean (see the 68–95–99.7 rule).
Worked examples
Section titled “Worked examples”Example 1: Population standard deviation by hand
Section titled “Example 1: Population standard deviation by hand”Five siblings timed how long they took to get ready for school, in minutes: . Treat the five siblings as the whole population. Find the range, the variance, and the standard deviation.
Solution. Range: minutes.
Mean: minutes.
| Sum |
Check: the deviations add to . ✓
Example 2: Sample standard deviation by hand
Section titled “Example 2: Sample standard deviation by hand”A gardener picks tomatoes at random from a large crop and weighs them, in grams: . Find the sample standard deviation.
Solution. These are a sample from a larger crop, so use .
Mean: g.
| Sum |
(If you had wrongly used the population formula, you’d get g, which is too small.)
Example 3: Using a spreadsheet
Section titled “Example 3: Using a spreadsheet”The masses, in kilograms, of pumpkins picked at random from a farm field are entered in cells A1 to A12:
Which formula should you use, and what does it give?
Solution. The pumpkins are a sample of the whole field, so use the sample standard deviation:
=AVERAGE(A1:A12)gives kg.=STDEV.S(A1:A12)gives kg.
For comparison, =STDEV.P(A1:A12) gives about kg. With values, the difference is small; with only or values, it would be much bigger.
So a typical pumpkin is within about half a kilogram of the mean mass of kg.
Example 4: Grouped data
Section titled “Example 4: Grouped data”Estimate the mean and standard deviation of the masses of the apples from continuous random variables. Treat the apples as the population.
Solution. Use the midpoints. First the mean:
| Mass (g) | ||||
|---|---|---|---|---|
| – | ||||
| – | ||||
| – | ||||
| – | ||||
| – | ||||
| – | ||||
| Sum |
On a TI-84, enter the midpoints in L1 and the frequencies in L2, then run 1-Var Stats with L1 and frequency list L2.
Common mistakes
Section titled “Common mistakes”Using the wrong formula. Ask: is this the whole group, or a sample from a bigger group? A class’s own test scores, analysed for that class, are a population. Thirty people surveyed to learn about all of Ontario are a sample.
Forgetting to square the deviations. The deviations always add to , so their average tells you nothing. Square first, then add.
Forgetting the square root. The variance is in squared units (like ). The standard deviation is the square root, back in the original units.
Squaring a negative deviation wrong on a calculator. Typing -6² gives on many calculators. Use brackets: (-6)² . Squared deviations are never negative.
Rounding the mean too early. If the mean is , don’t round it to before finding deviations. Keep full accuracy (or use technology) and round only the final answer.
Practice
Section titled “Practice”1. (Warm-up) Find the range of these daily high temperatures, in °C: .
Solution
°C.
2. (Warm-up) Should you use the population () or sample () formula?
- (a) A teacher finds the spread of the test marks of all students in her class.
- (b) A factory weighs cereal boxes from today’s production of .
- (c) Health Canada measures the heights of teenagers to learn about all Canadian teens.
Solution
(a) Population: the class is the whole group of interest.
(b) Sample: the boxes stand for all .
(c) Sample: teens stand for millions.
3. (Warm-up) Treat as a population. Find and .
Solution
. Deviations: . Squares: , with sum .
4. (Core) A cyclist records the time for randomly chosen trips to work, in minutes: . Find the mean and the sample standard deviation.
Solution
minutes.
Deviations: . Squares: , with sum .
5. (Core) Two Grade 12 classes wrote the same test. Both had a mean of . Class A had a standard deviation of and class B had . Describe how the marks in the two classes differ.
Solution
Class A’s marks are tightly clustered: most students scored close to , perhaps roughly to . Class B’s marks are much more spread out, with some students scoring far above and others far below . The means are equal, but class B is much less consistent.
6. (Core) For the data in Example 1 (), find the sample standard deviation, and explain why it’s larger than .
Solution
The sum of squared deviations is still , but you divide by :
It’s larger than because dividing by instead of gives a bigger result. The sample formula is deliberately a little larger to make up for samples tending to underestimate the spread.
7. (Core) Estimate the mean and standard deviation of these grouped commute times. Treat the data as a population.
| Time (min) | – | – | – | – | – |
|---|---|---|---|---|---|
| Frequency |
Solution
Midpoints: . With :
Deviations from : .
8. (Challenge) In Example 1, minutes.
- (a) Every sibling takes minutes longer one morning. What is the new standard deviation?
- (b) Instead, every sibling takes twice as long. What is the new standard deviation?
Solution
(a) Adding to every value moves the mean up by too, so every deviation stays the same. The standard deviation is unchanged: about minutes.
(b) Doubling every value doubles the mean and every deviation, so the standard deviation doubles: minutes.
9. (Challenge) The population has a mean of . Find , then find .
Solution
The sum must be : , so .
Deviations: . Squares: , with sum .