Skip to content
Family Table Math

Standard Deviation

Two bakeries can both sell loaves with a mean mass of 500500 g, but if one bakery’s loaves range from 450450 g to 550550 g and the other’s from 495495 g to 505505 g, they’re not the same at all. The mean tells you the centre; the standard deviation tells you how spread out the data are around it. It’s the most important measure of spread in statistics, and you need it for the normal distribution and z-scores.

The range is the simplest measure of spread:

range=maximum−minimum\text{range} = \text{maximum} - \text{minimum}

It’s quick, but it depends on only two values, so one unusual value (an outlier) can make it huge. The interquartile range is less affected by outliers.

Deviations, variance, and standard deviation

Section titled “Deviations, variance, and standard deviation”

The deviation of a value is how far it is from the mean: x−μx - \mu. Some deviations are positive and some negative, and they always add to 00, so you can’t just average them. Instead, you square them first.

  • The variance is the mean of the squared deviations.
  • The standard deviation is the square root of the variance. Taking the square root brings the units back to the units of the data (cm, g, minutes).

Use the population formulas when your data are the whole group you care about. Use the sample formulas when your data are a sample used to estimate a larger population.

PopulationSample
Meanμ=∑xN\mu = \dfrac{\sum x}{N}xˉ=∑xn\bar{x} = \dfrac{\sum x}{n}
Varianceσ2=∑(x−μ)2N\sigma^2 = \dfrac{\sum (x - \mu)^2}{N}s2=∑(x−xˉ)2n−1s^2 = \dfrac{\sum (x - \bar{x})^2}{n - 1}
Standard deviationσ=∑(x−μ)2N\sigma = \sqrt{\dfrac{\sum (x - \mu)^2}{N}}s=∑(x−xˉ)2n−1s = \sqrt{\dfrac{\sum (x - \bar{x})^2}{n - 1}}

Why n−1n - 1 for a sample? A sample’s values tend to be a little closer to their own mean xˉ\bar{x} than to the true population mean, so dividing by nn would underestimate the spread. Dividing by the slightly smaller n−1n - 1 corrects for this. That’s why ss is a bit bigger than σ\sigma for the same numbers (unless all the values are equal, when both are 00).

ToolPopulationSample
Spreadsheet (Excel, Google Sheets)=STDEV.P(A1:A20)=STDEV.S(A1:A20)
Variance in a spreadsheet=VAR.P(A1:A20)=VAR.S(A1:A20)
TI-83/84 (STAT → CALC → 1-Var Stats)σx\sigma xSxSx

When data are in a frequency table with intervals, use the midpoint mm of each interval to stand for every value in it, and weight by the frequency ff:

μ≈∑fmN,σ≈∑f(m−μ)2N\mu \approx \frac{\sum f m}{N}, \qquad \sigma \approx \sqrt{\frac{\sum f (m - \mu)^2}{N}}

The answers are estimates, because the actual values aren’t all at the midpoints.

  • A small standard deviation means the values are clustered close to the mean (consistent).
  • A large standard deviation means the values are spread out (variable).
  • A standard deviation of 00 means every value is the same.
  • For mound-shaped data, most values are within about 22 standard deviations of the mean (see the 68–95–99.7 rule).

Example 1: Population standard deviation by hand

Section titled “Example 1: Population standard deviation by hand”

Five siblings timed how long they took to get ready for school, in minutes: 12,15,18,20,2512, 15, 18, 20, 25. Treat the five siblings as the whole population. Find the range, the variance, and the standard deviation.

Solution. Range: 25−12=1325 - 12 = 13 minutes.

Mean: μ=12+15+18+20+255=905=18\mu = \dfrac{12 + 15 + 18 + 20 + 25}{5} = \dfrac{90}{5} = 18 minutes.

xxx−μx - \mu(x−μ)2(x - \mu)^2
1212−6-63636
1515−3-399
18180000
20202244
2525774949
Sum009898

Check: the deviations add to 00. ✓

σ2=985=19.6,σ=19.6≈4.43 minutes\sigma^2 = \frac{98}{5} = 19.6, \qquad \sigma = \sqrt{19.6} \approx 4.43 \text{ minutes}

Example 2: Sample standard deviation by hand

Section titled “Example 2: Sample standard deviation by hand”

A gardener picks 66 tomatoes at random from a large crop and weighs them, in grams: 112,125,118,130,121,114112, 125, 118, 130, 121, 114. Find the sample standard deviation.

Solution. These are a sample from a larger crop, so use n−1n - 1.

Mean: xˉ=7206=120\bar{x} = \dfrac{720}{6} = 120 g.

xxx−xˉx - \bar{x}(x−xˉ)2(x - \bar{x})^2
112112−8-86464
125125552525
118118−2-244
1301301010100100
1211211111
114114−6-63636
Sum00230230
s2=2306−1=46,s=46≈6.78 gs^2 = \frac{230}{6 - 1} = 46, \qquad s = \sqrt{46} \approx 6.78 \text{ g}

(If you had wrongly used the population formula, you’d get 230/6≈6.19\sqrt{230/6} \approx 6.19 g, which is too small.)

The masses, in kilograms, of 1212 pumpkins picked at random from a farm field are entered in cells A1 to A12:

4.2, 5.1, 3.8, 4.6, 5.5, 4.9, 4.0, 5.3, 4.4, 4.7, 5.0, 4.54.2,\ 5.1,\ 3.8,\ 4.6,\ 5.5,\ 4.9,\ 4.0,\ 5.3,\ 4.4,\ 4.7,\ 5.0,\ 4.5

Which formula should you use, and what does it give?

Solution. The pumpkins are a sample of the whole field, so use the sample standard deviation:

  • =AVERAGE(A1:A12) gives xˉ≈4.67\bar{x} \approx 4.67 kg.
  • =STDEV.S(A1:A12) gives s≈0.52s \approx 0.52 kg.

For comparison, =STDEV.P(A1:A12) gives about 0.500.50 kg. With 1212 values, the difference is small; with only 33 or 44 values, it would be much bigger.

So a typical pumpkin is within about half a kilogram of the mean mass of 4.674.67 kg.

Estimate the mean and standard deviation of the masses of the 5050 apples from continuous random variables. Treat the 5050 apples as the population.

Solution. Use the midpoints. First the mean:

μ≈3(145)+8(155)+14(165)+13(175)+9(185)+3(195)50=851050=170.2 g\mu \approx \frac{3(145) + 8(155) + 14(165) + 13(175) + 9(185) + 3(195)}{50} = \frac{8510}{50} = 170.2 \text{ g}
Mass (g)mmffm−μm - \muf(m−μ)2f(m - \mu)^2
140140–15015014514533−25.2-25.21905.121905.12
150150–16016015515588−15.2-15.21848.321848.32
160160–1701701651651414−5.2-5.2378.56378.56
170170–18018017517513134.84.8299.52299.52
180180–1901901851859914.814.81971.361971.36
190190–2002001951953324.824.81845.121845.12
Sum505082488248
σ≈824850=164.96≈12.84 g\sigma \approx \sqrt{\frac{8248}{50}} = \sqrt{164.96} \approx 12.84 \text{ g}

On a TI-84, enter the midpoints in L1 and the frequencies in L2, then run 1-Var Stats with L1 and frequency list L2.

Using the wrong formula. Ask: is this the whole group, or a sample from a bigger group? A class’s own test scores, analysed for that class, are a population. Thirty people surveyed to learn about all of Ontario are a sample.

Forgetting to square the deviations. The deviations always add to 00, so their average tells you nothing. Square first, then add.

Forgetting the square root. The variance is in squared units (like g2\text{g}^2). The standard deviation is the square root, back in the original units.

Squaring a negative deviation wrong on a calculator. Typing -6² gives −36-36 on many calculators. Use brackets: (-6)² =36= 36. Squared deviations are never negative.

Rounding the mean too early. If the mean is 4.66674.6667, don’t round it to 4.74.7 before finding deviations. Keep full accuracy (or use technology) and round only the final answer.

1. (Warm-up) Find the range of these daily high temperatures, in °C: 14,19,11,22,17,9,1614, 19, 11, 22, 17, 9, 16.

Solution

22−9=1322 - 9 = 13 °C.

2. (Warm-up) Should you use the population (σ\sigma) or sample (ss) formula?

  • (a) A teacher finds the spread of the test marks of all 2828 students in her class.
  • (b) A factory weighs 4040 cereal boxes from today’s production of 10 00010\,000.
  • (c) Health Canada measures the heights of 20002000 teenagers to learn about all Canadian teens.
Solution

(a) Population: the class is the whole group of interest.

(b) Sample: the 4040 boxes stand for all 10 00010\,000.

(c) Sample: 20002000 teens stand for millions.

3. (Warm-up) Treat 4,7,7,10,124, 7, 7, 10, 12 as a population. Find μ\mu and σ\sigma.

Solution

μ=405=8\mu = \dfrac{40}{5} = 8. Deviations: −4,−1,−1,2,4-4, -1, -1, 2, 4. Squares: 16,1,1,4,1616, 1, 1, 4, 16, with sum 3838.

σ=385=7.6≈2.76\sigma = \sqrt{\frac{38}{5}} = \sqrt{7.6} \approx 2.76

4. (Core) A cyclist records the time for 66 randomly chosen trips to work, in minutes: 38,45,41,50,36,4238, 45, 41, 50, 36, 42. Find the mean and the sample standard deviation.

Solution

xˉ=2526=42\bar{x} = \dfrac{252}{6} = 42 minutes.

Deviations: −4,3,−1,8,−6,0-4, 3, -1, 8, -6, 0. Squares: 16,9,1,64,36,016, 9, 1, 64, 36, 0, with sum 126126.

s=1266−1=25.2≈5.02 minutess = \sqrt{\frac{126}{6 - 1}} = \sqrt{25.2} \approx 5.02 \text{ minutes}

5. (Core) Two Grade 12 classes wrote the same test. Both had a mean of 72%72\%. Class A had a standard deviation of 4%4\% and class B had 15%15\%. Describe how the marks in the two classes differ.

Solution

Class A’s marks are tightly clustered: most students scored close to 72%72\%, perhaps roughly 64%64\% to 80%80\%. Class B’s marks are much more spread out, with some students scoring far above and others far below 72%72\%. The means are equal, but class B is much less consistent.

6. (Core) For the data in Example 1 (12,15,18,20,2512, 15, 18, 20, 25), find the sample standard deviation, and explain why it’s larger than σ\sigma.

Solution

The sum of squared deviations is still 9898, but you divide by n−1=4n - 1 = 4:

s=984=24.5≈4.95 minutess = \sqrt{\frac{98}{4}} = \sqrt{24.5} \approx 4.95 \text{ minutes}

It’s larger than σ≈4.43\sigma \approx 4.43 because dividing by 44 instead of 55 gives a bigger result. The sample formula is deliberately a little larger to make up for samples tending to underestimate the spread.

7. (Core) Estimate the mean and standard deviation of these grouped commute times. Treat the data as a population.

Time (min)55–10101010–15151515–20202020–25252525–3030
Frequency2244884422
Solution

Midpoints: 7.5,12.5,17.5,22.5,27.57.5, 12.5, 17.5, 22.5, 27.5. With N=20N = 20:

μ≈2(7.5)+4(12.5)+8(17.5)+4(22.5)+2(27.5)20=35020=17.5 minutes\mu \approx \frac{2(7.5) + 4(12.5) + 8(17.5) + 4(22.5) + 2(27.5)}{20} = \frac{350}{20} = 17.5 \text{ minutes}

Deviations from 17.517.5: −10,−5,0,5,10-10, -5, 0, 5, 10.

∑f(m−μ)2=2(100)+4(25)+8(0)+4(25)+2(100)=600\sum f(m - \mu)^2 = 2(100) + 4(25) + 8(0) + 4(25) + 2(100) = 600σ≈60020=30≈5.48 minutes\sigma \approx \sqrt{\frac{600}{20}} = \sqrt{30} \approx 5.48 \text{ minutes}

8. (Challenge) In Example 1, σ≈4.43\sigma \approx 4.43 minutes.

  • (a) Every sibling takes 33 minutes longer one morning. What is the new standard deviation?
  • (b) Instead, every sibling takes twice as long. What is the new standard deviation?
Solution

(a) Adding 33 to every value moves the mean up by 33 too, so every deviation stays the same. The standard deviation is unchanged: about 4.434.43 minutes.

(b) Doubling every value doubles the mean and every deviation, so the standard deviation doubles: σ=4×19.6=219.6≈8.85\sigma = \sqrt{4 \times 19.6} = 2\sqrt{19.6} \approx 8.85 minutes.

9. (Challenge) The population 6,8,10,12,x6, 8, 10, 12, x has a mean of 1010. Find xx, then find σ\sigma.

Solution

The sum must be 5×10=505 \times 10 = 50: 6+8+10+12+x=506 + 8 + 10 + 12 + x = 50, so x=14x = 14.

Deviations: −4,−2,0,2,4-4, -2, 0, 2, 4. Squares: 16,4,0,4,1616, 4, 0, 4, 16, with sum 4040.

σ=405=8≈2.83\sigma = \sqrt{\frac{40}{5}} = \sqrt{8} \approx 2.83