Skip to content
Family Table Math
Auto

The Central Limit Theorem

Take one measurement and it could land almost anywhere in the population. Take the mean of 4040 measurements and something remarkable happens: whatever shape the population has, the sample mean behaves like a normal variable, and it is much less spread out than a single value. This is the central limit theorem, and it’s the reason the normal distribution appears all over statistics, from quality control to opinion polls to confidence intervals.

From linear combinations of random variables, you already know how to find the mean and variance of a combination. When the variables are normal and independent, there’s a bonus: the combination is normal too.

If X1,X2,…,XnX_1, X_2, \dots, X_n are independent and each Xi∼N(μi,σi2)X_i \sim N(\mu_i, \sigma_i^2), then

a1X1±a2X2±⋯±anXn∼N(a1μ1±a2μ2±⋯±anμn,  a12σ12+a22σ22+⋯+an2σn2)a_1X_1 \pm a_2X_2 \pm \dots \pm a_nX_n \sim N\big(a_1\mu_1 \pm a_2\mu_2 \pm \dots \pm a_n\mu_n,\ \ a_1^2\sigma_1^2 + a_2^2\sigma_2^2 + \dots + a_n^2\sigma_n^2\big)

So to find a probability about a sum or difference, you:

  1. Define the new variable (for example D=A−BD = A - B or T=X1+⋯+X12T = X_1 + \dots + X_{12}).
  2. Find its mean and variance with the usual rules.
  3. Use your GDC’s normal cdf with that mean and standard deviation (the square root of the variance).

A common trick: to find P(A>B)P(A \gt B), rewrite it as P(A−B>0)P(A - B \gt 0).

Take a random sample of size nn from a population and work out its mean. Different samples give different means, so the sample mean Xˉ\bar{X} is itself a random variable. Its distribution is called the sampling distribution of the mean.

If the population is normal, X∼N(μ,σ2)X \sim N(\mu, \sigma^2), then

Xˉ∼N(μ, σ2n)\bar{X} \sim N\left(\mu,\ \frac{\sigma^2}{n}\right)

This is exact, for any sample size. The centre stays at μ\mu, but the variance is divided by nn. The standard deviation of Xˉ\bar{X} is σn\dfrac{\sigma}{\sqrt{n}} (often called the standard error of the mean). Bigger samples give means that cluster more tightly around μ\mu.

The total of the sample, X1+X2+⋯+XnX_1 + X_2 + \dots + X_n, is normal too, with mean nμn\mu and variance nσ2n\sigma^2.

What if the population isn’t normal? The central limit theorem says:

For a random sample of size nn from any population with mean μ\mu and variance σ2\sigma^2, the distribution of Xˉ\bar{X} approaches N(μ,σ2n)N\left(\mu, \dfrac{\sigma^2}{n}\right) as nn gets large.

How large is “large” depends on the population: for a roughly symmetric population, quite small samples are fine, while for a very skewed one you need more. In IB examinations, n>30n \gt 30 is considered sufficient. The same applies to the sample total, which is approximately N(nμ,nσ2)N(n\mu, n\sigma^2).

The figure shows this for a skewed population (waiting times with mean 44 minutes and variance 88). The mean of 55 values is less skewed; the mean of 3030 values is almost a perfect bell curve, centred at 44 and much narrower.

A skewed population and the distributions of the sample mean for n = 5 and n = 30 2 4 6 8 10 0.2 0.4 0.6 0 population mean of 5 mean of 30 waiting time (minutes) probability density
A skewed population (mean 44, variance 88) and the exact distributions of Xˉ\bar{X} for n=5n = 5 and n=30n = 30. As nn grows, Xˉ\bar{X} becomes normal and narrower.
PopulationSample sizeDistribution of the sample mean
Normalany nnexactly N(μ,σ2n)N\left(\mu, \dfrac{\sigma^2}{n}\right)
Not normal (or unknown)n>30n \gt 30approximately N(μ,σ2n)N\left(\mu, \dfrac{\sigma^2}{n}\right) by the CLT
Not normal (or unknown)small nncan’t assume normal

The mass of an egg is N(60,42)N(60, 4^2) grams, and the mass of an empty carton is N(30,22)N(30, 2^2) grams, all independent. A carton holds 1212 eggs. Find the probability that a full carton has a mass of more than 770770 g.

Solution. Let T=C+E1+E2+⋯+E12T = C + E_1 + E_2 + \dots + E_{12} (twelve different eggs, so a sum, not 12E12E).

E(T)=30+12(60)=750,Var(T)=22+12(42)=4+192=196E(T) = 30 + 12(60) = 750, \qquad \mathrm{Var}(T) = 2^2 + 12(4^2) = 4 + 192 = 196

A sum of independent normal variables is normal, so T∼N(750,196)T \sim N(750, 196), with standard deviation 196=14\sqrt{196} = 14.

GDC normal cdf with lower bound 770770, upper bound a very large number, μ=750\mu = 750, σ=14\sigma = 14:

P(T>770)=0.0766 (3 s.f.)P(T \gt 770) = 0.0766 \text{ (3 s.f.)}

In the long jump, Aisha’s distances are A∼N(5.2,0.32)A \sim N(5.2, 0.3^2) metres and Bea’s are B∼N(5.0,0.42)B \sim N(5.0, 0.4^2) metres, independently. Find the probability that Bea jumps further than Aisha on a given attempt.

Solution. P(B>A)=P(A−B<0)P(B \gt A) = P(A - B \lt 0). Let D=A−BD = A - B.

E(D)=5.2−5.0=0.2,Var(D)=0.32+0.42=0.25E(D) = 5.2 - 5.0 = 0.2, \qquad \mathrm{Var}(D) = 0.3^2 + 0.4^2 = 0.25

So D∼N(0.2,0.25)D \sim N(0.2, 0.25), with standard deviation 0.50.5.

P(D<0)=0.345 (3 s.f.)P(D \lt 0) = 0.345 \text{ (3 s.f.)}

Bea wins about a third of the time, even though Aisha’s mean is higher. (Check: P(D>0)=0.655P(D \gt 0) = 0.655, and the two add to 11.)

Example 3: One bottle or the mean of nine?

Section titled “Example 3: One bottle or the mean of nine?”

The volume of juice in a bottle is X∼N(250,122)X \sim N(250, 12^2) ml.

  • (a) Find the probability that one bottle contains less than 245245 ml.
  • (b) A random sample of 99 bottles is taken. Find the probability that their mean volume is less than 245245 ml.

Solution.

(a) P(X<245)=0.338P(X \lt 245) = 0.338 (3 s.f.), using μ=250\mu = 250, σ=12\sigma = 12.

(b) The population is normal, so Xˉ∼N(250,1229)=N(250,16)\bar{X} \sim N\left(250, \dfrac{12^2}{9}\right) = N(250, 16), with standard deviation 129=4\dfrac{12}{\sqrt{9}} = 4.

P(Xˉ<245)=0.106 (3 s.f.)P(\bar{X} \lt 245) = 0.106 \text{ (3 s.f.)}

A single bottle is quite likely to be 55 ml short, but the average of nine bottles being 55 ml short is much less likely, because the short and full bottles balance out.

Example 4: Using the CLT for a skewed population

Section titled “Example 4: Using the CLT for a skewed population”

The daily screen time of teenagers in a city is skewed to the right, with mean 3.23.2 hours and standard deviation 1.51.5 hours. A random sample of 4040 teenagers is taken.

  • (a) Find the probability that the sample mean is more than 3.53.5 hours.
  • (b) Find the probability that the total screen time of the 4040 teenagers is more than 135135 hours.

Solution.

(a) The population isn’t normal, but n=40>30n = 40 \gt 30, so by the central limit theorem

Xˉ≈N(3.2, 1.5240),standard deviation 1.540=0.23717…\bar{X} \approx N\left(3.2,\ \frac{1.5^2}{40}\right), \qquad \text{standard deviation } \frac{1.5}{\sqrt{40}} = 0.23717\ldots P(Xˉ>3.5)≈0.103 (3 s.f.)P(\bar{X} \gt 3.5) \approx 0.103 \text{ (3 s.f.)}

(b) The total S=X1+⋯+X40S = X_1 + \dots + X_{40} is approximately normal with mean 40(3.2)=12840(3.2) = 128 and variance 40(1.52)=9040(1.5^2) = 90, so its standard deviation is 90=9.4868…\sqrt{90} = 9.4868\ldots

P(S>135)≈0.230 (3 s.f.)P(S \gt 135) \approx 0.230 \text{ (3 s.f.)}

Check: S>135S \gt 135 is the same as Xˉ>13540=3.375\bar{X} \gt \dfrac{135}{40} = 3.375, and the normal cdf for Xˉ\bar{X} above 3.3753.375 gives the same 0.2300.230.

Using σ instead of σ/√n for the sample mean. Questions about one value use σ\sigma. Questions about a mean of nn values use σn\dfrac{\sigma}{\sqrt{n}}. Underline the word “mean” when you see it.

Typing the variance into the normal cdf. Most GDCs ask for the standard deviation. If Xˉ∼N(250,16)\bar{X} \sim N(250, 16), enter σ=4\sigma = 4, not 1616.

Writing nX for a total. The total of 1212 eggs is E1+⋯+E12E_1 + \dots + E_{12}, with variance 12σ212\sigma^2. Using 12E12E gives variance 144σ2144\sigma^2, which is far too big.

Subtracting variances for a difference. Var(A−B)=Var(A)+Var(B)\mathrm{Var}(A - B) = \mathrm{Var}(A) + \mathrm{Var}(B), always (for independent variables).

Using the CLT for small samples. With a skewed population and n=8n = 8, you can’t assume Xˉ\bar{X} is normal. The CLT needs a large sample (n>30n \gt 30 in exams), unless the population itself is normal.

Thinking the CLT makes the population normal. The CLT is about the distribution of the sample mean (or total). The individual values keep whatever shape the population has.

1. (Warm-up) X∼N(80,36)X \sim N(80, 36). A random sample of 1616 values is taken. State the distribution of Xˉ\bar{X}, including its standard deviation.

SolutionXˉ∼N(80,3616)=N(80,2.25)\bar{X} \sim N\left(80, \frac{36}{16}\right) = N(80, 2.25)

The standard deviation is 2.25=1.5\sqrt{2.25} = 1.5 (or 616=1.5\dfrac{6}{\sqrt{16}} = 1.5).

2. (Warm-up) In each case, can you assume that Xˉ\bar{X} is (at least approximately) normal? Give a reason.

  • (a) A sample of 55 from a normal population.
  • (b) A sample of 5050 from a skewed population.
  • (c) A sample of 88 from a skewed population.
Solution

(a) Yes, exactly normal: the mean of a sample from a normal population is normal for any nn.

(b) Yes, approximately normal by the central limit theorem, since n=50>30n = 50 \gt 30.

(c) No: the population isn’t normal and the sample is too small for the central limit theorem.

3. (Warm-up) X∼N(20,9)X \sim N(20, 9) and Y∼N(15,16)Y \sim N(15, 16) are independent. State the distributions of:

  • (a) X+YX + Y
  • (b) X−YX - Y
  • (c) 2X−Y2X - Y
Solution

(a) X+Y∼N(35,25)X + Y \sim N(35, 25)

(b) X−Y∼N(5,25)X - Y \sim N(5, 25)

(c) 2X−Y∼N(2(20)−15, 4(9)+16)=N(25,52)2X - Y \sim N(2(20) - 15,\ 4(9) + 16) = N(25, 52)

4. (Core) The masses of adults using an elevator are normally distributed with mean 7575 kg and standard deviation 1212 kg. The elevator’s safe load is 10001000 kg. Find the probability that 1212 randomly chosen adults have a total mass of more than 10001000 kg.

Solution

T=X1+⋯+X12T = X_1 + \dots + X_{12} has mean 12(75)=90012(75) = 900 and variance 12(122)=172812(12^2) = 1728, so T∼N(900,1728)T \sim N(900, 1728) with standard deviation 1728=41.569…\sqrt{1728} = 41.569\ldots

P(T>1000)=0.00807 (3 s.f.)P(T \gt 1000) = 0.00807 \text{ (3 s.f.)}

5. (Core) A coffee machine pours F∼N(250,82)F \sim N(250, 8^2) ml into a cup. The capacity of a cup is C∼N(260,32)C \sim N(260, 3^2) ml, independently. Find the probability that a cup overflows.

Solution

The cup overflows when F>CF \gt C, that is, F−C>0F - C \gt 0.

E(F−C)=250−260=−10,Var(F−C)=64+9=73E(F - C) = 250 - 260 = -10, \qquad \mathrm{Var}(F - C) = 64 + 9 = 73

So F−C∼N(−10,73)F - C \sim N(-10, 73), with standard deviation 73=8.544…\sqrt{73} = 8.544\ldots

P(F−C>0)=0.121 (3 s.f.)P(F - C \gt 0) = 0.121 \text{ (3 s.f.)}

6. (Core) A fair six-sided die has mean score 3.53.5 and variance 3512\dfrac{35}{12}. It is rolled 5050 times. Find the approximate probability that the mean score is between 3.33.3 and 3.73.7.

Solution

The scores are not normal, but n=50>30n = 50 \gt 30, so by the CLT

Xˉ≈N(3.5, 35/1250),standard deviation 35600=0.24152…\bar{X} \approx N\left(3.5,\ \frac{35/12}{50}\right), \qquad \text{standard deviation } \sqrt{\frac{35}{600}} = 0.24152\ldotsP(3.3<Xˉ<3.7)≈0.592 (3 s.f.)P(3.3 \lt \bar{X} \lt 3.7) \approx 0.592 \text{ (3 s.f.)}

7. (Core) The masses of chocolate bars are normally distributed with standard deviation 2.52.5 g. A sample of nn bars is taken. Find the smallest nn for which the probability that the sample mean is within 0.50.5 g of the population mean is at least 0.950.95.

Solution

Xˉ∼N(μ,2.52n)\bar{X} \sim N\left(\mu, \dfrac{2.5^2}{n}\right). We need P(∣Xˉ−μ∣<0.5)≥0.95P(|\bar{X} - \mu| \lt 0.5) \ge 0.95. By symmetry, 0.50.5 must be at least 1.95996…1.95996\ldots standard deviations (the inverse normal of 0.9750.975):

1.95996…×2.5n≤0.5⇒n≥9.7998…⇒n≥96.04…1.95996\ldots \times \frac{2.5}{\sqrt{n}} \le 0.5 \quad\Rightarrow\quad \sqrt{n} \ge 9.7998\ldots \quad\Rightarrow\quad n \ge 96.04\ldots

The smallest sample size is n=97n = 97.

8. (Challenge) The lengths of wooden planks are L∼N(2.4,0.022)L \sim N(2.4, 0.02^2) metres, independently.

  • (a) Three planks are laid end to end. Find the probability that their total length is more than 7.257.25 m.
  • (b) Find the probability that the total length of two planks is more than 0.050.05 m longer than twice the length of a third plank.
Solution

(a) The total has mean 3(2.4)=7.23(2.4) = 7.2 and variance 3(0.022)=0.00123(0.02^2) = 0.0012, so standard deviation 0.0012=0.034641…\sqrt{0.0012} = 0.034641\ldots

P(total>7.25)=0.0745 (3 s.f.)P(\text{total} \gt 7.25) = 0.0745 \text{ (3 s.f.)}

(b) Let D=L1+L2−2L3D = L_1 + L_2 - 2L_3. Then E(D)=2.4+2.4−2(2.4)=0E(D) = 2.4 + 2.4 - 2(2.4) = 0 and

Var(D)=0.022+0.022+22(0.022)=6(0.0004)=0.0024\mathrm{Var}(D) = 0.02^2 + 0.02^2 + 2^2(0.02^2) = 6(0.0004) = 0.0024

So D∼N(0,0.0024)D \sim N(0, 0.0024), with standard deviation 0.048989…0.048989\ldots

P(D>0.05)=0.154 (3 s.f.)P(D \gt 0.05) = 0.154 \text{ (3 s.f.)}

Notice that 2L32L_3 is one plank doubled, so its variance is 4(0.022)4(0.02^2), while L1+L2L_1 + L_2 is two different planks.

9. (Challenge) The amount spent by a customer at a hardware store has mean $42 and standard deviation $18, and the distribution is skewed. On one day, 6464 customers visit.

  • (a) Find the probability that the total amount spent is more than $2900.
  • (b) Find the amount, to the nearest dollar, that the total exceeds with probability 0.90.9.
Solution

(a) n=64>30n = 64 \gt 30, so by the CLT the total SS is approximately normal, with mean 64×42=268864 \times 42 = 2688 dollars and standard deviation 1864=14418\sqrt{64} = 144 dollars.

P(S>2900)≈0.0705 (3 s.f.)P(S \gt 2900) \approx 0.0705 \text{ (3 s.f.)}

(b) We need kk with P(S>k)=0.9P(S \gt k) = 0.9, so P(S<k)=0.1P(S \lt k) = 0.1. Inverse normal with area 0.10.1, μ=2688\mu = 2688, σ=144\sigma = 144:

k=2503.45…k = 2503.45\ldots

The total exceeds about $2503 with probability 0.90.9.