Skip to content
Family Table Math
Auto

Confidence Intervals for the Mean

A sample mean of 182182 g is a useful estimate of a population mean, but on its own it hides how uncertain it is. A confidence interval gives a range of plausible values instead, such as ”176176 g to 188188 g, with 95%95\% confidence”. In this topic you’ll build these intervals for the mean of a normal population, both when the population standard deviation is known and when you have to estimate it, and learn to interpret them correctly. It’s the same idea as the margin of error in a poll, made precise.

The sample mean xˉ\bar{x} is a point estimate of the population mean μ\mu (see linear combinations of random variables). A confidence interval surrounds it with a margin that reflects sampling variability:

xˉ−margintoxˉ+margin\bar{x} - \text{margin} \quad \text{to} \quad \bar{x} + \text{margin}

The margin depends on how spread out the population is, how big the sample is, and how confident you want to be.

When σ is known: use the normal distribution

Section titled “When σ is known: use the normal distribution”

If the population is normal with known standard deviation σ\sigma, then Xˉ∼N(μ,σ2n)\bar{X} \sim N\left(\mu, \dfrac{\sigma^2}{n}\right) (see the central limit theorem). A confidence interval for μ\mu is

xˉ−zσn < μ < xˉ+zσn\bar{x} - z\frac{\sigma}{\sqrt{n}} \ \lt\ \mu \ \lt\ \bar{x} + z\frac{\sigma}{\sqrt{n}}

where zz is chosen so that the middle area under the standard normal curve equals the confidence level:

Confidence level90%90\%95%95\%99%99\%
zz1.6451.6451.9601.9602.5762.576

For 95%95\%, you need 2.5%2.5\% in each tail, so zz is the inverse normal of 0.9750.975, which is 1.95996…1.95996\ldots On a GDC, the z-interval function (for example ZInterval) does all of this from σ\sigma, xˉ\bar{x}, nn, and the level, or straight from a list of data.

When σ is unknown: use the t-distribution

Section titled “When σ is unknown: use the t-distribution”

Usually you don’t know σ\sigma, so you estimate it with sn−1s_{n-1}, the sample standard deviation. Using an estimate adds extra uncertainty, so instead of zz you use a value from the t-distribution with n−1n - 1 degrees of freedom:

xˉ−tsn−1n < μ < xˉ+tsn−1n\bar{x} - t\frac{s_{n-1}}{\sqrt{n}} \ \lt\ \mu \ \lt\ \bar{x} + t\frac{s_{n-1}}{\sqrt{n}}

The t-distribution looks like the standard normal but with fatter tails, so tt is a bit bigger than zz and the interval is a bit wider. In practice you use your GDC’s t-interval function (for example TInterval), entering the data or xˉ\bar{x}, sn−1s_{n-1} (labelled SxSx), and nn.

The IB rule is simple: σ known → normal (z); σ unknown → t, regardless of the sample size. Both methods assume the population is normal.

A 95%95\% confidence interval comes from a method that captures the true mean in 95%95\% of samples. If you took many samples and built an interval from each, about 95%95\% of the intervals would contain μ\mu and about 5%5\% would miss it.

Twenty 95% confidence intervals from repeated samples; nineteen contain the true mean μ 40 42 44 46 48 50 52 54 56 58 60
Twenty 95%95\% confidence intervals from twenty different samples. The true mean μ\mu is fixed; the intervals move. Here 1919 of the 2020 capture μ\mu.

Once you’ve calculated a particular interval, like (176.1,187.9)(176.1, 187.9), μ\mu is either in it or not; you just don’t know which. So the correct interpretation is:

“We are 95%95\% confident that the population mean is between 176.1176.1 g and 187.9187.9 g.”

meaning the method you used works 95%95\% of the time. It does not mean that 95%95\% of the individual values lie in the interval, and it’s not about the sample mean (which is definitely in the middle of the interval).

The width of the interval is 2×zσn2 \times z\dfrac{\sigma}{\sqrt{n}} (or the same with tt and sn−1s_{n-1}).

  • Higher confidence → wider interval. To be more confident of catching μ\mu, you need a bigger net: zz goes from 1.6451.645 to 1.9601.960 to 2.5762.576.
  • Bigger sample → narrower interval. The width is proportional to 1n\dfrac{1}{\sqrt{n}}, so multiplying nn by 44 halves the width.
  • More variable population → wider interval.

If a claimed value of μ\mu lies outside a confidence interval, the data give evidence against the claim. If it lies inside, the claim is plausible (but not proven). This connects confidence intervals to hypothesis testing.

The masses of apples from an orchard are normally distributed with standard deviation 1515 g. A random sample of 2525 apples has mean mass 182182 g. Find a 95%95\% confidence interval for the mean mass of all the apples.

Solution. The standard deviation of Xˉ\bar{X} is 1525=3\dfrac{15}{\sqrt{25}} = 3 g.

margin=1.95996…×3=5.8798…interval=182±5.8798…\begin{aligned} \text{margin} &= 1.95996\ldots \times 3 = 5.8798\ldots \\ \text{interval} &= 182 \pm 5.8798\ldots \end{aligned}

The 95%95\% confidence interval is 176.1<μ<187.9176.1 \lt \mu \lt 187.9 g (to 1 d.p.), or (176,188)(176, 188) to 3 s.f. A GDC z-interval with σ=15\sigma = 15, xˉ=182\bar{x} = 182, n=25n = 25 gives the same result.

A random sample of 1010 students recorded how many hours they slept on a school night:

7.2, 6.5, 8.1, 7.8, 6.9, 7.4, 8.4, 6.2, 7.0, 7.57.2,\ 6.5,\ 8.1,\ 7.8,\ 6.9,\ 7.4,\ 8.4,\ 6.2,\ 7.0,\ 7.5

Assuming sleep times are normally distributed, find a 95%95\% confidence interval for the mean sleep time of all students at the school.

Solution. σ\sigma is unknown, so use the t-distribution with 10−1=910 - 1 = 9 degrees of freedom.

From one-variable statistics: xˉ=7.3\bar{x} = 7.3 and sn−1=0.68799…s_{n-1} = 0.68799\ldots

GDC t-interval (data, 95%95\%): the interval is 6.81<μ<7.796.81 \lt \mu \lt 7.79 hours (3 s.f.).

By formula, as a check: t=2.262…t = 2.262\ldots (the t-value with 99 degrees of freedom that leaves 2.5%2.5\% in the upper tail), so

7.3±2.262…×0.68799…10=7.3±0.4922…7.3 \pm 2.262\ldots \times \frac{0.68799\ldots}{\sqrt{10}} = 7.3 \pm 0.4922\ldots

which matches. We are 95%95\% confident that the mean sleep time of all students at the school is between 6.816.81 and 7.797.79 hours.

A company claims its bags of rice have a mean mass of 500500 g. A consumer group weighs a random sample and calculates a 99%99\% confidence interval for the mean of (493.2,498.6)(493.2, 498.6) g.

  • (a) Interpret the interval.
  • (b) Does the interval support the company’s claim?
  • (c) Would a 95%95\% interval from the same sample be wider or narrower?

Solution.

(a) We are 99%99\% confident that the mean mass of all the company’s bags of rice is between 493.2493.2 g and 498.6498.6 g.

(b) No. The claimed value 500500 g lies above the whole interval, so the data suggest the true mean is less than 500500 g.

(c) Narrower: a lower confidence level uses a smaller zz (or tt) value. Since 500500 is outside the 99%99\% interval, it’s also outside the narrower 95%95\% interval.

A researcher wants to estimate the mean commute time of workers in a city. Commute times are normally distributed with standard deviation 1212 minutes. How large a sample is needed for a 95%95\% confidence interval to have a width of at most 55 minutes?

Solution. The width is 2×1.95996…×12n2 \times 1.95996\ldots \times \dfrac{12}{\sqrt{n}}.

2×1.95996…×12n≤5n≥2×1.95996…×125=9.4078…n≥88.507…\begin{aligned} 2 \times 1.95996\ldots \times \frac{12}{\sqrt{n}} &\le 5 \\ \sqrt{n} &\ge \frac{2 \times 1.95996\ldots \times 12}{5} = 9.4078\ldots \\ n &\ge 88.507\ldots \end{aligned}

The sample must contain at least 8989 workers. (Always round up: 8888 would give an interval slightly wider than 55 minutes.)

Using z when σ is unknown. If you’ve calculated the standard deviation from the sample, use the t-distribution with n−1n - 1 degrees of freedom, even for a large sample.

Forgetting to divide by √n. The margin uses the standard deviation of Xˉ\bar{X}, which is σn\dfrac{\sigma}{\sqrt{n}}, not σ\sigma.

Using sₙ instead of sₙ₋₁. For a t-interval, use the unbiased estimate sn−1s_{n-1} (SxSx on the GDC), not σx\sigma x.

Saying “there is a 95% probability that μ is in this interval”. Once the interval is calculated, μ\mu is either in it or not. Say “we are 95%95\% confident”, and know that it describes the long-run success rate of the method.

Thinking the interval contains 95% of the data. A confidence interval estimates the mean. Individual values are much more spread out than the interval.

Rounding the sample size down. If the calculation gives n≥88.5n \ge 88.5, you need n=89n = 89.

1. (Warm-up) A normal population has σ=5\sigma = 5. A random sample of 1616 values has mean 64.064.0. Find a 95%95\% confidence interval for μ\mu.

Solution64.0±1.95996…×516=64.0±2.4499…64.0 \pm 1.95996\ldots \times \frac{5}{\sqrt{16}} = 64.0 \pm 2.4499\ldots

The interval is 61.6<μ<66.461.6 \lt \mu \lt 66.4 (3 s.f.).

2. (Warm-up) A confidence interval for a mean is (12.4,15.0)(12.4, 15.0). Find the sample mean and the margin.

Solution

The sample mean is the midpoint: 12.4+15.02=13.7\dfrac{12.4 + 15.0}{2} = 13.7. The margin is half the width: 15.0−12.42=1.3\dfrac{15.0 - 12.4}{2} = 1.3.

3. (Warm-up) A 90%90\% confidence interval for the mean height of Grade 12 students is (168.2,173.6)(168.2, 173.6) cm. Which statements are correct?

  • (a) 90%90\% of Grade 12 students are between 168.2168.2 cm and 173.6173.6 cm tall.
  • (b) We are 90%90\% confident that the mean height of all Grade 12 students is between 168.2168.2 cm and 173.6173.6 cm.
  • (c) If many samples were taken, about 90%90\% of the intervals made this way would contain the population mean.
  • (d) The sample mean is between 168.2168.2 cm and 173.6173.6 cm with probability 0.90.9.
Solution

(b) and (c) are correct. (a) is wrong because the interval is about the mean, not individual heights. (d) is wrong because the sample mean is known exactly: it’s 170.9170.9 cm, the centre of the interval.

4. (Core) A random sample of 2020 bags of chips from a normal population has mean mass 48.348.3 g and sn−1=6.2s_{n-1} = 6.2 g. Find a 95%95\% confidence interval for the population mean.

Solution

σ\sigma is unknown, so use tt with 1919 degrees of freedom: t=2.0930…t = 2.0930\ldots

48.3±2.0930…×6.220=48.3±2.9016…48.3 \pm 2.0930\ldots \times \frac{6.2}{\sqrt{20}} = 48.3 \pm 2.9016\ldots

The interval is 45.4<μ<51.245.4 \lt \mu \lt 51.2 g (3 s.f.). A GDC t-interval from the summary statistics gives the same.

5. (Core) The lifetime of a type of phone battery is normally distributed with standard deviation 0.80.8 years. A sample of 4040 batteries has mean lifetime 3.423.42 years. Find a 90%90\% confidence interval for the mean lifetime.

Solution

σ\sigma is known, so use z=1.6448…z = 1.6448\ldots (inverse normal of 0.950.95).

3.42±1.6448…×0.840=3.42±0.20806…3.42 \pm 1.6448\ldots \times \frac{0.8}{\sqrt{40}} = 3.42 \pm 0.20806\ldots

The interval is 3.21<μ<3.633.21 \lt \mu \lt 3.63 years (3 s.f.).

6. (Core) The times, in seconds, for 88 randomly chosen swimmers to complete a 5050 m length were:

31.5, 29.8, 33.2, 30.4, 32.7, 28.9, 31.1, 30.631.5,\ 29.8,\ 33.2,\ 30.4,\ 32.7,\ 28.9,\ 31.1,\ 30.6

Assuming the times are normally distributed, find a 99%99\% confidence interval for the mean time, and interpret it.

Solution

σ\sigma is unknown, so use a t-interval with 77 degrees of freedom. From the GDC: xˉ=31.025\bar{x} = 31.025, sn−1=1.4320…s_{n-1} = 1.4320\ldots

GDC t-interval (99%99\%): 29.3<μ<32.829.3 \lt \mu \lt 32.8 seconds (3 s.f.).

(Check: t=3.4994…t = 3.4994\ldots and 31.025±3.4994…×1.4320…8=31.025±1.7718…31.025 \pm 3.4994\ldots \times \dfrac{1.4320\ldots}{\sqrt{8}} = 31.025 \pm 1.7718\ldots.)

We are 99%99\% confident that the mean time for all such swimmers is between 29.329.3 and 32.832.8 seconds.

7. (Core) A normal population has σ=10\sigma = 10.

  • (a) Find the width of a 95%95\% confidence interval for μ\mu from a sample of 2525.
  • (b) Find the width if the sample size is 100100 instead. How does it compare with (a)?
  • (c) Find the width of a 99%99\% interval from a sample of 2525.
Solution

(a) 2×1.95996…×1025=7.842 \times 1.95996\ldots \times \dfrac{10}{\sqrt{25}} = 7.84 (3 s.f.)

(b) 2×1.95996…×10100=3.922 \times 1.95996\ldots \times \dfrac{10}{\sqrt{100}} = 3.92 (3 s.f.). Four times the sample size gives half the width.

(c) 2×2.5758…×1025=10.32 \times 2.5758\ldots \times \dfrac{10}{\sqrt{25}} = 10.3 (3 s.f.). Higher confidence makes the interval wider.

8. (Challenge) A 95%95\% confidence interval for the mean of a normal population with known σ\sigma is (71.2,76.8)(71.2, 76.8). Find the 99%99\% confidence interval from the same sample.

Solution

The centre is xˉ=74\bar{x} = 74 and the margin is 2.82.8. So

1.95996…×σn=2.8⇒σn=1.42859…1.95996\ldots \times \frac{\sigma}{\sqrt{n}} = 2.8 \quad\Rightarrow\quad \frac{\sigma}{\sqrt{n}} = 1.42859\ldots

For 99%99\%: margin =2.5758…×1.42859…=3.6798…= 2.5758\ldots \times 1.42859\ldots = 3.6798\ldots

The 99%99\% interval is 70.3<μ<77.770.3 \lt \mu \lt 77.7 (3 s.f.).

9. (Challenge) A machine fills bottles with a volume that is normally distributed with standard deviation 0.150.15 litres. Quality control wants a 99%99\% confidence interval for the mean volume with a width of no more than 0.050.05 litres. Find the smallest sample size needed.

Solution2×2.5758…×0.15n≤0.05n≥2×2.5758…×0.150.05=15.454…n≥238.85…\begin{aligned} 2 \times 2.5758\ldots \times \frac{0.15}{\sqrt{n}} &\le 0.05 \\ \sqrt{n} &\ge \frac{2 \times 2.5758\ldots \times 0.15}{0.05} = 15.454\ldots \\ n &\ge 238.85\ldots \end{aligned}

The smallest sample size is 239239 bottles.