Confidence Intervals for the Mean
A sample mean of g is a useful estimate of a population mean, but on its own it hides how uncertain it is. A confidence interval gives a range of plausible values instead, such as ” g to g, with confidence”. In this topic you’ll build these intervals for the mean of a normal population, both when the population standard deviation is known and when you have to estimate it, and learn to interpret them correctly. It’s the same idea as the margin of error in a poll, made precise.
Key ideas
Section titled “Key ideas”Point estimates and interval estimates
Section titled “Point estimates and interval estimates”The sample mean is a point estimate of the population mean (see linear combinations of random variables). A confidence interval surrounds it with a margin that reflects sampling variability:
The margin depends on how spread out the population is, how big the sample is, and how confident you want to be.
When σ is known: use the normal distribution
Section titled “When σ is known: use the normal distribution”If the population is normal with known standard deviation , then (see the central limit theorem). A confidence interval for is
where is chosen so that the middle area under the standard normal curve equals the confidence level:
| Confidence level | |||
|---|---|---|---|
For , you need in each tail, so is the inverse normal of , which is On a GDC, the z-interval function (for example ZInterval) does all of this from , , , and the level, or straight from a list of data.
When σ is unknown: use the t-distribution
Section titled “When σ is unknown: use the t-distribution”Usually you don’t know , so you estimate it with , the sample standard deviation. Using an estimate adds extra uncertainty, so instead of you use a value from the t-distribution with degrees of freedom:
The t-distribution looks like the standard normal but with fatter tails, so is a bit bigger than and the interval is a bit wider. In practice you use your GDC’s t-interval function (for example TInterval), entering the data or , (labelled ), and .
The IB rule is simple: σ known → normal (z); σ unknown → t, regardless of the sample size. Both methods assume the population is normal.
What the confidence level means
Section titled “What the confidence level means”A confidence interval comes from a method that captures the true mean in of samples. If you took many samples and built an interval from each, about of the intervals would contain and about would miss it.
Once you’ve calculated a particular interval, like , is either in it or not; you just don’t know which. So the correct interpretation is:
“We are confident that the population mean is between g and g.”
meaning the method you used works of the time. It does not mean that of the individual values lie in the interval, and it’s not about the sample mean (which is definitely in the middle of the interval).
What affects the width
Section titled “What affects the width”The width of the interval is (or the same with and ).
- Higher confidence → wider interval. To be more confident of catching , you need a bigger net: goes from to to .
- Bigger sample → narrower interval. The width is proportional to , so multiplying by halves the width.
- More variable population → wider interval.
Using an interval to judge a claim
Section titled “Using an interval to judge a claim”If a claimed value of lies outside a confidence interval, the data give evidence against the claim. If it lies inside, the claim is plausible (but not proven). This connects confidence intervals to hypothesis testing.
Worked examples
Section titled “Worked examples”Example 1: σ known
Section titled “Example 1: σ known”The masses of apples from an orchard are normally distributed with standard deviation g. A random sample of apples has mean mass g. Find a confidence interval for the mean mass of all the apples.
Solution. The standard deviation of is g.
The confidence interval is g (to 1 d.p.), or to 3 s.f. A GDC z-interval with , , gives the same result.
Example 2: σ unknown, from data
Section titled “Example 2: σ unknown, from data”A random sample of students recorded how many hours they slept on a school night:
Assuming sleep times are normally distributed, find a confidence interval for the mean sleep time of all students at the school.
Solution. is unknown, so use the t-distribution with degrees of freedom.
From one-variable statistics: and
GDC t-interval (data, ): the interval is hours (3 s.f.).
By formula, as a check: (the t-value with degrees of freedom that leaves in the upper tail), so
which matches. We are confident that the mean sleep time of all students at the school is between and hours.
Example 3: Checking a claim
Section titled “Example 3: Checking a claim”A company claims its bags of rice have a mean mass of g. A consumer group weighs a random sample and calculates a confidence interval for the mean of g.
- (a) Interpret the interval.
- (b) Does the interval support the company’s claim?
- (c) Would a interval from the same sample be wider or narrower?
Solution.
(a) We are confident that the mean mass of all the company’s bags of rice is between g and g.
(b) No. The claimed value g lies above the whole interval, so the data suggest the true mean is less than g.
(c) Narrower: a lower confidence level uses a smaller (or ) value. Since is outside the interval, it’s also outside the narrower interval.
Example 4: Choosing the sample size
Section titled “Example 4: Choosing the sample size”A researcher wants to estimate the mean commute time of workers in a city. Commute times are normally distributed with standard deviation minutes. How large a sample is needed for a confidence interval to have a width of at most minutes?
Solution. The width is .
The sample must contain at least workers. (Always round up: would give an interval slightly wider than minutes.)
Common mistakes
Section titled “Common mistakes”Using z when σ is unknown. If you’ve calculated the standard deviation from the sample, use the t-distribution with degrees of freedom, even for a large sample.
Forgetting to divide by √n. The margin uses the standard deviation of , which is , not .
Using sₙ instead of sₙ₋₁. For a t-interval, use the unbiased estimate ( on the GDC), not .
Saying “there is a 95% probability that μ is in this interval”. Once the interval is calculated, is either in it or not. Say “we are confident”, and know that it describes the long-run success rate of the method.
Thinking the interval contains 95% of the data. A confidence interval estimates the mean. Individual values are much more spread out than the interval.
Rounding the sample size down. If the calculation gives , you need .
Practice
Section titled “Practice”1. (Warm-up) A normal population has . A random sample of values has mean . Find a confidence interval for .
Solution
The interval is (3 s.f.).
2. (Warm-up) A confidence interval for a mean is . Find the sample mean and the margin.
Solution
The sample mean is the midpoint: . The margin is half the width: .
3. (Warm-up) A confidence interval for the mean height of Grade 12 students is cm. Which statements are correct?
- (a) of Grade 12 students are between cm and cm tall.
- (b) We are confident that the mean height of all Grade 12 students is between cm and cm.
- (c) If many samples were taken, about of the intervals made this way would contain the population mean.
- (d) The sample mean is between cm and cm with probability .
Solution
(b) and (c) are correct. (a) is wrong because the interval is about the mean, not individual heights. (d) is wrong because the sample mean is known exactly: it’s cm, the centre of the interval.
4. (Core) A random sample of bags of chips from a normal population has mean mass g and g. Find a confidence interval for the population mean.
Solution
is unknown, so use with degrees of freedom:
The interval is g (3 s.f.). A GDC t-interval from the summary statistics gives the same.
5. (Core) The lifetime of a type of phone battery is normally distributed with standard deviation years. A sample of batteries has mean lifetime years. Find a confidence interval for the mean lifetime.
Solution
is known, so use (inverse normal of ).
The interval is years (3 s.f.).
6. (Core) The times, in seconds, for randomly chosen swimmers to complete a m length were:
Assuming the times are normally distributed, find a confidence interval for the mean time, and interpret it.
Solution
is unknown, so use a t-interval with degrees of freedom. From the GDC: ,
GDC t-interval (): seconds (3 s.f.).
(Check: and .)
We are confident that the mean time for all such swimmers is between and seconds.
7. (Core) A normal population has .
- (a) Find the width of a confidence interval for from a sample of .
- (b) Find the width if the sample size is instead. How does it compare with (a)?
- (c) Find the width of a interval from a sample of .
Solution
(a) (3 s.f.)
(b) (3 s.f.). Four times the sample size gives half the width.
(c) (3 s.f.). Higher confidence makes the interval wider.
8. (Challenge) A confidence interval for the mean of a normal population with known is . Find the confidence interval from the same sample.
Solution
The centre is and the margin is . So
For : margin
The interval is (3 s.f.).
9. (Challenge) A machine fills bottles with a volume that is normally distributed with standard deviation litres. Quality control wants a confidence interval for the mean volume with a width of no more than litres. Find the smallest sample size needed.
Solution
The smallest sample size is bottles.