The t-Test
Students who used a revision app scored an average of , and students who didn’t scored . Is the app really better, or could a gap like that easily happen by chance with small groups? When you want to compare the means of two populations using a sample from each, you use a two-sample t-test. It follows the same steps as every test in introduction to hypothesis testing; your GDC does the calculation, and your job is to set it up and interpret it.
Key ideas
Section titled “Key ideas”What the test compares
Section titled “What the test compares”You have two independent (unpaired) samples, one from each population, such as plants grown in two soils, or students in two different classes. Let and be the two population means. The null hypothesis is always that they’re equal:
The alternative hypothesis depends on the question:
| Question asks whether the means are… | Type of test | |
|---|---|---|
| different | two-tailed | |
| is greater | one-tailed | |
| is smaller | one-tailed |
The idea behind the test statistic
Section titled “The idea behind the test statistic”The test asks: is the gap between the two sample means large compared with how much the samples vary? The t-statistic measures the gap in units of its standard error. A close to means the gap is small compared with the natural variation; a far from (positive or negative) is evidence that the population means differ.
If is true, follows a t-distribution with
degrees of freedom, where and are the sample sizes. A t-distribution is bell-shaped and symmetric about , like the normal curve, but with slightly heavier tails.
Doing the test with technology
Section titled “Doing the test with technology”In IB exams you always use technology for the t-test. On your GDC, choose the two-sample t-test, then:
- Enter the data as two lists, or enter the summary statistics: each sample’s mean , standard deviation and size .
- Choose the alternative hypothesis: , or .
- Choose pooled: yes.
The GDC gives , the p-value and the degrees of freedom. Then compare the p-value with the significance level and write your conclusion in context.
If you enter summary statistics, use the sample standard deviation (often shown as on a GDC), the one that divides by .
Assumptions
Section titled “Assumptions”The t-test is only valid if:
- the two samples are independent and randomly chosen (at SL, the samples are always unpaired: different individuals in each group)
- the variable is normally distributed in both populations
- the two populations have equal variances. That’s what “pooled” means: the GDC combines the two samples to estimate one common variance. In IB exams, assume the variances are equal and use the pooled test.
The population variances are unknown, which is exactly why you use a -test rather than a normal distribution: the sample standard deviations are estimates.
Worked examples
Section titled “Worked examples”Example 1: From summary statistics
Section titled “Example 1: From summary statistics”A researcher records the nightly sleep of students at each of two schools.
| Mean (hours) | (hours) | ||
|---|---|---|---|
| School P | |||
| School Q |
Assuming sleep times are normally distributed with equal variances, test at the significance level whether the mean sleep time is different at the two schools.
Solution. Let and be the population mean sleep times.
“Different” means a two-tailed test. Enter the summary statistics into the GDC’s two-sample t-test (pooled, ):
, so do not reject . There is insufficient evidence at the significance level that the mean sleep time is different at the two schools.
Example 2: From raw data
Section titled “Example 2: From raw data”A biology class grew sunflower seedlings in two types of soil and measured their heights (in cm) after two weeks.
| Soil A | ||||||||
|---|---|---|---|---|---|---|---|---|
| Soil B |
Assume heights are normally distributed with equal variance. Test at the significance level whether the mean height is different for the two soils.
Solution. Let and be the population mean heights.
Enter the two lists and run the pooled two-sample t-test with . The sample means are cm and cm (to 3 s.f.), and
, so reject . There is sufficient evidence at the significance level that the mean height of seedlings is different in the two soils. (The figure above shows this p-value as the two shaded tails.)
At the level, the conclusion would be different: .
Example 3: A one-tailed test
Section titled “Example 3: A one-tailed test”A teacher believes a revision app increases mean test scores. Ten students used the app and nine didn’t.
| With app | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Without |
Assuming scores are normally distributed with equal variances, test the teacher’s belief at the significance level.
Solution. Let and be the population mean scores with and without the app. “Increases” has a direction, so the test is one-tailed:
Enter the lists (app first) and run the pooled two-sample t-test with . The means are and (to 3 s.f.), and
, so reject . There is sufficient evidence at the significance level that the mean test score is higher for students who use the app.
This shows the scores are higher, not necessarily that the app caused it. If students chose whether to use the app, the more motivated ones might have picked it. Randomly assigning students to the two groups would make a causal claim much stronger.
Example 4: No significant difference
Section titled “Example 4: No significant difference”Six batteries of brand X and six of brand Y were tested, with these lifetimes (in hours).
| Brand X | ||||||
|---|---|---|---|---|---|---|
| Brand Y |
- (a) State the assumptions needed for a two-sample t-test.
- (b) Test at the significance level whether the mean lifetimes are different.
Solution.
(a) The lifetimes of each brand are normally distributed, the two populations have equal variances, and the two samples are independent random samples.
(b) , . The pooled two-sample t-test gives
, so do not reject . There is insufficient evidence at the significance level that the mean lifetimes of the two brands are different.
The sample means are h and h. A gap of h is small compared with how much individual batteries vary, so it could easily be due to chance.
Common mistakes
Section titled “Common mistakes”Choosing the wrong alternative hypothesis. Read the question for direction words. “Different” or “changed” means (two-tailed); “higher”, “faster”, “improves” mean or (one-tailed). And make sure the order on your GDC matches: if is , enter group as the first list.
Forgetting to choose pooled. IB exam questions expect you to assume equal variances and use the pooled test. The unpooled version gives slightly different values of , and the degrees of freedom.
Using the wrong standard deviation. When entering summary statistics, use the sample standard deviation , not .
Using a two-sample t-test on paired data. If the same people are measured twice (before and after), the samples aren’t independent. The SL two-sample test is for two separate groups.
Ignoring the normality assumption. The t-test assumes each population is normally distributed. For strongly skewed data (like household incomes) with small samples, the p-value can’t be trusted.
Concluding that the means are equal. A large p-value means there’s insufficient evidence of a difference, not proof that the means are the same. Write “do not reject ”, with the significance level and the context.
Practice
Section titled “Practice”1. (Warm-up) Write and for each test, defining your symbols.
- (a) Is the mean time to finish a puzzle different for students who listen to music and students who don’t?
- (b) Do tomato plants given fertilizer have a greater mean yield than plants without it?
Solution
(a) Let and be the population mean times with and without music. , (two-tailed).
(b) Let and be the population mean yields with and without fertilizer. , (one-tailed).
2. (Warm-up) A two-sample t-test comparing the mean heights of two varieties of sunflower gives . Write the conclusion at
- (a) the significance level
- (b) the significance level.
Solution
(a) , so reject . There is sufficient evidence at the significance level that the mean heights of the two varieties are different.
(b) , so do not reject . There is insufficient evidence at the significance level that the mean heights are different.
3. (Core) The fuel consumption (in L/100 km) of two car models was measured.
| Model 1 | |||||||
|---|---|---|---|---|---|---|---|
| Model 2 |
Assuming normal distributions with equal variances, test at the significance level whether the mean fuel consumption of the two models is different.
Solution
, , where and are the population mean fuel consumptions.
The pooled two-sample t-test gives
, so reject . There is sufficient evidence at the significance level that the mean fuel consumption of the two models is different. (Model 1’s sample mean, , is lower than Model 2’s, .)
4. (Core) A researcher thinks that drinking coffee reduces mean reaction time. Reaction times (in milliseconds) were recorded for two groups.
| Coffee | ||||||||
|---|---|---|---|---|---|---|---|---|
| No coffee |
Test the researcher’s claim at the significance level, assuming normal distributions with equal variances.
Solution
Let and be the population mean reaction times with and without coffee. “Reduces” gives a one-tailed test:
The pooled two-sample t-test (coffee as the first list, ) gives
, so reject . There is sufficient evidence at the significance level that the mean reaction time is lower for people who drink coffee.
5. (Core) Two classes wrote the same test.
| Mean | |||
|---|---|---|---|
| Class A | |||
| Class B |
Assuming normal distributions with equal variances, test at the significance level whether the mean marks of the two classes are different.
Solution
, .
Entering the summary statistics into the pooled two-sample t-test:
, so do not reject . There is insufficient evidence at the significance level that the mean marks of the two classes are different.
6. (Core) Explain why the SL two-sample t-test is not appropriate in each case.
- (a) Twelve students’ marks are compared before and after a revision course.
- (b) The annual incomes of people in each of two towns are compared. Income data are strongly right-skewed.
Solution
(a) The same students are measured twice, so the two sets of marks are paired, not independent. The two-sample test assumes two separate, independent groups.
(b) The t-test assumes the variable is normally distributed in each population. Strongly skewed data like incomes aren’t normal, and with only values per group the p-value wouldn’t be reliable.
7. (Core) Two groups of plants were grown under different lights, and their growth (in cm) recorded.
| Light A | |||||||
|---|---|---|---|---|---|---|---|
| Light B |
Assuming normal distributions with equal variances, find the p-value and state the conclusion at the level for
- (a) : the mean growth is different under the two lights
- (b) : the mean growth is greater under light A.
Solution
In both cases , the pooled test gives (to 3 s.f.), and .
(a) Two-tailed: (to 3 s.f.). , so do not reject . There is insufficient evidence at the level that the mean growth is different under the two lights.
(b) One-tailed: (to 3 s.f.), half the two-tailed value. , so reject . There is sufficient evidence at the level that the mean growth is greater under light A.
The one-tailed test only “uses up” the significance level in one direction, so it’s easier to reject in that direction. That’s why the choice must be made from the question, before seeing the data.
8. (Challenge) Using the data from question 7, a student tests : the mean growth is less under light A. Find the p-value and explain why it’s so large.
Solution
The pooled test with gives (to 3 s.f.), so do not reject .
The p-value is the area in the lower tail, below . Since is positive (light A’s sample mean is larger), almost all of the distribution is below it: . The data point in the opposite direction to , so they give no evidence at all for it.
9. (Challenge) Two groups have sample means and , and both have sample standard deviation . Assume normal distributions with equal variances.
- (a) Test at the level whether the population means are different if each group has members.
- (b) Repeat with members in each group.
- (c) Explain why the conclusions differ.
Solution
, .
(a) Summary statistics with : , (to 3 s.f.), . , so do not reject : insufficient evidence at the level that the means differ.
(b) With : , (to 3 s.f.), . , so reject : sufficient evidence at the level that the means differ.
(c) The gap between the sample means () and the spread are the same in both cases. But sample means from larger samples vary less from sample to sample, so a gap of is much less likely to be due to chance with in each group than with . Larger samples make it easier to detect a real difference.