Skip to content
Family Table Math
Auto

The t-Test

Students who used a revision app scored an average of 76.576.5, and students who didn’t scored 69.969.9. Is the app really better, or could a gap like that easily happen by chance with small groups? When you want to compare the means of two populations using a sample from each, you use a two-sample t-test. It follows the same steps as every test in introduction to hypothesis testing; your GDC does the calculation, and your job is to set it up and interpret it.

You have two independent (unpaired) samples, one from each population, such as plants grown in two soils, or students in two different classes. Let μ1\mu_1 and μ2\mu_2 be the two population means. The null hypothesis is always that they’re equal:

H0:μ1=μ2H_0: \mu_1 = \mu_2

The alternative hypothesis depends on the question:

Question asks whether the means are…H1H_1Type of test
differentμ1≠μ2\mu_1 \ne \mu_2two-tailed
μ1\mu_1 is greaterμ1>μ2\mu_1 \gt \mu_2one-tailed
μ1\mu_1 is smallerμ1<μ2\mu_1 \lt \mu_2one-tailed

The test asks: is the gap between the two sample means xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 large compared with how much the samples vary? The t-statistic measures the gap in units of its standard error. A tt close to 00 means the gap is small compared with the natural variation; a tt far from 00 (positive or negative) is evidence that the population means differ.

If H0H_0 is true, tt follows a t-distribution with

ν=n1+n2−2\nu = n_1 + n_2 - 2

degrees of freedom, where n1n_1 and n2n_2 are the sample sizes. A t-distribution is bell-shaped and symmetric about 00, like the normal curve, but with slightly heavier tails.

A t-distribution with 13 degrees of freedom. The areas beyond t = 2.93 and t = -2.93 are shaded, about 0.00586 each, giving a two-tailed p-value of 0.0117. −4 −3 −2 −1 0 1 2 3 4 t = 2.93 t = −2.93 0.00586 0.00586 13 degrees of freedom p = 0.0117
For a two-tailed test with t=2.93t = 2.93 and 1313 degrees of freedom, the p-value is the total shaded area in both tails: 0.01170.0117.

In IB exams you always use technology for the t-test. On your GDC, choose the two-sample t-test, then:

  1. Enter the data as two lists, or enter the summary statistics: each sample’s mean xˉ\bar{x}, standard deviation sn−1s_{n-1} and size nn.
  2. Choose the alternative hypothesis: ≠\ne, >\gt or <\lt.
  3. Choose pooled: yes.

The GDC gives tt, the p-value and the degrees of freedom. Then compare the p-value with the significance level and write your conclusion in context.

If you enter summary statistics, use the sample standard deviation sn−1s_{n-1} (often shown as sxs_x on a GDC), the one that divides by n−1n - 1.

The t-test is only valid if:

  • the two samples are independent and randomly chosen (at SL, the samples are always unpaired: different individuals in each group)
  • the variable is normally distributed in both populations
  • the two populations have equal variances. That’s what “pooled” means: the GDC combines the two samples to estimate one common variance. In IB exams, assume the variances are equal and use the pooled test.

The population variances are unknown, which is exactly why you use a tt-test rather than a normal distribution: the sample standard deviations are estimates.

A researcher records the nightly sleep of 1515 students at each of two schools.

nnMean (hours)sn−1s_{n-1} (hours)
School P15157.407.400.900.90
School Q15157.857.850.780.78

Assuming sleep times are normally distributed with equal variances, test at the 5%5\% significance level whether the mean sleep time is different at the two schools.

Solution. Let μP\mu_P and μQ\mu_Q be the population mean sleep times.

H0:μP=μQH1:μP≠μQH_0: \mu_P = \mu_Q \qquad H_1: \mu_P \ne \mu_Q

“Different” means a two-tailed test. Enter the summary statistics into the GDC’s two-sample t-test (pooled, ≠\ne):

t=−1.46,p=0.154 (to 3 s.f.),ν=15+15−2=28t = -1.46, \qquad p = 0.154 \text{ (to 3 s.f.)}, \qquad \nu = 15 + 15 - 2 = 28

0.154>0.050.154 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the mean sleep time is different at the two schools.

A biology class grew sunflower seedlings in two types of soil and measured their heights (in cm) after two weeks.

Soil A12.412.414.114.113.213.215.015.012.812.813.913.914.614.613.513.5
Soil B11.811.812.912.912.112.113.413.411.511.512.612.613.013.0

Assume heights are normally distributed with equal variance. Test at the 5%5\% significance level whether the mean height is different for the two soils.

Solution. Let μA\mu_A and μB\mu_B be the population mean heights.

H0:μA=μBH1:μA≠μBH_0: \mu_A = \mu_B \qquad H_1: \mu_A \ne \mu_B

Enter the two lists and run the pooled two-sample t-test with ≠\ne. The sample means are xˉA=13.7\bar{x}_A = 13.7 cm and xˉB=12.5\bar{x}_B = 12.5 cm (to 3 s.f.), and

t=2.93,p=0.0117 (to 3 s.f.),ν=8+7−2=13t = 2.93, \qquad p = 0.0117 \text{ (to 3 s.f.)}, \qquad \nu = 8 + 7 - 2 = 13

0.0117<0.050.0117 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean height of seedlings is different in the two soils. (The figure above shows this p-value as the two shaded tails.)

At the 1%1\% level, the conclusion would be different: 0.0117>0.010.0117 \gt 0.01.

A teacher believes a revision app increases mean test scores. Ten students used the app and nine didn’t.

With app7272818168687777858574747979707083837676
Without707065657474686871717777636369697272

Assuming scores are normally distributed with equal variances, test the teacher’s belief at the 5%5\% significance level.

Solution. Let μA\mu_A and μW\mu_W be the population mean scores with and without the app. “Increases” has a direction, so the test is one-tailed:

H0:μA=μWH1:μA>μWH_0: \mu_A = \mu_W \qquad H_1: \mu_A \gt \mu_W

Enter the lists (app first) and run the pooled two-sample t-test with >\gt. The means are 76.576.5 and 69.969.9 (to 3 s.f.), and

t=2.86,p=0.00546 (to 3 s.f.),ν=10+9−2=17t = 2.86, \qquad p = 0.00546 \text{ (to 3 s.f.)}, \qquad \nu = 10 + 9 - 2 = 17

0.00546<0.050.00546 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean test score is higher for students who use the app.

This shows the scores are higher, not necessarily that the app caused it. If students chose whether to use the app, the more motivated ones might have picked it. Randomly assigning students to the two groups would make a causal claim much stronger.

Six batteries of brand X and six of brand Y were tested, with these lifetimes (in hours).

Brand X18.218.219.519.517.817.820.120.118.918.919.319.3
Brand Y18.818.819.919.918.418.419.219.220.420.418.618.6
  • (a) State the assumptions needed for a two-sample t-test.
  • (b) Test at the 10%10\% significance level whether the mean lifetimes are different.

Solution.

(a) The lifetimes of each brand are normally distributed, the two populations have equal variances, and the two samples are independent random samples.

(b) H0:μX=μYH_0: \mu_X = \mu_Y, H1:μX≠μYH_1: \mu_X \ne \mu_Y. The pooled two-sample t-test gives

t=−0.528,p=0.609 (to 3 s.f.),ν=10t = -0.528, \qquad p = 0.609 \text{ (to 3 s.f.)}, \qquad \nu = 10

0.609>0.100.609 \gt 0.10, so do not reject H0H_0. There is insufficient evidence at the 10%10\% significance level that the mean lifetimes of the two brands are different.

The sample means are 18.9718.97 h and 19.2219.22 h. A gap of 0.250.25 h is small compared with how much individual batteries vary, so it could easily be due to chance.

Choosing the wrong alternative hypothesis. Read the question for direction words. “Different” or “changed” means ≠\ne (two-tailed); “higher”, “faster”, “improves” mean >\gt or <\lt (one-tailed). And make sure the order on your GDC matches: if H1H_1 is μ1>μ2\mu_1 \gt \mu_2, enter group 11 as the first list.

Forgetting to choose pooled. IB exam questions expect you to assume equal variances and use the pooled test. The unpooled version gives slightly different values of tt, pp and the degrees of freedom.

Using the wrong standard deviation. When entering summary statistics, use the sample standard deviation sn−1s_{n-1}, not σn\sigma_n.

Using a two-sample t-test on paired data. If the same people are measured twice (before and after), the samples aren’t independent. The SL two-sample test is for two separate groups.

Ignoring the normality assumption. The t-test assumes each population is normally distributed. For strongly skewed data (like household incomes) with small samples, the p-value can’t be trusted.

Concluding that the means are equal. A large p-value means there’s insufficient evidence of a difference, not proof that the means are the same. Write “do not reject H0H_0”, with the significance level and the context.

1. (Warm-up) Write H0H_0 and H1H_1 for each test, defining your symbols.

  • (a) Is the mean time to finish a puzzle different for students who listen to music and students who don’t?
  • (b) Do tomato plants given fertilizer have a greater mean yield than plants without it?
Solution

(a) Let μM\mu_M and μN\mu_N be the population mean times with and without music. H0:μM=μNH_0: \mu_M = \mu_N, H1:μM≠μNH_1: \mu_M \ne \mu_N (two-tailed).

(b) Let μF\mu_F and μN\mu_N be the population mean yields with and without fertilizer. H0:μF=μNH_0: \mu_F = \mu_N, H1:μF>μNH_1: \mu_F \gt \mu_N (one-tailed).

2. (Warm-up) A two-sample t-test comparing the mean heights of two varieties of sunflower gives p=0.034p = 0.034. Write the conclusion at

  • (a) the 5%5\% significance level
  • (b) the 1%1\% significance level.
Solution

(a) 0.034<0.050.034 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean heights of the two varieties are different.

(b) 0.034>0.010.034 \gt 0.01, so do not reject H0H_0. There is insufficient evidence at the 1%1\% significance level that the mean heights are different.

3. (Core) The fuel consumption (in L/100 km) of two car models was measured.

Model 17.87.88.28.27.57.58.08.08.48.47.97.9
Model 28.38.38.68.68.18.18.98.98.48.48.78.78.28.2

Assuming normal distributions with equal variances, test at the 5%5\% significance level whether the mean fuel consumption of the two models is different.

Solution

H0:μ1=μ2H_0: \mu_1 = \mu_2, H1:μ1≠μ2H_1: \mu_1 \ne \mu_2, where μ1\mu_1 and μ2\mu_2 are the population mean fuel consumptions.

The pooled two-sample t-test gives

t=−2.94,p=0.0135 (to 3 s.f.),ν=6+7−2=11t = -2.94, \qquad p = 0.0135 \text{ (to 3 s.f.)}, \qquad \nu = 6 + 7 - 2 = 11

0.0135<0.050.0135 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean fuel consumption of the two models is different. (Model 1’s sample mean, 7.977.97, is lower than Model 2’s, 8.468.46.)

4. (Core) A researcher thinks that drinking coffee reduces mean reaction time. Reaction times (in milliseconds) were recorded for two groups.

Coffee248248262262239239255255251251244244258258247247
No coffee259259270270252252266266249249263263271271258258

Test the researcher’s claim at the 1%1\% significance level, assuming normal distributions with equal variances.

Solution

Let μC\mu_C and μN\mu_N be the population mean reaction times with and without coffee. “Reduces” gives a one-tailed test:

H0:μC=μNH1:μC<μNH_0: \mu_C = \mu_N \qquad H_1: \mu_C \lt \mu_N

The pooled two-sample t-test (coffee as the first list, <\lt) gives

t=−2.70,p=0.00871 (to 3 s.f.),ν=14t = -2.70, \qquad p = 0.00871 \text{ (to 3 s.f.)}, \qquad \nu = 14

0.00871<0.010.00871 \lt 0.01, so reject H0H_0. There is sufficient evidence at the 1%1\% significance level that the mean reaction time is lower for people who drink coffee.

5. (Core) Two classes wrote the same test.

nnMeansn−1s_{n-1}
Class A121264.264.25.15.1
Class B141460.560.54.64.6

Assuming normal distributions with equal variances, test at the 5%5\% significance level whether the mean marks of the two classes are different.

Solution

H0:μA=μBH_0: \mu_A = \mu_B, H1:μA≠μBH_1: \mu_A \ne \mu_B.

Entering the summary statistics into the pooled two-sample t-test:

t=1.95,p=0.0636 (to 3 s.f.),ν=12+14−2=24t = 1.95, \qquad p = 0.0636 \text{ (to 3 s.f.)}, \qquad \nu = 12 + 14 - 2 = 24

0.0636>0.050.0636 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the mean marks of the two classes are different.

6. (Core) Explain why the SL two-sample t-test is not appropriate in each case.

  • (a) Twelve students’ marks are compared before and after a revision course.
  • (b) The annual incomes of 88 people in each of two towns are compared. Income data are strongly right-skewed.
Solution

(a) The same students are measured twice, so the two sets of marks are paired, not independent. The two-sample test assumes two separate, independent groups.

(b) The t-test assumes the variable is normally distributed in each population. Strongly skewed data like incomes aren’t normal, and with only 88 values per group the p-value wouldn’t be reliable.

7. (Core) Two groups of 77 plants were grown under different lights, and their growth (in cm) recorded.

Light A5.25.26.16.15.85.86.46.45.55.56.06.05.75.7
Light B5.05.05.65.65.35.35.95.95.15.15.85.85.45.4

Assuming normal distributions with equal variances, find the p-value and state the conclusion at the 5%5\% level for

  • (a) H1H_1: the mean growth is different under the two lights
  • (b) H1H_1: the mean growth is greater under light A.
Solution

In both cases H0:μA=μBH_0: \mu_A = \mu_B, the pooled test gives t=1.88t = 1.88 (to 3 s.f.), and ν=12\nu = 12.

(a) Two-tailed: p=0.0851p = 0.0851 (to 3 s.f.). 0.0851>0.050.0851 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% level that the mean growth is different under the two lights.

(b) One-tailed: p=0.0426p = 0.0426 (to 3 s.f.), half the two-tailed value. 0.0426<0.050.0426 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% level that the mean growth is greater under light A.

The one-tailed test only “uses up” the significance level in one direction, so it’s easier to reject H0H_0 in that direction. That’s why the choice must be made from the question, before seeing the data.

8. (Challenge) Using the data from question 7, a student tests H1H_1: the mean growth is less under light A. Find the p-value and explain why it’s so large.

Solution

The pooled test with <\lt gives p=0.957p = 0.957 (to 3 s.f.), so do not reject H0H_0.

The p-value is the area in the lower tail, below t=1.88t = 1.88. Since tt is positive (light A’s sample mean is larger), almost all of the distribution is below it: p=1−0.0426=0.957p = 1 - 0.0426 = 0.957. The data point in the opposite direction to H1H_1, so they give no evidence at all for it.

9. (Challenge) Two groups have sample means 52.052.0 and 47.047.0, and both have sample standard deviation 6.06.0. Assume normal distributions with equal variances.

  • (a) Test at the 5%5\% level whether the population means are different if each group has 66 members.
  • (b) Repeat with 1515 members in each group.
  • (c) Explain why the conclusions differ.
Solution

H0:μ1=μ2H_0: \mu_1 = \mu_2, H1:μ1≠μ2H_1: \mu_1 \ne \mu_2.

(a) Summary statistics with n=6n = 6: t=1.44t = 1.44, p=0.180p = 0.180 (to 3 s.f.), ν=10\nu = 10. 0.180>0.050.180 \gt 0.05, so do not reject H0H_0: insufficient evidence at the 5%5\% level that the means differ.

(b) With n=15n = 15: t=2.28t = 2.28, p=0.0303p = 0.0303 (to 3 s.f.), ν=28\nu = 28. 0.0303<0.050.0303 \lt 0.05, so reject H0H_0: sufficient evidence at the 5%5\% level that the means differ.

(c) The gap between the sample means (5.05.0) and the spread are the same in both cases. But sample means from larger samples vary less from sample to sample, so a gap of 5.05.0 is much less likely to be due to chance with 1515 in each group than with 66. Larger samples make it easier to detect a real difference.