Skip to content
Family Table Math
Auto

Introduction to Hypothesis Testing

Does a new revision app really raise marks, or did the group that used it just get lucky? Is a die fair, or loaded? A hypothesis test is a careful way to decide whether data give real evidence for a claim, or whether the pattern could easily be down to chance. Every test you meet in IB Applications and Interpretation SL (the chi-squared tests and the t-test) follows the same steps, so it’s worth learning them once, properly.

Every test compares two statements about a population:

  • The null hypothesis H0H_0 is the “nothing special is going on” statement: no difference, no relationship, no change, or the claimed model is right. You assume it’s true unless the data give strong evidence against it.
  • The alternative hypothesis H1H_1 is what you’re looking for evidence of: there is a difference, a relationship, or a change.

You can write them in words or with symbols, whichever suits the test. For example, comparing the mean marks μA\mu_A and μB\mu_B of two classes:

H0:μA=μBH1:μA≠μBH_0: \mu_A = \mu_B \qquad H_1: \mu_A \ne \mu_B

And for a χ2\chi^2 test of independence, in words:

  • H0H_0: favourite sport is independent of grade.
  • H1H_1: favourite sport is not independent of grade.

The null hypothesis always contains the “equals” (or “independent”, or “fits the model”).

The significance level is how strong the evidence must be before you reject H0H_0. It’s a probability, usually 10%10\%, 5%5\% or 1%1\%, and it’s chosen before you look at the data. A smaller significance level means you need stronger evidence: rejecting H0H_0 at the 1%1\% level is a stronger result than rejecting it at the 5%5\% level.

Your GDC turns the data into a test statistic (like χ2\chi^2 or tt) and a p-value:

The p-value is the probability of getting a result at least as extreme as the one observed, if H0H_0 is true.

A small p-value means the data would be surprising if H0H_0 were true, so it counts as evidence against H0H_0.

For example, suppose you toss a coin 1010 times and get 99 heads. If the coin is fair, the probability of 99 or more heads is

P(X≥9)=(109)(0.5)10+(1010)(0.5)10=111024=0.0107 (to 3 s.f.)P(X \ge 9) = \binom{10}{9}(0.5)^{10} + \binom{10}{10}(0.5)^{10} = \frac{11}{1024} = 0.0107 \text{ (to 3 s.f.)}

That’s the p-value for testing “the coin is fair” against “the coin is biased toward heads”. Results this extreme happen only about 1%1\% of the time with a fair coin, so it’s fairly strong evidence of bias. (You won’t be asked to carry out this particular test at SL; it’s here to show what a p-value means.)

Compare the p-value with the significance level:

Decision
p-value <\lt significance levelReject H0H_0. There is sufficient evidence for H1H_1.
p-value >\gt significance levelDo not reject H0H_0. There is insufficient evidence for H1H_1.

Equivalently, you can compare the test statistic with a critical value. The critical value marks the edge of the critical region: the set of test-statistic values extreme enough to reject H0H_0. If the test statistic lands in the critical region, reject H0H_0. In IB χ2\chi^2 questions, the critical value is given to you when you need it.

The wording of H1H_1 decides where the critical region is.

  • Two-tailed test: H1H_1 says “different”, with no direction, such as μA≠μB\mu_A \ne \mu_B. Extreme results in either direction count as evidence, so the significance level is split between the two tails.
  • One-tailed test: H1H_1 has a direction, such as μA>μB\mu_A \gt \mu_B (“higher”, “faster”, “improves”). Only results in that direction count, so the whole significance level sits in one tail.
Two bell-shaped curves. Left: a two-tailed test at the 5% level shades 2.5% in each tail. Right: a one-tailed test at the 5% level shades 5% in the upper tail only. The shaded areas are the critical regions where H0 is rejected. Two-tailed test, 5% level 2.5% 2.5% do not reject H₀ test statistic One-tailed test, 5% level 5% do not reject H₀ test statistic
At the 5%5\% level, a two-tailed test puts 2.5%2.5\% in each tail; a one-tailed test puts all 5%5\% in one tail.

Choose one-tailed or two-tailed from the question (what you want to show), before seeing the data, never because of which way the data happen to point. The χ2\chi^2 tests you’ll meet are always upper-tail tests: only a large χ2\chi^2 is evidence against H0H_0.

A conclusion has three parts: the comparison, the decision, and what it means in context.

Since p=0.012<0.05p = 0.012 \lt 0.05, reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean marks of the two classes are different.

Since p=0.31>0.05p = 0.31 \gt 0.05, do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that favourite sport depends on grade.

Never write “accept H0H_0” or “this proves H0H_0”. Not rejecting H0H_0 just means the data aren’t strong enough to rule it out, in the same way that a “not guilty” verdict doesn’t prove someone is innocent.

  1. Write H0H_0 and H1H_1 (in words or symbols).
  2. State the significance level, and whether the test is one- or two-tailed.
  3. Use technology to find the test statistic and the p-value (and the degrees of freedom, where needed).
  4. Compare the p-value with the significance level (or the test statistic with the critical value).
  5. Make a decision and write a conclusion in context.

Write suitable null and alternative hypotheses for each situation.

  • (a) A researcher wants to know whether the mean reaction time of teenagers is different from that of adults.
  • (b) A farmer wants to know whether a new feed increases the mean mass of chickens.
  • (c) A teacher wants to know whether choice of lunch option is related to grade level.

Solution.

(a) Let μT\mu_T and μA\mu_A be the mean reaction times of teenagers and adults. “Different” has no direction, so the test is two-tailed:

H0:μT=μAH1:μT≠μAH_0: \mu_T = \mu_A \qquad H_1: \mu_T \ne \mu_A

(b) Let μN\mu_N and μO\mu_O be the mean masses of chickens on the new and old feed. “Increases” has a direction, so the test is one-tailed:

H0:μN=μOH1:μN>μOH_0: \mu_N = \mu_O \qquad H_1: \mu_N \gt \mu_O

(c) Words work best here:

  • H0H_0: choice of lunch option is independent of grade level.
  • H1H_1: choice of lunch option is not independent of grade level.

A test of H0H_0: mean commute times in two cities are equal, against H1H_1: they are different, gives p=0.027p = 0.027.

  • (a) State the conclusion at the 5%5\% significance level.
  • (b) State the conclusion at the 1%1\% significance level.

Solution.

(a) 0.027<0.050.027 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the mean commute times in the two cities are different.

(b) 0.027>0.010.027 \gt 0.01, so do not reject H0H_0. There is insufficient evidence at the 1%1\% significance level that the mean commute times are different.

The same data can lead to different conclusions at different significance levels. That’s why the level must be chosen before the test, not after.

A χ2\chi^2 test at the 5%5\% significance level gives a test statistic of χ2=7.21\chi^2 = 7.21. The critical value is 9.4889.488. The hypotheses are H0H_0: the type of music a student prefers is independent of their age group, and H1H_1: it is not independent. Write a conclusion.

Solution. Only large values of χ2\chi^2 are in the critical region, and 7.21<9.4887.21 \lt 9.488, so the test statistic is not in the critical region.

Do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that music preference depends on age group.

A school tested whether a new timetable changed the mean number of late arrivals per day. At the 5%5\% level the p-value was 0.180.18. A student wrote:

”p=0.18p = 0.18, so we accept H0H_0. This proves the new timetable makes no difference to late arrivals.”

Explain what’s wrong, and write a better conclusion.

Solution. Three problems:

  • It doesn’t compare the p-value with the significance level.
  • “Accept H0H_0” and “proves” are too strong. A large p-value only means the data are consistent with H0H_0; a real (perhaps small) effect could still exist that this sample didn’t detect.
  • It should state the significance level.

Better: since p=0.18>0.05p = 0.18 \gt 0.05, do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the new timetable changed the mean number of late arrivals per day.

Putting the claim you want to show in H₀. H0H_0 is always the “no difference / no relationship / fits the model” statement, and it contains the equals sign. What you’re looking for evidence of goes in H1H_1.

Writing “accept H₀”. You either reject H0H_0 or you do not reject it. Not rejecting it doesn’t prove it’s true; it just means the evidence isn’t strong enough.

Mixing up the decision rule. A small p-value is evidence against H0H_0. Reject when the p-value is less than the significance level. A quick memory aid: “if p is low, H0H_0 must go”.

Choosing one- or two-tailed after looking at the data. The direction comes from the question (“increases”, “different”, “faster”), not from which way the sample means happen to point. Deciding after you’ve seen the data makes a “significant” result too easy to get.

Leaving out the context. “Reject H0H_0” on its own doesn’t answer the question. Say what it means for the situation: which means differ, or which variables are related, and at what significance level.

Reading the p-value as “the probability that H₀ is true”. The p-value is the probability of data at least this extreme assuming H0H_0 is true. It is not the probability that H0H_0 is true.

1. (Warm-up) Write H0H_0 and H1H_1 for each situation, and say whether the test is one-tailed or two-tailed.

  • (a) A company wants to know whether the mean battery life of its phones is different from a rival brand’s.
  • (b) A coach wants to know whether athletes who stretch have a lower mean injury recovery time than those who don’t.
Solution

(a) Let μC\mu_C and μR\mu_R be the mean battery lives. H0:μC=μRH_0: \mu_C = \mu_R, H1:μC≠μRH_1: \mu_C \ne \mu_R. Two-tailed (“different”).

(b) Let μS\mu_S and μN\mu_N be the mean recovery times of athletes who stretch and who don’t. H0:μS=μNH_0: \mu_S = \mu_N, H1:μS<μNH_1: \mu_S \lt \mu_N. One-tailed (“lower”).

2. (Warm-up) For each p-value, state whether you would reject H0H_0 at the 5%5\% significance level.

  • (a) p=0.003p = 0.003
  • (b) p=0.072p = 0.072
  • (c) p=0.049p = 0.049
Solution

(a) 0.003<0.050.003 \lt 0.05: reject H0H_0.

(b) 0.072>0.050.072 \gt 0.05: do not reject H0H_0.

(c) 0.049<0.050.049 \lt 0.05: reject H0H_0 (only just).

3. (Warm-up) Write the hypotheses, in words, for a test of whether a six-sided die is fair.

Solution
  • H0H_0: the die is fair (each score has probability 16\dfrac{1}{6}).
  • H1H_1: the die is not fair.

4. (Core) A test of H0H_0: preferred social media app is independent of age group, against H1H_1: it is not independent, gives p=0.064p = 0.064. Write a conclusion at

  • (a) the 10%10\% significance level
  • (b) the 5%5\% significance level.
Solution

(a) 0.064<0.100.064 \lt 0.10, so reject H0H_0. There is sufficient evidence at the 10%10\% significance level that preferred social media app depends on age group.

(b) 0.064>0.050.064 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that preferred social media app depends on age group.

5. (Core) A χ2\chi^2 test at the 1%1\% significance level gives χ2=14.2\chi^2 = 14.2, with critical value 13.27713.277. The hypotheses are H0H_0: a spinner is fair, and H1H_1: the spinner is not fair. Write a conclusion.

Solution

14.2>13.27714.2 \gt 13.277, so the test statistic is in the critical region. Reject H0H_0. There is sufficient evidence at the 1%1\% significance level that the spinner is not fair.

6. (Core) Explain, in context, what p=0.02p = 0.02 means for a test of H0H_0: a coin is fair, against H1H_1: the coin is biased.

Solution

If the coin really were fair, there would only be a 2%2\% probability of getting a result at least as far from “half heads” as the one observed. That’s unlikely, so it counts as evidence that the coin is biased (sufficient evidence at the 5%5\% level, but not at the 1%1\% level).

It does not mean there’s a 2%2\% chance the coin is fair.

7. (Core) A dietitian claims a new breakfast cereal lowers mean cholesterol. Her colleague writes H1:μnew≠μoldH_1: \mu_{\text{new}} \ne \mu_{\text{old}}.

  • (a) Is this the right alternative hypothesis for the claim? Explain.
  • (b) Write the correct hypotheses.
Solution

(a) No. The claim has a direction (“lowers”), so the test should be one-tailed. ≠\ne is for a two-tailed test, which looks for a change in either direction.

(b) H0:μnew=μoldH_0: \mu_{\text{new}} = \mu_{\text{old}}, H1:μnew<μoldH_1: \mu_{\text{new}} \lt \mu_{\text{old}}, where μ\mu is the mean cholesterol level of people eating each cereal.

8. (Challenge) A two-tailed test of H0:μA=μBH_0: \mu_A = \mu_B against H1:μA≠μBH_1: \mu_A \ne \mu_B gives p=0.064p = 0.064, and the sample mean for AA is larger than the sample mean for BB. The test statistic has a symmetric distribution.

  • (a) What would the p-value be for the one-tailed test with H1:μA>μBH_1: \mu_A \gt \mu_B?
  • (b) Compare the conclusions of the two tests at the 5%5\% level.
  • (c) Why must the choice between them be made before the data are collected?
Solution

(a) The two-tailed p-value counts both tails equally, so the one-tailed p-value (in the direction the data point) is half of it: 0.0642=0.032\dfrac{0.064}{2} = 0.032.

(b) Two-tailed: 0.064>0.050.064 \gt 0.05, do not reject H0H_0; insufficient evidence that the means differ. One-tailed: 0.032<0.050.032 \lt 0.05, reject H0H_0; sufficient evidence that μA>μB\mu_A \gt \mu_B.

(c) If you looked at the data first and then picked the one-tailed test in the direction the data point, you would be halving the p-value after the fact. That makes it too easy to get a “significant” result just by chance, so the test would no longer have the significance level it claims.

9. (Challenge, enrichment — not an AI SL exam skill) A coin is tossed 1010 times and lands heads 88 times. Find the p-value for testing H0H_0: the coin is fair, against H1H_1: the coin is biased toward heads, and state the conclusion at the 5%5\% significance level.

Solution

If the coin is fair, the number of heads is X∼B(10,0.5)X \sim B(10, 0.5). The p-value is the probability of a result at least as extreme as 88 heads:

P(X≥8)=[(108)+(109)+(1010)](0.5)10=45+10+11024=561024=0.0547 (to 3 s.f.)\begin{aligned} P(X \ge 8) &= \left[\binom{10}{8} + \binom{10}{9} + \binom{10}{10}\right](0.5)^{10} \\ &= \frac{45 + 10 + 1}{1024} = \frac{56}{1024} = 0.0547 \text{ (to 3 s.f.)} \end{aligned}

0.0547>0.050.0547 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the coin is biased toward heads. (Compare 99 heads, where p=0.0107p = 0.0107 was enough.)