Introduction to Hypothesis Testing
Does a new revision app really raise marks, or did the group that used it just get lucky? Is a die fair, or loaded? A hypothesis test is a careful way to decide whether data give real evidence for a claim, or whether the pattern could easily be down to chance. Every test you meet in IB Applications and Interpretation SL (the chi-squared tests and the t-test) follows the same steps, so it’s worth learning them once, properly.
Key ideas
Section titled “Key ideas”The two hypotheses
Section titled “The two hypotheses”Every test compares two statements about a population:
- The null hypothesis is the “nothing special is going on” statement: no difference, no relationship, no change, or the claimed model is right. You assume it’s true unless the data give strong evidence against it.
- The alternative hypothesis is what you’re looking for evidence of: there is a difference, a relationship, or a change.
You can write them in words or with symbols, whichever suits the test. For example, comparing the mean marks and of two classes:
And for a test of independence, in words:
- : favourite sport is independent of grade.
- : favourite sport is not independent of grade.
The null hypothesis always contains the “equals” (or “independent”, or “fits the model”).
The significance level
Section titled “The significance level”The significance level is how strong the evidence must be before you reject . It’s a probability, usually , or , and it’s chosen before you look at the data. A smaller significance level means you need stronger evidence: rejecting at the level is a stronger result than rejecting it at the level.
The p-value
Section titled “The p-value”Your GDC turns the data into a test statistic (like or ) and a p-value:
The p-value is the probability of getting a result at least as extreme as the one observed, if is true.
A small p-value means the data would be surprising if were true, so it counts as evidence against .
For example, suppose you toss a coin times and get heads. If the coin is fair, the probability of or more heads is
That’s the p-value for testing “the coin is fair” against “the coin is biased toward heads”. Results this extreme happen only about of the time with a fair coin, so it’s fairly strong evidence of bias. (You won’t be asked to carry out this particular test at SL; it’s here to show what a p-value means.)
The decision rule
Section titled “The decision rule”Compare the p-value with the significance level:
| Decision | |
|---|---|
| p-value significance level | Reject . There is sufficient evidence for . |
| p-value significance level | Do not reject . There is insufficient evidence for . |
Equivalently, you can compare the test statistic with a critical value. The critical value marks the edge of the critical region: the set of test-statistic values extreme enough to reject . If the test statistic lands in the critical region, reject . In IB questions, the critical value is given to you when you need it.
One-tailed and two-tailed tests
Section titled “One-tailed and two-tailed tests”The wording of decides where the critical region is.
- Two-tailed test: says “different”, with no direction, such as . Extreme results in either direction count as evidence, so the significance level is split between the two tails.
- One-tailed test: has a direction, such as (“higher”, “faster”, “improves”). Only results in that direction count, so the whole significance level sits in one tail.
Choose one-tailed or two-tailed from the question (what you want to show), before seeing the data, never because of which way the data happen to point. The tests you’ll meet are always upper-tail tests: only a large is evidence against .
Writing the conclusion
Section titled “Writing the conclusion”A conclusion has three parts: the comparison, the decision, and what it means in context.
Since , reject . There is sufficient evidence at the significance level that the mean marks of the two classes are different.
Since , do not reject . There is insufficient evidence at the significance level that favourite sport depends on grade.
Never write “accept ” or “this proves ”. Not rejecting just means the data aren’t strong enough to rule it out, in the same way that a “not guilty” verdict doesn’t prove someone is innocent.
The structure of every test
Section titled “The structure of every test”- Write and (in words or symbols).
- State the significance level, and whether the test is one- or two-tailed.
- Use technology to find the test statistic and the p-value (and the degrees of freedom, where needed).
- Compare the p-value with the significance level (or the test statistic with the critical value).
- Make a decision and write a conclusion in context.
Worked examples
Section titled “Worked examples”Example 1: Writing hypotheses
Section titled “Example 1: Writing hypotheses”Write suitable null and alternative hypotheses for each situation.
- (a) A researcher wants to know whether the mean reaction time of teenagers is different from that of adults.
- (b) A farmer wants to know whether a new feed increases the mean mass of chickens.
- (c) A teacher wants to know whether choice of lunch option is related to grade level.
Solution.
(a) Let and be the mean reaction times of teenagers and adults. “Different” has no direction, so the test is two-tailed:
(b) Let and be the mean masses of chickens on the new and old feed. “Increases” has a direction, so the test is one-tailed:
(c) Words work best here:
- : choice of lunch option is independent of grade level.
- : choice of lunch option is not independent of grade level.
Example 2: Using a p-value
Section titled “Example 2: Using a p-value”A test of : mean commute times in two cities are equal, against : they are different, gives .
- (a) State the conclusion at the significance level.
- (b) State the conclusion at the significance level.
Solution.
(a) , so reject . There is sufficient evidence at the significance level that the mean commute times in the two cities are different.
(b) , so do not reject . There is insufficient evidence at the significance level that the mean commute times are different.
The same data can lead to different conclusions at different significance levels. That’s why the level must be chosen before the test, not after.
Example 3: Using a critical value
Section titled “Example 3: Using a critical value”A test at the significance level gives a test statistic of . The critical value is . The hypotheses are : the type of music a student prefers is independent of their age group, and : it is not independent. Write a conclusion.
Solution. Only large values of are in the critical region, and , so the test statistic is not in the critical region.
Do not reject . There is insufficient evidence at the significance level that music preference depends on age group.
Example 4: Spotting a weak conclusion
Section titled “Example 4: Spotting a weak conclusion”A school tested whether a new timetable changed the mean number of late arrivals per day. At the level the p-value was . A student wrote:
”, so we accept . This proves the new timetable makes no difference to late arrivals.”
Explain what’s wrong, and write a better conclusion.
Solution. Three problems:
- It doesn’t compare the p-value with the significance level.
- “Accept ” and “proves” are too strong. A large p-value only means the data are consistent with ; a real (perhaps small) effect could still exist that this sample didn’t detect.
- It should state the significance level.
Better: since , do not reject . There is insufficient evidence at the significance level that the new timetable changed the mean number of late arrivals per day.
Common mistakes
Section titled “Common mistakes”Putting the claim you want to show in H₀. is always the “no difference / no relationship / fits the model” statement, and it contains the equals sign. What you’re looking for evidence of goes in .
Writing “accept H₀”. You either reject or you do not reject it. Not rejecting it doesn’t prove it’s true; it just means the evidence isn’t strong enough.
Mixing up the decision rule. A small p-value is evidence against . Reject when the p-value is less than the significance level. A quick memory aid: “if p is low, must go”.
Choosing one- or two-tailed after looking at the data. The direction comes from the question (“increases”, “different”, “faster”), not from which way the sample means happen to point. Deciding after you’ve seen the data makes a “significant” result too easy to get.
Leaving out the context. “Reject ” on its own doesn’t answer the question. Say what it means for the situation: which means differ, or which variables are related, and at what significance level.
Reading the p-value as “the probability that H₀ is true”. The p-value is the probability of data at least this extreme assuming is true. It is not the probability that is true.
Practice
Section titled “Practice”1. (Warm-up) Write and for each situation, and say whether the test is one-tailed or two-tailed.
- (a) A company wants to know whether the mean battery life of its phones is different from a rival brand’s.
- (b) A coach wants to know whether athletes who stretch have a lower mean injury recovery time than those who don’t.
Solution
(a) Let and be the mean battery lives. , . Two-tailed (“different”).
(b) Let and be the mean recovery times of athletes who stretch and who don’t. , . One-tailed (“lower”).
2. (Warm-up) For each p-value, state whether you would reject at the significance level.
- (a)
- (b)
- (c)
Solution
(a) : reject .
(b) : do not reject .
(c) : reject (only just).
3. (Warm-up) Write the hypotheses, in words, for a test of whether a six-sided die is fair.
Solution
- : the die is fair (each score has probability ).
- : the die is not fair.
4. (Core) A test of : preferred social media app is independent of age group, against : it is not independent, gives . Write a conclusion at
- (a) the significance level
- (b) the significance level.
Solution
(a) , so reject . There is sufficient evidence at the significance level that preferred social media app depends on age group.
(b) , so do not reject . There is insufficient evidence at the significance level that preferred social media app depends on age group.
5. (Core) A test at the significance level gives , with critical value . The hypotheses are : a spinner is fair, and : the spinner is not fair. Write a conclusion.
Solution
, so the test statistic is in the critical region. Reject . There is sufficient evidence at the significance level that the spinner is not fair.
6. (Core) Explain, in context, what means for a test of : a coin is fair, against : the coin is biased.
Solution
If the coin really were fair, there would only be a probability of getting a result at least as far from “half heads” as the one observed. That’s unlikely, so it counts as evidence that the coin is biased (sufficient evidence at the level, but not at the level).
It does not mean there’s a chance the coin is fair.
7. (Core) A dietitian claims a new breakfast cereal lowers mean cholesterol. Her colleague writes .
- (a) Is this the right alternative hypothesis for the claim? Explain.
- (b) Write the correct hypotheses.
Solution
(a) No. The claim has a direction (“lowers”), so the test should be one-tailed. is for a two-tailed test, which looks for a change in either direction.
(b) , , where is the mean cholesterol level of people eating each cereal.
8. (Challenge) A two-tailed test of against gives , and the sample mean for is larger than the sample mean for . The test statistic has a symmetric distribution.
- (a) What would the p-value be for the one-tailed test with ?
- (b) Compare the conclusions of the two tests at the level.
- (c) Why must the choice between them be made before the data are collected?
Solution
(a) The two-tailed p-value counts both tails equally, so the one-tailed p-value (in the direction the data point) is half of it: .
(b) Two-tailed: , do not reject ; insufficient evidence that the means differ. One-tailed: , reject ; sufficient evidence that .
(c) If you looked at the data first and then picked the one-tailed test in the direction the data point, you would be halving the p-value after the fact. That makes it too easy to get a “significant” result just by chance, so the test would no longer have the significance level it claims.
9. (Challenge, enrichment — not an AI SL exam skill) A coin is tossed times and lands heads times. Find the p-value for testing : the coin is fair, against : the coin is biased toward heads, and state the conclusion at the significance level.
Solution
If the coin is fair, the number of heads is . The p-value is the probability of a result at least as extreme as heads:
, so do not reject . There is insufficient evidence at the significance level that the coin is biased toward heads. (Compare heads, where was enough.)