Chi-Squared Tests
A die lands on more often than you’d expect. In a survey, older students seem to choose different transport from younger ones. Is that real, or just chance? Chi-squared () tests answer questions like these for counts (frequencies). They compare what you observed with what you’d expect if were true, and measure how big the gap is. This page uses the testing steps from introduction to hypothesis testing.
Key ideas
Section titled “Key ideas”Observed and expected frequencies
Section titled “Observed and expected frequencies”- The observed frequencies are the counts in your data.
- The expected frequencies are the counts you’d expect, on average, if were true.
Expected frequencies don’t have to be whole numbers. If you roll a fair die times, you expect of each score.
The chi-squared statistic
Section titled “The chi-squared statistic”The statistic adds up the squared gaps between observed and expected, each divided by the expected frequency:
- If the data match well, every gap is small and is close to .
- The bigger the gaps, the bigger . So only a large is evidence against , and tests are always upper-tail (one-tailed) tests.
In exams you find and the p-value with your GDC. Working out a few expected values and terms by hand, as in the examples below, helps you understand what the GDC is doing.
Degrees of freedom and the critical value
Section titled “Degrees of freedom and the critical value”The degrees of freedom count how many frequencies are “free to vary” once the totals are fixed. The distribution you compare with depends on . The critical value is the point with an area equal to the significance level to its right. In IB exam questions it’s given to you when you need it.
The goodness-of-fit test
Section titled “The goodness-of-fit test”A goodness-of-fit test checks whether data fit a claimed distribution.
- : the data fit the distribution (for example, “the die is fair” or “the colours are in the ratio ”).
- : the data do not fit the distribution.
For each category, the expected frequency is
With categories, the degrees of freedom are
(once you know of the counts, the last one is fixed by the total). On your GDC, enter the observed and expected frequencies in two lists and use the goodness-of-fit test with degrees of freedom.
The test for independence
Section titled “The test for independence”A test for independence checks whether two categorical variables in a contingency table are related.
- : the two variables are independent.
- : the two variables are not independent.
If the variables are independent, the expected frequency in each cell is
and for a table with rows and columns (not counting the totals),
On your GDC, enter the observed frequencies as a matrix and use the test. It gives , the p-value and the degrees of freedom, and it stores the matrix of expected frequencies so you can check them.
When the test is reliable
Section titled “When the test is reliable”The test is unreliable when expected frequencies are small. Every expected frequency should be greater than (IB exam questions will always satisfy this). If some expected frequencies are or less, combine categories that make sense together (such as ” pets” and ” or more pets”), then recalculate. Combining categories reduces the degrees of freedom, so recompute too. Note that it’s the expected frequencies that matter, not the observed ones.
Worked examples
Section titled “Worked examples”Example 1: Is the die fair?
Section titled “Example 1: Is the die fair?”A die is rolled times.
| Score | ||||||
|---|---|---|---|---|---|---|
| Frequency |
Test at the significance level whether the die is fair. The critical value is .
Solution.
: the die is fair. : the die is not fair.
If the die is fair, each score has probability , so each expected frequency is (all greater than ).
| Score | ||||||
|---|---|---|---|---|---|---|
There are categories, so . A GDC gives (to 3 s.f.).
(equivalently, ), so do not reject . There is insufficient evidence at the significance level that the die is not fair.
Example 2: Given proportions
Section titled “Example 2: Given proportions”A sweet company says its bags contain red, green, yellow and purple sweets. A student counts sweets.
| Colour | Red | Green | Yellow | Purple |
|---|---|---|---|---|
| Observed |
Test the company’s claim at the significance level.
Solution.
: the colours are in the proportions , , , . : they are not in these proportions.
Expected frequencies: , , and . All are greater than .
, and a GDC gives (to 3 s.f.).
, so reject . There is sufficient evidence at the significance level that the colours are not in the proportions the company claims. (Notice that it’s not significant at the level, since .)
Example 3: A test for independence
Section titled “Example 3: A test for independence”A school surveyed students about how they usually get to school.
| Bus | Walk or cycle | Car | Total | |
|---|---|---|---|---|
| Grade 10 | ||||
| Grade 11 | ||||
| Grade 12 | ||||
| Total |
Test at the significance level whether the way students get to school is independent of grade. The critical value is .
Solution.
: way of getting to school is independent of grade. : way of getting to school is not independent of grade.
Expected frequencies, using row total × column total ÷ grand total. For example, for Grade 10 and Bus:
| Bus | Walk or cycle | Car | |
|---|---|---|---|
| Grade 10 | |||
| Grade 11 | |||
| Grade 12 |
All are greater than . (Check: each row of expected values adds to the same row total as the data, e.g. .)
Degrees of freedom: .
With the observed table entered as a matrix, the GDC’s test gives
(and ), so reject . There is sufficient evidence at the significance level that the way students get to school depends on their grade. Looking at the table, Grade 12 students travel by car much more than expected ( observed, expected), which makes sense: many of them can drive.
Example 4: Combining categories
Section titled “Example 4: Combining categories”A store manager believes that of customers pay by card, by phone, with cash and with a gift card. In a sample of customers:
| Payment | Card | Phone | Cash | Gift card |
|---|---|---|---|---|
| Observed |
Test the manager’s belief at the significance level.
Solution.
: payment methods are in the proportions , , , . : they are not.
Expected frequencies: , , and .
The gift-card expected frequency is , which is too small. Combine the two smallest categories into “cash or gift card”: observed , expected .
| Payment | Card | Phone | Cash or gift card |
|---|---|---|---|
There are now categories, so (not ). A GDC gives (to 3 s.f.).
, so do not reject . There is insufficient evidence at the significance level that the payment methods differ from the manager’s proportions.
Common mistakes
Section titled “Common mistakes”Using the formula on percentages instead of counts. must be calculated from frequencies (counts). If the data are given as percentages, convert them to counts first; the same percentages from a bigger sample give a bigger .
Getting the degrees of freedom wrong. Goodness of fit: , where is the number of categories. Independence: . Don’t count the “Total” row or column, and recount after combining categories.
Checking the observed frequencies instead of the expected ones. The “greater than ” condition is about the expected frequencies. A small observed count is fine.
Treating a small χ² as evidence against H₀. A small means the data fit well. Only a large (beyond the critical value, or with a small p-value) leads to rejecting .
Concluding that one variable causes the other. A test for independence can show that two variables are related, not why. In Example 3, being in Grade 12 doesn’t “cause” car travel; having a licence is a likely explanation.
Writing “accept H₀” or leaving out the context. Say “do not reject ” and explain what that means for the situation, with the significance level.
Practice
Section titled “Practice”1. (Warm-up) Find the number of degrees of freedom for
- (a) a test for independence on a table with rows and columns
- (b) a goodness-of-fit test with categories
- (c) a test for independence on a table with rows and columns.
Solution
(a)
(b)
(c)
2. (Warm-up) In a contingency table with grand total , a cell is in a row with total and a column with total . Find the expected frequency for that cell under the hypothesis of independence.
Solution
3. (Core) A school recorded the number of student absences on each weekday over a term.
| Day | Mon | Tue | Wed | Thu | Fri |
|---|---|---|---|---|---|
| Absences |
- (a) Test at the significance level whether absences are equally likely on every weekday.
- (b) Would your conclusion change at the level?
Solution
(a) : absences are uniformly distributed over the five weekdays. : they are not.
The total is , so each expected frequency is .
, and a GDC gives (to 3 s.f.).
, so do not reject . There is insufficient evidence at the significance level that absences are not equally likely on every weekday.
(b) Yes: , so at the level you would reject and conclude there is sufficient evidence that absences are not equally likely on every weekday.
4. (Core) A garden centre says its seeds produce red, pink and white flowers in the ratio . Of plants grown, were red, pink and white. Test the claim at the significance level. The critical value is .
Solution
: the colours are in the ratio . : they are not.
The probabilities are , so the expected frequencies are , and .
( to 3 s.f.).
, so do not reject . There is insufficient evidence at the significance level that the colours are not in the ratio .
5. (Core) Students at two schools were asked which lunch option they prefer.
| Hot meal | Sandwich | Salad | Total | |
|---|---|---|---|---|
| School X | ||||
| School Y | ||||
| Total |
- (a) Find the expected frequencies.
- (b) Test at the significance level whether lunch preference is independent of school. The critical value is .
Solution
(a) Using row total × column total ÷ :
| Hot meal | Sandwich | Salad | |
|---|---|---|---|
| School X | |||
| School Y |
(For example, .) All are greater than .
(b) : lunch preference is independent of school. : lunch preference is not independent of school.
. A GDC gives and (to 3 s.f.).
, so do not reject . There is insufficient evidence at the significance level that lunch preference depends on school.
6. (Core) A cinema chain surveyed people about how they most often watch films.
| Streaming | Cinema | Broadcast TV | Disc | Total | |
|---|---|---|---|---|---|
| Age 13–24 | |||||
| Age 25–49 | |||||
| Age 50+ | |||||
| Total |
Test at the significance level whether the way people watch films is independent of age group.
Solution
: the way people watch films is independent of age group. : it is not independent of age group.
The smallest expected frequency is , so all are greater than .
. A GDC gives
, so reject . There is sufficient evidence at the significance level that the way people watch films depends on age group. (For example, the youngest group streams more than expected and the oldest group watches more broadcast TV.)
7. (Core) A contingency table has row totals , and , column totals , and , and grand total .
- (a) Find the expected frequencies for the first row.
- (b) Find the expected frequencies for the other two rows. How many of the nine expected frequencies did you really need to calculate with the formula?
Solution
(a) , , .
(b)
| Col 1 | Col 2 | Col 3 | Total | |
|---|---|---|---|---|
| Row 1 | ||||
| Row 2 | ||||
| Row 3 | ||||
| Total |
Only were really needed: once the top-left block is known, the rest follow from the row and column totals (e.g. ). That’s exactly why .
8. (Challenge) A survey of households recorded where they live and how many pets they have.
| 0 pets | 1 pet | 2 pets | 3 or more | Total | |
|---|---|---|---|---|---|
| Urban | |||||
| Suburban | |||||
| Rural | |||||
| Total |
- (a) Show that a test on this table isn’t reliable.
- (b) Combine suitable columns and test at the significance level whether number of pets is independent of where a household lives. The critical value is .
Solution
(a) The “2 pets” column has expected frequencies , and , all or less.
(b) Combine “2 pets” and “3 or more” into “2 or more”:
| 0 pets | 1 pet | 2 or more | Total | |
|---|---|---|---|---|
| Urban | ||||
| Suburban | ||||
| Rural | ||||
| Total |
The smallest expected frequency is now , so all are greater than .
: number of pets is independent of where a household lives. : it is not independent.
. A GDC gives and (to 3 s.f.).
, so do not reject . There is insufficient evidence at the significance level that number of pets depends on where a household lives.
9. (Challenge) A spinner with three equal sections is spun times, landing on A times, B times and C times.
- (a) Test at the significance level whether the spinner is fair. The critical value is .
- (b) A second experiment with spins gives exactly three times as many of each: , and . Repeat the test.
- (c) Explain why the conclusions differ even though the proportions are the same.
Solution
(a) : the spinner is fair. : it is not. Each expected frequency is .
. (), so do not reject : insufficient evidence at the level that the spinner is unfair.
(b) Each expected frequency is .
(), so reject : sufficient evidence at the level that the spinner is not fair.
(c) Multiplying every count by multiplies each by and each by , so is multiplied by . With more data, the same proportions are much less likely to be down to chance. A larger sample makes the test better at detecting a real difference.