Skip to content
Family Table Math
Auto

Chi-Squared Tests

A die lands on 44 more often than you’d expect. In a survey, older students seem to choose different transport from younger ones. Is that real, or just chance? Chi-squared (χ2\chi^2) tests answer questions like these for counts (frequencies). They compare what you observed with what you’d expect if H0H_0 were true, and measure how big the gap is. This page uses the testing steps from introduction to hypothesis testing.

  • The observed frequencies fof_o are the counts in your data.
  • The expected frequencies fef_e are the counts you’d expect, on average, if H0H_0 were true.

Expected frequencies don’t have to be whole numbers. If you roll a fair die 5050 times, you expect 506=8.33\dfrac{50}{6} = 8.33 of each score.

The χ2\chi^2 statistic adds up the squared gaps between observed and expected, each divided by the expected frequency:

χcalc2=∑(fo−fe)2fe\chi^2_{\text{calc}} = \sum \frac{(f_o - f_e)^2}{f_e}
  • If the data match H0H_0 well, every gap is small and χ2\chi^2 is close to 00.
  • The bigger the gaps, the bigger χ2\chi^2. So only a large χ2\chi^2 is evidence against H0H_0, and χ2\chi^2 tests are always upper-tail (one-tailed) tests.

In exams you find χ2\chi^2 and the p-value with your GDC. Working out a few expected values and terms by hand, as in the examples below, helps you understand what the GDC is doing.

The degrees of freedom ν\nu count how many frequencies are “free to vary” once the totals are fixed. The χ2\chi^2 distribution you compare with depends on ν\nu. The critical value is the point with an area equal to the significance level to its right. In IB exam questions it’s given to you when you need it.

The chi-squared distribution with 4 degrees of freedom, a right-skewed curve. The area to the right of the critical value 9.488 is shaded: it is 0.05, the critical region. The test statistic 17.1 lies inside the critical region. 0 2 4 6 8 10 12 14 16 18 20 critical value 9.488 critical region area 0.05 (reject H₀) χ2 = 17.1 4 degrees of freedom
With 44 degrees of freedom, the 5%5\% critical value is 9.4889.488. A test statistic of 17.117.1 is in the critical region, so H0H_0 is rejected.

A goodness-of-fit test checks whether data fit a claimed distribution.

  • H0H_0: the data fit the distribution (for example, “the die is fair” or “the colours are in the ratio 40:30:20:1040 : 30 : 20 : 10”).
  • H1H_1: the data do not fit the distribution.

For each category, the expected frequency is

fe=(total number of observations)×(probability of that category under H0)f_e = (\text{total number of observations}) \times (\text{probability of that category under } H_0)

With nn categories, the degrees of freedom are

ν=n−1\nu = n - 1

(once you know n−1n - 1 of the counts, the last one is fixed by the total). On your GDC, enter the observed and expected frequencies in two lists and use the χ2\chi^2 goodness-of-fit test with n−1n - 1 degrees of freedom.

A test for independence checks whether two categorical variables in a contingency table are related.

  • H0H_0: the two variables are independent.
  • H1H_1: the two variables are not independent.

If the variables are independent, the expected frequency in each cell is

fe=row total×column totalgrand totalf_e = \frac{\text{row total} \times \text{column total}}{\text{grand total}}

and for a table with rr rows and cc columns (not counting the totals),

ν=(r−1)(c−1)\nu = (r - 1)(c - 1)

On your GDC, enter the observed frequencies as a matrix and use the χ2\chi^2 test. It gives χ2\chi^2, the p-value and the degrees of freedom, and it stores the matrix of expected frequencies so you can check them.

The χ2\chi^2 test is unreliable when expected frequencies are small. Every expected frequency should be greater than 55 (IB exam questions will always satisfy this). If some expected frequencies are 55 or less, combine categories that make sense together (such as ”22 pets” and ”33 or more pets”), then recalculate. Combining categories reduces the degrees of freedom, so recompute ν\nu too. Note that it’s the expected frequencies that matter, not the observed ones.

A die is rolled 120120 times.

Score112233445566
Frequency141422221818272716162323

Test at the 5%5\% significance level whether the die is fair. The critical value is 11.07011.070.

Solution.

H0H_0: the die is fair. H1H_1: the die is not fair.

If the die is fair, each score has probability 16\dfrac{1}{6}, so each expected frequency is 120×16=20120 \times \dfrac{1}{6} = 20 (all greater than 55).

Score112233445566
fof_o141422221818272716162323
fef_e202020202020202020202020
(fo−fe)2fe\dfrac{(f_o - f_e)^2}{f_e}1.81.80.20.20.20.22.452.450.80.80.450.45
χcalc2=1.8+0.2+0.2+2.45+0.8+0.45=5.9\chi^2_{\text{calc}} = 1.8 + 0.2 + 0.2 + 2.45 + 0.8 + 0.45 = 5.9

There are 66 categories, so ν=6−1=5\nu = 6 - 1 = 5. A GDC gives p=0.316p = 0.316 (to 3 s.f.).

5.9<11.0705.9 \lt 11.070 (equivalently, 0.316>0.050.316 \gt 0.05), so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the die is not fair.

A sweet company says its bags contain 40%40\% red, 30%30\% green, 20%20\% yellow and 10%10\% purple sweets. A student counts 200200 sweets.

ColourRedGreenYellowPurple
Observed6262727238382828

Test the company’s claim at the 5%5\% significance level.

Solution.

H0H_0: the colours are in the proportions 40%40\%, 30%30\%, 20%20\%, 10%10\%. H1H_1: they are not in these proportions.

Expected frequencies: 200×0.4=80200 \times 0.4 = 80, 200×0.3=60200 \times 0.3 = 60, 200×0.2=40200 \times 0.2 = 40 and 200×0.1=20200 \times 0.1 = 20. All are greater than 55.

χcalc2=(62−80)280+(72−60)260+(38−40)240+(28−20)220=4.05+2.4+0.1+3.2=9.75\chi^2_{\text{calc}} = \frac{(62 - 80)^2}{80} + \frac{(72 - 60)^2}{60} + \frac{(38 - 40)^2}{40} + \frac{(28 - 20)^2}{20} = 4.05 + 2.4 + 0.1 + 3.2 = 9.75

ν=4−1=3\nu = 4 - 1 = 3, and a GDC gives p=0.0208p = 0.0208 (to 3 s.f.).

0.0208<0.050.0208 \lt 0.05, so reject H0H_0. There is sufficient evidence at the 5%5\% significance level that the colours are not in the proportions the company claims. (Notice that it’s not significant at the 1%1\% level, since 0.0208>0.010.0208 \gt 0.01.)

A school surveyed 250250 students about how they usually get to school.

BusWalk or cycleCarTotal
Grade 104242282820209090
Grade 113030252525258080
Grade 121818222240408080
Total909075758585250250

Test at the 1%1\% significance level whether the way students get to school is independent of grade. The critical value is 13.27713.277.

Solution.

H0H_0: way of getting to school is independent of grade. H1H_1: way of getting to school is not independent of grade.

Expected frequencies, using row total × column total ÷ grand total. For example, for Grade 10 and Bus:

fe=90×90250=32.4f_e = \frac{90 \times 90}{250} = 32.4
BusWalk or cycleCar
Grade 1032.432.4272730.630.6
Grade 1128.828.8242427.227.2
Grade 1228.828.8242427.227.2

All are greater than 55. (Check: each row of expected values adds to the same row total as the data, e.g. 32.4+27+30.6=9032.4 + 27 + 30.6 = 90.)

Degrees of freedom: ν=(3−1)(3−1)=4\nu = (3 - 1)(3 - 1) = 4.

With the observed table entered as a matrix, the GDC’s χ2\chi^2 test gives

χcalc2=17.1,p=0.00188 (to 3 s.f.)\chi^2_{\text{calc}} = 17.1, \qquad p = 0.00188 \text{ (to 3 s.f.)}

17.1>13.27717.1 \gt 13.277 (and 0.00188<0.010.00188 \lt 0.01), so reject H0H_0. There is sufficient evidence at the 1%1\% significance level that the way students get to school depends on their grade. Looking at the table, Grade 12 students travel by car much more than expected (4040 observed, 27.227.2 expected), which makes sense: many of them can drive.

A store manager believes that 50%50\% of customers pay by card, 30%30\% by phone, 15%15\% with cash and 5%5\% with a gift card. In a sample of 8080 customers:

PaymentCardPhoneCashGift card
Observed32322929131366

Test the manager’s belief at the 10%10\% significance level.

Solution.

H0H_0: payment methods are in the proportions 50%50\%, 30%30\%, 15%15\%, 5%5\%. H1H_1: they are not.

Expected frequencies: 80×0.5=4080 \times 0.5 = 40, 80×0.3=2480 \times 0.3 = 24, 80×0.15=1280 \times 0.15 = 12 and 80×0.05=480 \times 0.05 = 4.

The gift-card expected frequency is 44, which is too small. Combine the two smallest categories into “cash or gift card”: observed 13+6=1913 + 6 = 19, expected 12+4=1612 + 4 = 16.

PaymentCardPhoneCash or gift card
fof_o323229291919
fef_e404024241616
χcalc2=(32−40)240+(29−24)224+(19−16)216=1.6+1.0417+0.5625=3.20 (to 3 s.f.)\chi^2_{\text{calc}} = \frac{(32 - 40)^2}{40} + \frac{(29 - 24)^2}{24} + \frac{(19 - 16)^2}{16} = 1.6 + 1.0417 + 0.5625 = 3.20 \text{ (to 3 s.f.)}

There are now 33 categories, so ν=3−1=2\nu = 3 - 1 = 2 (not 33). A GDC gives p=0.201p = 0.201 (to 3 s.f.).

0.201>0.100.201 \gt 0.10, so do not reject H0H_0. There is insufficient evidence at the 10%10\% significance level that the payment methods differ from the manager’s proportions.

Using the formula on percentages instead of counts. χ2\chi^2 must be calculated from frequencies (counts). If the data are given as percentages, convert them to counts first; the same percentages from a bigger sample give a bigger χ2\chi^2.

Getting the degrees of freedom wrong. Goodness of fit: ν=n−1\nu = n - 1, where nn is the number of categories. Independence: ν=(r−1)(c−1)\nu = (r - 1)(c - 1). Don’t count the “Total” row or column, and recount after combining categories.

Checking the observed frequencies instead of the expected ones. The “greater than 55” condition is about the expected frequencies. A small observed count is fine.

Treating a small χ² as evidence against H₀. A small χ2\chi^2 means the data fit H0H_0 well. Only a large χ2\chi^2 (beyond the critical value, or with a small p-value) leads to rejecting H0H_0.

Concluding that one variable causes the other. A test for independence can show that two variables are related, not why. In Example 3, being in Grade 12 doesn’t “cause” car travel; having a licence is a likely explanation.

Writing “accept H₀” or leaving out the context. Say “do not reject H0H_0” and explain what that means for the situation, with the significance level.

1. (Warm-up) Find the number of degrees of freedom for

  • (a) a test for independence on a table with 33 rows and 44 columns
  • (b) a goodness-of-fit test with 55 categories
  • (c) a test for independence on a table with 22 rows and 44 columns.
Solution

(a) ν=(3−1)(4−1)=6\nu = (3 - 1)(4 - 1) = 6

(b) ν=5−1=4\nu = 5 - 1 = 4

(c) ν=(2−1)(4−1)=3\nu = (2 - 1)(4 - 1) = 3

2. (Warm-up) In a contingency table with grand total 300300, a cell is in a row with total 120120 and a column with total 8585. Find the expected frequency for that cell under the hypothesis of independence.

Solutionfe=120×85300=34f_e = \frac{120 \times 85}{300} = 34

3. (Core) A school recorded the number of student absences on each weekday over a term.

DayMonTueWedThuFri
Absences52523838353541415959
  • (a) Test at the 5%5\% significance level whether absences are equally likely on every weekday.
  • (b) Would your conclusion change at the 10%10\% level?
Solution

(a) H0H_0: absences are uniformly distributed over the five weekdays. H1H_1: they are not.

The total is 225225, so each expected frequency is 2255=45\dfrac{225}{5} = 45.

χcalc2=49+49+100+16+19645=41045=9.11 (to 3 s.f.)\chi^2_{\text{calc}} = \frac{49 + 49 + 100 + 16 + 196}{45} = \frac{410}{45} = 9.11 \text{ (to 3 s.f.)}

ν=5−1=4\nu = 5 - 1 = 4, and a GDC gives p=0.0584p = 0.0584 (to 3 s.f.).

0.0584>0.050.0584 \gt 0.05, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that absences are not equally likely on every weekday.

(b) Yes: 0.0584<0.100.0584 \lt 0.10, so at the 10%10\% level you would reject H0H_0 and conclude there is sufficient evidence that absences are not equally likely on every weekday.

4. (Core) A garden centre says its seeds produce red, pink and white flowers in the ratio 1:2:11 : 2 : 1. Of 160160 plants grown, 3131 were red, 9292 pink and 3737 white. Test the claim at the 5%5\% significance level. The critical value is 5.9915.991.

Solution

H0H_0: the colours are in the ratio 1:2:11 : 2 : 1. H1H_1: they are not.

The probabilities are 14,24,14\dfrac{1}{4}, \dfrac{2}{4}, \dfrac{1}{4}, so the expected frequencies are 4040, 8080 and 4040.

χcalc2=(31−40)240+(92−80)280+(37−40)240=2.025+1.8+0.225=4.05\chi^2_{\text{calc}} = \frac{(31 - 40)^2}{40} + \frac{(92 - 80)^2}{80} + \frac{(37 - 40)^2}{40} = 2.025 + 1.8 + 0.225 = 4.05

ν=3−1=2\nu = 3 - 1 = 2 (p=0.132p = 0.132 to 3 s.f.).

4.05<5.9914.05 \lt 5.991, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that the colours are not in the ratio 1:2:11 : 2 : 1.

5. (Core) Students at two schools were asked which lunch option they prefer.

Hot mealSandwichSaladTotal
School X3434262620208080
School Y262644443030100100
Total606070705050180180
  • (a) Find the expected frequencies.
  • (b) Test at the 5%5\% significance level whether lunch preference is independent of school. The critical value is 5.9915.991.
Solution

(a) Using row total × column total ÷ 180180:

Hot mealSandwichSalad
School X26.726.731.131.122.222.2
School Y33.333.338.938.927.827.8

(For example, 80×60180=26.6‾\dfrac{80 \times 60}{180} = 26.\overline{6}.) All are greater than 55.

(b) H0H_0: lunch preference is independent of school. H1H_1: lunch preference is not independent of school.

ν=(2−1)(3−1)=2\nu = (2 - 1)(3 - 1) = 2. A GDC gives χcalc2=5.54\chi^2_{\text{calc}} = 5.54 and p=0.0626p = 0.0626 (to 3 s.f.).

5.54<5.9915.54 \lt 5.991, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that lunch preference depends on school.

6. (Core) A cinema chain surveyed 340340 people about how they most often watch films.

StreamingCinemaBroadcast TVDiscTotal
Age 13–244545303015151010100100
Age 25–493535404025252020120120
Age 50+2020303040403030120120
Total10010010010080806060340340

Test at the 1%1\% significance level whether the way people watch films is independent of age group.

Solution

H0H_0: the way people watch films is independent of age group. H1H_1: it is not independent of age group.

The smallest expected frequency is 100×60340=17.6\dfrac{100 \times 60}{340} = 17.6, so all are greater than 55.

ν=(3−1)(4−1)=6\nu = (3 - 1)(4 - 1) = 6. A GDC gives

χcalc2=31.7,p=1.83×10−5 (to 3 s.f.)\chi^2_{\text{calc}} = 31.7, \qquad p = 1.83 \times 10^{-5} \text{ (to 3 s.f.)}

1.83×10−5<0.011.83 \times 10^{-5} \lt 0.01, so reject H0H_0. There is sufficient evidence at the 1%1\% significance level that the way people watch films depends on age group. (For example, the youngest group streams more than expected and the oldest group watches more broadcast TV.)

7. (Core) A contingency table has row totals 5050, 7070 and 8080, column totals 6060, 9090 and 5050, and grand total 200200.

  • (a) Find the expected frequencies for the first row.
  • (b) Find the expected frequencies for the other two rows. How many of the nine expected frequencies did you really need to calculate with the formula?
Solution

(a) 50×60200=15\dfrac{50 \times 60}{200} = 15, 50×90200=22.5\dfrac{50 \times 90}{200} = 22.5, 50×50200=12.5\dfrac{50 \times 50}{200} = 12.5.

(b)

Col 1Col 2Col 3Total
Row 1151522.522.512.512.55050
Row 2212131.531.517.517.57070
Row 32424363620208080
Total606090905050200200

Only 44 were really needed: once the top-left 2×22 \times 2 block is known, the rest follow from the row and column totals (e.g. 12.5=50−15−22.512.5 = 50 - 15 - 22.5). That’s exactly why ν=(3−1)(3−1)=4\nu = (3 - 1)(3 - 1) = 4.

8. (Challenge) A survey of 180180 households recorded where they live and how many pets they have.

0 pets1 pet2 pets3 or moreTotal
Urban1818222244665050
Suburban252528283314147070
Rural171730303310106060
Total6060808010103030180180
  • (a) Show that a χ2\chi^2 test on this table isn’t reliable.
  • (b) Combine suitable columns and test at the 5%5\% significance level whether number of pets is independent of where a household lives. The critical value is 9.4889.488.
Solution

(a) The “2 pets” column has expected frequencies 50×10180=2.78\dfrac{50 \times 10}{180} = 2.78, 70×10180=3.89\dfrac{70 \times 10}{180} = 3.89 and 60×10180=3.33\dfrac{60 \times 10}{180} = 3.33, all 55 or less.

(b) Combine “2 pets” and “3 or more” into “2 or more”:

0 pets1 pet2 or moreTotal
Urban1818222210105050
Suburban2525282817177070
Rural1717303013136060
Total606080804040180180

The smallest expected frequency is now 50×40180=11.1\dfrac{50 \times 40}{180} = 11.1, so all are greater than 55.

H0H_0: number of pets is independent of where a household lives. H1H_1: it is not independent.

ν=(3−1)(3−1)=4\nu = (3 - 1)(3 - 1) = 4. A GDC gives χcalc2=1.66\chi^2_{\text{calc}} = 1.66 and p=0.798p = 0.798 (to 3 s.f.).

1.66<9.4881.66 \lt 9.488, so do not reject H0H_0. There is insufficient evidence at the 5%5\% significance level that number of pets depends on where a household lives.

9. (Challenge) A spinner with three equal sections is spun 6060 times, landing on A 2626 times, B 2020 times and C 1414 times.

  • (a) Test at the 5%5\% significance level whether the spinner is fair. The critical value is 5.9915.991.
  • (b) A second experiment with 180180 spins gives exactly three times as many of each: 7878, 6060 and 4242. Repeat the test.
  • (c) Explain why the conclusions differ even though the proportions are the same.
Solution

(a) H0H_0: the spinner is fair. H1H_1: it is not. Each expected frequency is 2020.

χcalc2=36+0+3620=3.6\chi^2_{\text{calc}} = \frac{36 + 0 + 36}{20} = 3.6

ν=2\nu = 2. 3.6<5.9913.6 \lt 5.991 (p=0.165p = 0.165), so do not reject H0H_0: insufficient evidence at the 5%5\% level that the spinner is unfair.

(b) Each expected frequency is 6060.

χcalc2=324+0+32460=10.8\chi^2_{\text{calc}} = \frac{324 + 0 + 324}{60} = 10.8

10.8>5.99110.8 \gt 5.991 (p=0.00452p = 0.00452), so reject H0H_0: sufficient evidence at the 5%5\% level that the spinner is not fair.

(c) Multiplying every count by 33 multiplies each (fo−fe)2(f_o - f_e)^2 by 99 and each fef_e by 33, so χ2\chi^2 is multiplied by 33. With more data, the same proportions are much less likely to be down to chance. A larger sample makes the test better at detecting a real difference.