Skip to content
Family Table Math
Auto

Hypothesis Tests for Parameters

In hypothesis testing you met the general structure: state H0H_0 and H1H_1, find a p-value, compare it with the significance level, and conclude in context. This page applies that structure to parameters: a population mean, a proportion, and a Poisson rate. You’ll also learn a second way to decide, using a critical region, and how to calculate the chance of each kind of wrong conclusion.

The critical region is the set of values of the test statistic that lead you to reject H0H_0. The boundary of the region is the critical value. You can decide a test either way:

  • p-value method: reject H0H_0 if the p-value is less than the significance level.
  • Critical region method: reject H0H_0 if the observed value lies in the critical region.

The two methods always give the same conclusion. The critical region method is useful because you can work it out before collecting data, and you need it to find error probabilities.

Testing a mean: σ known (normal distribution)

Section titled “Testing a mean: σ known (normal distribution)”

If the population is normal with known σ\sigma (or nn is large, by the central limit theorem), then under H0:μ=μ0H_0: \mu = \mu_0,

Xˉ∼N(μ0, σ2n)andz=xˉ−μ0σ/n\bar{X} \sim N\left(\mu_0,\ \frac{\sigma^2}{n}\right) \qquad\text{and}\qquad z = \frac{\bar{x} - \mu_0}{\sigma/\sqrt{n}}

For a one-tailed test with H1:μ>μ0H_1: \mu \gt \mu_0 at the 5%5\% level, the critical value of zz is 1.6451.645, so the critical region is z>1.645z \gt 1.645, or equivalently xˉ>μ0+1.645σn\bar{x} \gt \mu_0 + 1.645\dfrac{\sigma}{\sqrt{n}}. For a two-tailed test at 5%5\%, split the 5%5\%: reject if z<−1.960z \lt -1.960 or z>1.960z \gt 1.960. A GDC’s z-test function gives zz and the p-value directly.

If σ\sigma is unknown, use a one-sample t-test with sn−1s_{n-1} and n−1n - 1 degrees of freedom, as in the t-test topic. Use your GDC’s t-test function to get the test statistic and p-value. As with confidence intervals: σ known → normal; σ unknown → t, regardless of the sample size. You won’t be asked to calculate critical regions for t-tests.

Paired samples. When each individual is measured twice (before and after, left and right), work out the differences dd for each pair and do a one-sample test on them, usually with H0:μd=0H_0: \mu_d = 0. A matched-pairs test is just a single-sample test on the differences.

Testing a proportion with the binomial distribution

Section titled “Testing a proportion with the binomial distribution”

To test a claimed probability pp, count the successes XX in nn trials. Under H0H_0, X∼B(n,p0)X \sim B(n, p_0). These tests are one-tailed only:

  • H1:p<p0H_1: p \lt p_0: the critical region is X≤kX \le k (small values).
  • H1:p>p0H_1: p \gt p_0: the critical region is X≥kX \ge k (large values).

Because XX is discrete, you usually can’t hit the significance level exactly. The IB rule: choose the critical region that makes P(X in critical region)P(X \text{ in critical region}) as large as possible while still less than the significance level. List cumulative probabilities with the GDC to find it.

Testing a mean with the Poisson distribution

Section titled “Testing a mean with the Poisson distribution”

To test a claimed Poisson rate, count the events XX in the observed interval. Under H0H_0, X∼Po(m0)X \sim \mathrm{Po}(m_0), where m0m_0 is the mean for the whole interval observed (see the Poisson distribution). These tests are also one-tailed only, and the critical region is found the same way as for the binomial.

For bivariate normal data, you can test whether there’s any linear correlation in the population. The population correlation coefficient is ρ\rho; the sample value is rr (see scatter plots and correlation).

  • H0:ρ=0H_0: \rho = 0 (no linear correlation in the population).
  • H1:ρ≠0H_1: \rho \ne 0, or ρ>0\rho \gt 0, or ρ<0\rho \lt 0 (decide before looking at the data).

Enter the data in two lists and use the GDC’s linear regression t-test (for example LinRegTTest), choosing the alternative hypothesis. It gives rr and the p-value. In exams, the data will be given.

H0H_0 trueH0H_0 false
Reject H0H_0Type I errorcorrect
Don’t reject H0H_0correctType II error
  • Type I error: rejecting H0H_0 when it is true. For a continuous test statistic (normal), P(Type I)=P(\text{Type I}) = the significance level. For a binomial or Poisson test, P(Type I)=P(X in critical region∣H0)P(\text{Type I}) = P(X \text{ in critical region} \mid H_0), which is a bit less than the significance level.
  • Type II error: not rejecting H0H_0 when it is false. To calculate it, you need a specific true value of the parameter: P(Type II)=P(X not in critical region∣that true value)P(\text{Type II}) = P(X \text{ not in critical region} \mid \text{that true value}).

Lowering the significance level makes Type I errors less likely but Type II errors more likely. Taking a bigger sample reduces P(Type II)P(\text{Type II}) without changing the significance level.

Critical region for the sample mean with Type I and Type II error areas 496 498 500 502 504 506 508 c = 502.47 H₀: μ = 500 true μ = 504 Type I Type II ← don't reject H₀ reject H₀ →
Example 1: the critical region is xˉ>502.47\bar{x} \gt 502.47. Orange: P(Type I)=0.05P(\text{Type I}) = 0.05 if μ=500\mu = 500. Blue: P(Type II)≈0.153P(\text{Type II}) \approx 0.153 if actually μ=504\mu = 504.

Example 1: A z-test with a critical region

Section titled “Example 1: A z-test with a critical region”

A machine fills bags of flour with masses that are normally distributed with mean 500500 g and standard deviation 66 g. After a repair, the manager suspects that the mean has increased. A random sample of 1616 bags has a mean mass of 503.1503.1 g. Assume the standard deviation is still 66 g.

  • (a) Find the critical region for xˉ\bar{x} for a test at the 5%5\% level.
  • (b) Carry out the test.
  • (c) Find the probability of a Type II error if the true mean is actually 504504 g.

Solution.

(a) H0:μ=500H_0: \mu = 500 and H1:μ>500H_1: \mu \gt 500. Under H0H_0, Xˉ∼N(500,6216)\bar{X} \sim N\left(500, \dfrac{6^2}{16}\right), with standard deviation 64=1.5\dfrac{6}{4} = 1.5.

The critical value cc has P(Xˉ>c)=0.05P(\bar{X} \gt c) = 0.05. Inverse normal with area 0.950.95, μ=500\mu = 500, σ=1.5\sigma = 1.5:

c=502.467…c = 502.467\ldots

The critical region is xˉ>502.47\bar{x} \gt 502.47 g (to 2 d.p.).

(b) 503.1503.1 is in the critical region, so reject H0H_0. There is significant evidence at the 5%5\% level that the mean mass has increased.

Check with the p-value: z=503.1−5001.5=2.07z = \dfrac{503.1 - 500}{1.5} = 2.07 (3 s.f.), and P(Xˉ≥503.1)=0.0194<0.05P(\bar{X} \ge 503.1) = 0.0194 \lt 0.05. Same conclusion.

(c) A Type II error means xˉ\bar{x} is not in the critical region even though μ=504\mu = 504. Now Xˉ∼N(504,1.52)\bar{X} \sim N(504, 1.5^2):

P(Type II)=P(Xˉ≤502.467…)=0.153 (3 s.f.)P(\text{Type II}) = P(\bar{X} \le 502.467\ldots) = 0.153 \text{ (3 s.f.)}

A website says that 30%30\% of students at a school walk to school. A student thinks the true proportion is lower. She asks a random sample of 2525 students and 44 of them walk.

  • (a) Find the critical region for a test at the 5%5\% level, and state the probability of a Type I error.
  • (b) Carry out the test.

Solution.

(a) Let XX be the number who walk. H0:p=0.3H_0: p = 0.3 and H1:p<0.3H_1: p \lt 0.3. Under H0H_0, X∼B(25,0.3)X \sim B(25, 0.3). The critical region is X≤kX \le k; list the cumulative probabilities:

kk223344
P(X≤k)P(X \le k)0.008960.008960.03320.03320.09050.0905

P(X≤3)=0.0332<0.05P(X \le 3) = 0.0332 \lt 0.05 but P(X≤4)=0.0905>0.05P(X \le 4) = 0.0905 \gt 0.05, so the critical region is X≤3X \le 3. The probability of a Type I error is P(X≤3)=0.0332P(X \le 3) = 0.0332 (3 s.f.).

Binomial distribution B(25, 0.3) with the critical region X at most 3 highlighted 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 0.05 0.1 0.15 0.2 critical region X ≤ 3: 0.0332 x (number who walk) P(X = x)
Under H0H_0, X∼B(25,0.3)X \sim B(25, 0.3). The orange bars (X≤3X \le 3) form the critical region, with total probability 0.0332<0.050.0332 \lt 0.05.

(b) The observed value 44 is not in the critical region, so don’t reject H0H_0. There is insufficient evidence at the 5%5\% level that fewer than 30%30\% of students walk to school. (The p-value is P(X≤4)=0.0905>0.05P(X \le 4) = 0.0905 \gt 0.05, which agrees.)

A help desk has historically received calls at an average rate of 44 per hour. After an advertising campaign, the manager thinks the rate has increased. In a 33-hour period, 2020 calls are received.

  • (a) Find the critical region for a test at the 5%5\% level, and carry out the test.
  • (b) Find the probability of a Type I error.
  • (c) The true rate is actually 66 calls per hour. Find the probability of a Type II error.

Solution.

(a) Let XX be the number of calls in 33 hours. H0:m=12H_0: m = 12 (that is, 44 per hour) and H1:m>12H_1: m \gt 12. Under H0H_0, X∼Po(12)X \sim \mathrm{Po}(12), and the critical region is X≥cX \ge c.

cc181819192020
P(X≥c)P(X \ge c)0.06300.06300.03740.03740.02130.0213

The critical region is X≥19X \ge 19. Since 20≥1920 \ge 19, reject H0H_0: there is significant evidence at the 5%5\% level that the call rate has increased.

(b) P(Type I)=P(X≥19∣m=12)=0.0374P(\text{Type I}) = P(X \ge 19 \mid m = 12) = 0.0374 (3 s.f.)

(c) At 66 calls per hour, the mean for 33 hours is 1818, so X∼Po(18)X \sim \mathrm{Po}(18). A Type II error means XX is not in the critical region:

P(Type II)=P(X≤18∣m=18)=0.562 (3 s.f.)P(\text{Type II}) = P(X \le 18 \mid m = 18) = 0.562 \text{ (3 s.f.)}

Even with a 50%50\% higher rate, this test misses the change more than half the time, because 33 hours is a short period. A longer observation period would reduce this.

A teacher records the number of hours 88 randomly chosen students spent studying and their test scores. The data can be assumed to come from a bivariate normal distribution.

Hours, xx2255118844663377
Score, yy62627070585881816666747469697272

Test at the 1%1\% level whether there is a positive correlation between study time and score.

Solution. H0:ρ=0H_0: \rho = 0 and H1:ρ>0H_1: \rho \gt 0.

GDC linear regression t-test (lists for xx and yy, alternative ρ>0\rho \gt 0): r=0.938r = 0.938 (3 s.f.) and p-value =0.000286= 0.000286 (3 s.f.).

Since 0.000286<0.010.000286 \lt 0.01, reject H0H_0. There is significant evidence at the 1%1\% level of a positive correlation between hours of study and test score in the population. (This is evidence of an association, not proof that studying causes higher scores.)

Using the wrong mean in a Poisson test. If the rate is 44 per hour and you observe for 33 hours, H0H_0 is m=12m = 12. Rescale before finding the critical region.

Choosing a discrete critical region that’s too big. The probability of the critical region must be less than the significance level. For B(25,0.3)B(25, 0.3) at 5%5\%, X≤4X \le 4 has probability 0.09050.0905, which is too large.

Saying P(Type I) is always the significance level. That’s only true for continuous tests. For binomial and Poisson tests, P(Type I)P(\text{Type I}) is the actual probability of the critical region under H0H_0, such as 0.03320.0332.

Finding a Type II error under H0. P(Type II)P(\text{Type II}) uses the true (alternative) parameter value, applied to the complement of the critical region. Find the critical region using H0H_0 first, then switch distributions.

Using a z-test when σ is unknown. If the standard deviation comes from the sample, use a t-test, even for large samples. For paired data, test the differences, not the two columns separately.

Concluding “H0 is true”. If you don’t reject H0H_0, say there is insufficient evidence for H1H_1. The data don’t prove H0H_0.

1. (Warm-up) A food inspector tests H0H_0: “the restaurant’s kitchen meets safety standards” against H1H_1: “it doesn’t”. Describe a Type I error and a Type II error in context.

Solution

Type I error: concluding that the kitchen doesn’t meet the standards when it actually does (rejecting a true H0H_0).

Type II error: concluding there’s not enough evidence of a problem when the kitchen actually fails the standards (not rejecting a false H0H_0).

2. (Warm-up) A population is normal with σ=10\sigma = 10. A test of H0:μ=80H_0: \mu = 80 against H1:μ<80H_1: \mu \lt 80 uses a sample of 2525 at the 1%1\% level. Find the critical region for xˉ\bar{x}.

Solution

Under H0H_0, Xˉ∼N(80,22)\bar{X} \sim N(80, 2^2), since 1025=2\dfrac{10}{\sqrt{25}} = 2. The critical value has P(Xˉ<c)=0.01P(\bar{X} \lt c) = 0.01. Inverse normal:

c=80−2.3263…×2=75.347…c = 80 - 2.3263\ldots \times 2 = 75.347\ldots

The critical region is xˉ<75.3\bar{x} \lt 75.3 (3 s.f.).

3. (Warm-up) A coin is tossed 2020 times to test H0:p=0.5H_0: p = 0.5 against H1:p>0.5H_1: p \gt 0.5, where pp is the probability of heads, at the 5%5\% level. Find the critical region and the probability of a Type I error.

Solution

Under H0H_0, X∼B(20,0.5)X \sim B(20, 0.5). P(X≥14)=0.0577>0.05P(X \ge 14) = 0.0577 \gt 0.05 and P(X≥15)=0.0207<0.05P(X \ge 15) = 0.0207 \lt 0.05. The critical region is X≥15X \ge 15, and P(Type I)=0.0207P(\text{Type I}) = 0.0207 (3 s.f.).

4. (Core) Bolts are made with lengths that are normally distributed with standard deviation 0.40.4 mm. The target mean is 12.012.0 mm. A sample of 1010 bolts has mean 12.2512.25 mm. Test at the 5%5\% level whether the mean length differs from 12.012.0 mm.

Solution

H0:μ=12.0H_0: \mu = 12.0 and H1:μ≠12.0H_1: \mu \ne 12.0 (two-tailed). σ\sigma is known, so use a z-test:

z=12.25−12.00.4/10=1.976…z = \frac{12.25 - 12.0}{0.4/\sqrt{10}} = 1.976\ldots

The critical values are ±1.960\pm 1.960. Since 1.976>1.9601.976 \gt 1.960, zz is in the critical region, so reject H0H_0. (p-value =0.0481<0.05= 0.0481 \lt 0.05.) There is significant evidence at the 5%5\% level that the mean bolt length is not 12.012.0 mm.

5. (Core) Eight adults had their systolic blood pressure (mmHg) measured before and after a 66-week exercise program.

AdultABCDEFGH
Before142142150150138138155155147147160160145145152152
After138138147147139139149149141141155155144144146146

Assuming the differences are normally distributed, test at the 5%5\% level whether the program reduces blood pressure.

Solution

This is paired data, so use the differences d=before−afterd = \text{before} - \text{after}:

4, 3, −1, 6, 6, 5, 1, 64,\ 3,\ -1,\ 6,\ 6,\ 5,\ 1,\ 6

H0:μd=0H_0: \mu_d = 0 and H1:μd>0H_1: \mu_d \gt 0. σ\sigma is unknown, so use a one-sample t-test on the differences (dˉ=3.75\bar{d} = 3.75, sn−1=2.6049…s_{n-1} = 2.6049\ldots, 77 degrees of freedom).

GDC t-test: t=4.07t = 4.07 and p-value =0.00237= 0.00237 (3 s.f.).

Since 0.00237<0.050.00237 \lt 0.05, reject H0H_0. There is significant evidence at the 5%5\% level that the program reduces systolic blood pressure on average.

6. (Core) A highway had an average of 2.52.5 accidents per month. After a new speed limit, there were 33 accidents in the next 33 months. Test at the 5%5\% level whether the accident rate has decreased, and state the probability of a Type I error.

Solution

Let XX be the number of accidents in 33 months. H0:m=7.5H_0: m = 7.5 and H1:m<7.5H_1: m \lt 7.5. Under H0H_0, X∼Po(7.5)X \sim \mathrm{Po}(7.5).

P(X≤2)=0.0203<0.05P(X \le 2) = 0.0203 \lt 0.05 and P(X≤3)=0.0591>0.05P(X \le 3) = 0.0591 \gt 0.05, so the critical region is X≤2X \le 2.

The observed value 33 is not in the critical region, so don’t reject H0H_0. There is insufficient evidence at the 5%5\% level that the accident rate has decreased.

P(Type I)=P(X≤2)=0.0203P(\text{Type I}) = P(X \le 2) = 0.0203 (3 s.f.).

7. (Core) A biologist measures the water temperature xx (°C) and a fish activity score yy at 77 sites. The data can be assumed to be bivariate normal.

xx1.21.22.02.02.82.83.53.54.14.15.05.05.65.6
yy3131292933333030353532323636
  • (a) Test H0:ρ=0H_0: \rho = 0 against H1:ρ≠0H_1: \rho \ne 0 at the 5%5\% level.
  • (b) Another biologist had decided before collecting the data to test H1:ρ>0H_1: \rho \gt 0. What would her p-value and conclusion be?
Solution

(a) GDC linear regression t-test (ρ≠0\rho \ne 0): r=0.673r = 0.673 and p-value =0.0976= 0.0976 (3 s.f.). Since 0.0976>0.050.0976 \gt 0.05, don’t reject H0H_0: there is insufficient evidence at the 5%5\% level of a linear correlation.

(b) With H1:ρ>0H_1: \rho \gt 0, the p-value is half as big: 0.04880.0488 (3 s.f.). Since 0.0488<0.050.0488 \lt 0.05, she would reject H0H_0 and conclude there is significant evidence of a positive correlation. This is why the alternative hypothesis must be chosen before seeing the data.

8. (Challenge) In question 3, suppose the coin is actually biased with p=0.7p = 0.7. Find the probability of a Type II error.

Solution

The critical region is X≥15X \ge 15. A Type II error means X≤14X \le 14 when X∼B(20,0.7)X \sim B(20, 0.7):

P(Type II)=P(X≤14)=0.584 (3 s.f.)P(\text{Type II}) = P(X \le 14) = 0.584 \text{ (3 s.f.)}

With only 2020 tosses, the test misses this bias more often than not.

9. (Challenge) IQ-style test scores are normally distributed with σ=15\sigma = 15. A researcher tests H0:μ=100H_0: \mu = 100 against H1:μ>100H_1: \mu \gt 100 at the 5%5\% level.

  • (a) With a sample of 3636, find the critical region for xˉ\bar{x} and the probability of a Type II error if the true mean is 106106.
  • (b) Find the smallest sample size for which the probability of a Type II error is less than 0.10.1 when the true mean is 106106.
Solution

(a) Under H0H_0, Xˉ∼N(100,2.52)\bar{X} \sim N\left(100, 2.5^2\right), since 1536=2.5\dfrac{15}{\sqrt{36}} = 2.5. Critical value: c=100+1.6448…×2.5=104.11…c = 100 + 1.6448\ldots \times 2.5 = 104.11\ldots, so the critical region is xˉ>104.1\bar{x} \gt 104.1 (to 1 d.p.).

If μ=106\mu = 106: P(Type II)=P(Xˉ≤104.11…)P(\text{Type II}) = P(\bar{X} \le 104.11\ldots) with Xˉ∼N(106,2.52)\bar{X} \sim N(106, 2.5^2), which is 0.2250.225 (3 s.f.).

(b) The critical value is c=100+1.6448…×15nc = 100 + 1.6448\ldots \times \dfrac{15}{\sqrt{n}}. We need P(Xˉ≤c)<0.1P(\bar{X} \le c) \lt 0.1 when μ=106\mu = 106, so cc must be more than 1.2815…1.2815\ldots standard deviations below 106106:

100+1.6448…15n<106−1.2815…15n(1.6448…+1.2815…)15n<6n>7.3159…n>53.52…\begin{aligned} 100 + 1.6448\ldots \frac{15}{\sqrt{n}} &\lt 106 - 1.2815\ldots \frac{15}{\sqrt{n}} \\ (1.6448\ldots + 1.2815\ldots)\frac{15}{\sqrt{n}} &\lt 6 \\ \sqrt{n} &\gt 7.3159\ldots \\ n &\gt 53.52\ldots \end{aligned}

The smallest sample size is n=54n = 54.