Hypothesis Tests for Parameters
In hypothesis testing you met the general structure: state and , find a p-value, compare it with the significance level, and conclude in context. This page applies that structure to parameters: a population mean, a proportion, and a Poisson rate. You’ll also learn a second way to decide, using a critical region, and how to calculate the chance of each kind of wrong conclusion.
Key ideas
Section titled “Key ideas”Critical values and critical regions
Section titled “Critical values and critical regions”The critical region is the set of values of the test statistic that lead you to reject . The boundary of the region is the critical value. You can decide a test either way:
- p-value method: reject if the p-value is less than the significance level.
- Critical region method: reject if the observed value lies in the critical region.
The two methods always give the same conclusion. The critical region method is useful because you can work it out before collecting data, and you need it to find error probabilities.
Testing a mean: σ known (normal distribution)
Section titled “Testing a mean: σ known (normal distribution)”If the population is normal with known (or is large, by the central limit theorem), then under ,
For a one-tailed test with at the level, the critical value of is , so the critical region is , or equivalently . For a two-tailed test at , split the : reject if or . A GDC’s z-test function gives and the p-value directly.
Testing a mean: σ unknown (t-test)
Section titled “Testing a mean: σ unknown (t-test)”If is unknown, use a one-sample t-test with and degrees of freedom, as in the t-test topic. Use your GDC’s t-test function to get the test statistic and p-value. As with confidence intervals: σ known → normal; σ unknown → t, regardless of the sample size. You won’t be asked to calculate critical regions for t-tests.
Paired samples. When each individual is measured twice (before and after, left and right), work out the differences for each pair and do a one-sample test on them, usually with . A matched-pairs test is just a single-sample test on the differences.
Testing a proportion with the binomial distribution
Section titled “Testing a proportion with the binomial distribution”To test a claimed probability , count the successes in trials. Under , . These tests are one-tailed only:
- : the critical region is (small values).
- : the critical region is (large values).
Because is discrete, you usually can’t hit the significance level exactly. The IB rule: choose the critical region that makes as large as possible while still less than the significance level. List cumulative probabilities with the GDC to find it.
Testing a mean with the Poisson distribution
Section titled “Testing a mean with the Poisson distribution”To test a claimed Poisson rate, count the events in the observed interval. Under , , where is the mean for the whole interval observed (see the Poisson distribution). These tests are also one-tailed only, and the critical region is found the same way as for the binomial.
Testing whether ρ = 0
Section titled “Testing whether ρ = 0”For bivariate normal data, you can test whether there’s any linear correlation in the population. The population correlation coefficient is ; the sample value is (see scatter plots and correlation).
- (no linear correlation in the population).
- , or , or (decide before looking at the data).
Enter the data in two lists and use the GDC’s linear regression t-test (for example LinRegTTest), choosing the alternative hypothesis. It gives and the p-value. In exams, the data will be given.
Type I and Type II errors
Section titled “Type I and Type II errors”| true | false | |
|---|---|---|
| Reject | Type I error | correct |
| Don’t reject | correct | Type II error |
- Type I error: rejecting when it is true. For a continuous test statistic (normal), the significance level. For a binomial or Poisson test, , which is a bit less than the significance level.
- Type II error: not rejecting when it is false. To calculate it, you need a specific true value of the parameter: .
Lowering the significance level makes Type I errors less likely but Type II errors more likely. Taking a bigger sample reduces without changing the significance level.
Worked examples
Section titled “Worked examples”Example 1: A z-test with a critical region
Section titled “Example 1: A z-test with a critical region”A machine fills bags of flour with masses that are normally distributed with mean g and standard deviation g. After a repair, the manager suspects that the mean has increased. A random sample of bags has a mean mass of g. Assume the standard deviation is still g.
- (a) Find the critical region for for a test at the level.
- (b) Carry out the test.
- (c) Find the probability of a Type II error if the true mean is actually g.
Solution.
(a) and . Under , , with standard deviation .
The critical value has . Inverse normal with area , , :
The critical region is g (to 2 d.p.).
(b) is in the critical region, so reject . There is significant evidence at the level that the mean mass has increased.
Check with the p-value: (3 s.f.), and . Same conclusion.
(c) A Type II error means is not in the critical region even though . Now :
Example 2: A test for a proportion
Section titled “Example 2: A test for a proportion”A website says that of students at a school walk to school. A student thinks the true proportion is lower. She asks a random sample of students and of them walk.
- (a) Find the critical region for a test at the level, and state the probability of a Type I error.
- (b) Carry out the test.
Solution.
(a) Let be the number who walk. and . Under , . The critical region is ; list the cumulative probabilities:
but , so the critical region is . The probability of a Type I error is (3 s.f.).
(b) The observed value is not in the critical region, so don’t reject . There is insufficient evidence at the level that fewer than of students walk to school. (The p-value is , which agrees.)
Example 3: A test for a Poisson mean
Section titled “Example 3: A test for a Poisson mean”A help desk has historically received calls at an average rate of per hour. After an advertising campaign, the manager thinks the rate has increased. In a -hour period, calls are received.
- (a) Find the critical region for a test at the level, and carry out the test.
- (b) Find the probability of a Type I error.
- (c) The true rate is actually calls per hour. Find the probability of a Type II error.
Solution.
(a) Let be the number of calls in hours. (that is, per hour) and . Under , , and the critical region is .
The critical region is . Since , reject : there is significant evidence at the level that the call rate has increased.
(b) (3 s.f.)
(c) At calls per hour, the mean for hours is , so . A Type II error means is not in the critical region:
Even with a higher rate, this test misses the change more than half the time, because hours is a short period. A longer observation period would reduce this.
Example 4: Testing ρ = 0 with technology
Section titled “Example 4: Testing ρ = 0 with technology”A teacher records the number of hours randomly chosen students spent studying and their test scores. The data can be assumed to come from a bivariate normal distribution.
| Hours, | ||||||||
|---|---|---|---|---|---|---|---|---|
| Score, |
Test at the level whether there is a positive correlation between study time and score.
Solution. and .
GDC linear regression t-test (lists for and , alternative ): (3 s.f.) and p-value (3 s.f.).
Since , reject . There is significant evidence at the level of a positive correlation between hours of study and test score in the population. (This is evidence of an association, not proof that studying causes higher scores.)
Common mistakes
Section titled “Common mistakes”Using the wrong mean in a Poisson test. If the rate is per hour and you observe for hours, is . Rescale before finding the critical region.
Choosing a discrete critical region that’s too big. The probability of the critical region must be less than the significance level. For at , has probability , which is too large.
Saying P(Type I) is always the significance level. That’s only true for continuous tests. For binomial and Poisson tests, is the actual probability of the critical region under , such as .
Finding a Type II error under H0. uses the true (alternative) parameter value, applied to the complement of the critical region. Find the critical region using first, then switch distributions.
Using a z-test when σ is unknown. If the standard deviation comes from the sample, use a t-test, even for large samples. For paired data, test the differences, not the two columns separately.
Concluding “H0 is true”. If you don’t reject , say there is insufficient evidence for . The data don’t prove .
Practice
Section titled “Practice”1. (Warm-up) A food inspector tests : “the restaurant’s kitchen meets safety standards” against : “it doesn’t”. Describe a Type I error and a Type II error in context.
Solution
Type I error: concluding that the kitchen doesn’t meet the standards when it actually does (rejecting a true ).
Type II error: concluding there’s not enough evidence of a problem when the kitchen actually fails the standards (not rejecting a false ).
2. (Warm-up) A population is normal with . A test of against uses a sample of at the level. Find the critical region for .
Solution
Under , , since . The critical value has . Inverse normal:
The critical region is (3 s.f.).
3. (Warm-up) A coin is tossed times to test against , where is the probability of heads, at the level. Find the critical region and the probability of a Type I error.
Solution
Under , . and . The critical region is , and (3 s.f.).
4. (Core) Bolts are made with lengths that are normally distributed with standard deviation mm. The target mean is mm. A sample of bolts has mean mm. Test at the level whether the mean length differs from mm.
Solution
and (two-tailed). is known, so use a z-test:
The critical values are . Since , is in the critical region, so reject . (p-value .) There is significant evidence at the level that the mean bolt length is not mm.
5. (Core) Eight adults had their systolic blood pressure (mmHg) measured before and after a -week exercise program.
| Adult | A | B | C | D | E | F | G | H |
|---|---|---|---|---|---|---|---|---|
| Before | ||||||||
| After |
Assuming the differences are normally distributed, test at the level whether the program reduces blood pressure.
Solution
This is paired data, so use the differences :
and . is unknown, so use a one-sample t-test on the differences (, , degrees of freedom).
GDC t-test: and p-value (3 s.f.).
Since , reject . There is significant evidence at the level that the program reduces systolic blood pressure on average.
6. (Core) A highway had an average of accidents per month. After a new speed limit, there were accidents in the next months. Test at the level whether the accident rate has decreased, and state the probability of a Type I error.
Solution
Let be the number of accidents in months. and . Under , .
and , so the critical region is .
The observed value is not in the critical region, so don’t reject . There is insufficient evidence at the level that the accident rate has decreased.
(3 s.f.).
7. (Core) A biologist measures the water temperature (°C) and a fish activity score at sites. The data can be assumed to be bivariate normal.
- (a) Test against at the level.
- (b) Another biologist had decided before collecting the data to test . What would her p-value and conclusion be?
Solution
(a) GDC linear regression t-test (): and p-value (3 s.f.). Since , don’t reject : there is insufficient evidence at the level of a linear correlation.
(b) With , the p-value is half as big: (3 s.f.). Since , she would reject and conclude there is significant evidence of a positive correlation. This is why the alternative hypothesis must be chosen before seeing the data.
8. (Challenge) In question 3, suppose the coin is actually biased with . Find the probability of a Type II error.
Solution
The critical region is . A Type II error means when :
With only tosses, the test misses this bias more often than not.
9. (Challenge) IQ-style test scores are normally distributed with . A researcher tests against at the level.
- (a) With a sample of , find the critical region for and the probability of a Type II error if the true mean is .
- (b) Find the smallest sample size for which the probability of a Type II error is less than when the true mean is .
Solution
(a) Under , , since . Critical value: , so the critical region is (to 1 d.p.).
If : with , which is (3 s.f.).
(b) The critical value is . We need when , so must be more than standard deviations below :
The smallest sample size is .