Skip to content
Family Table Math
Auto

Measures of Central Tendency

A measure of central tendency is a single number that describes the “middle” or “typical” value of a data set. The mean, median, and mode each describe the centre in a different way, and they don’t always agree. Knowing which one to use (and when a number is hiding something) is one of the most useful skills in statistics.

The mean is the sum of the values divided by how many there are. Statisticians use different symbols for a population and a sample:

μ=∑xNxˉ=∑xn\mu = \frac{\sum x}{N} \qquad\qquad \bar{x} = \frac{\sum x}{n}
SymbolMeaning
μ\mu (mu)mean of a whole population
xˉ\bar{x} (x-bar)mean of a sample
NNpopulation size
nnsample size
∑x\sum xthe sum of all the data values

The formula is the same; only the symbols change. The mean uses every value, which makes it sensitive to extreme values.

The median is the middle value when the data are in order.

  • If nn is odd, the median is the middle value: the one in position n+12\dfrac{n + 1}{2}.
  • If nn is even, the median is the mean of the two middle values.

The mode is the value that occurs most often. A data set can have one mode, more than one mode, or no mode (if every value occurs once). The mode is the only measure you can use for categorical data, like favourite colours.

When some values count more than others, use a weighted mean. Multiply each value by its weight, add, and divide by the total weight:

weighted mean=∑wx∑w\text{weighted mean} = \frac{\sum w x}{\sum w}

If the weights are percentages that add to 100%100\%, just multiply each value by its weight as a decimal and add.

When data are given in a frequency table with intervals, you don’t know the exact values. Estimate the mean by assuming every value in an interval sits at the interval’s midpoint mm:

xˉ≈∑fm∑f\bar{x} \approx \frac{\sum f m}{\sum f}

where ff is the frequency of each interval. The median interval is the interval that contains the middle value.

An outlier is a value far away from the rest of the data. Outliers pull the mean toward them, but they barely move the median. (A precise rule for spotting outliers is on the quartiles and percentiles page.)

MeasureBest when…Weakness
Meandata are numerical and roughly symmetric, with no outlierspulled by outliers
Mediandata are skewed or have outliers (incomes, house prices)ignores how far the other values are from the middle
Modedata are categorical, or you want the most common value (shoe sizes to stock)may not exist, or may not be near the centre

The centre is only half the story. Two data sets can have the same mean but very different spreads. To measure spread, see standard deviation.

In Desmos, type a list and use mean(...) and median(...), or name it first (L = [...], then mean(L)). For a frequency table, you don’t have to type every repeated value: multiply each value by its frequency and divide by the total count, e.g. values 2,3,42, 3, 4 with frequencies 3,5,23, 5, 2 give 6+15+810=2.9\dfrac{6 + 15 + 8}{10} = 2.9. Many SAT questions ask how the mean or median changes when a value is added or removed, or when there’s an outlier, and those are faster to reason out: an outlier pulls the mean toward it but barely moves the median. See using Desmos on the SAT.

A sample of 88 students reported how many hours they slept last night:

7, 8, 6, 9, 7, 8, 7, 107,\ 8,\ 6,\ 9,\ 7,\ 8,\ 7,\ 10
  • (a) Find the mean, median, and mode.
  • (b) A ninth student reports 22 hours. How does each measure change?

Solution. (a) The sum is 6262, so

xˉ=628=7.75 h\bar{x} = \frac{62}{8} = 7.75 \text{ h}

In order: 6,7,7,7,8,8,9,106, 7, 7, 7, 8, 8, 9, 10. With n=8n = 8 (even), the median is the mean of the 4th and 5th values:

median=7+82=7.5 h\text{median} = \frac{7 + 8}{2} = 7.5 \text{ h}

The value 77 occurs three times, so the mode is 77 h.

(b) Now n=9n = 9 and the sum is 6464, so xˉ=649≈7.11\bar{x} = \dfrac{64}{9} \approx 7.11 h. In order: 2,6,7,7,7,8,8,9,102, 6, 7, 7, 7, 8, 8, 9, 10. The median is the 5th value, 77 h. The mode is still 77 h.

The one low value dragged the mean down by about 0.640.64 h, but the median only moved by 0.50.5 h and the mode didn’t move at all.

Mei’s final mark is calculated like this: tests are worth 45%45\%, assignments 25%25\%, and the final exam 30%30\%. Her averages are 76%76\% on tests, 88%88\% on assignments, and 70%70\% on the exam. Find her final mark.

Solution. The weights add to 100%100\%, so multiply each mark by its weight as a decimal:

final mark=0.45(76)+0.25(88)+0.30(70)=34.2+22+21=77.2\begin{aligned} \text{final mark} &= 0.45(76) + 0.25(88) + 0.30(70) \\ &= 34.2 + 22 + 21 \\ &= 77.2 \end{aligned}

Her final mark is 77.2%77.2\%. Notice that a simple mean of the three marks, 76+88+703≈78.0\dfrac{76 + 88 + 70}{3} \approx 78.0, gives the wrong answer because it treats the three parts as equally important.

A class of 3030 students recorded their commute times to school.

Time (min)Frequency ffMidpoint mmfmf m
00 to under 101066553030
1010 to under 202011111515165165
2020 to under 3030882525200200
3030 to under 4040443535140140
4040 to under 50501145454545
Total3030580580

Estimate the mean commute time, and find the median interval.

Solution.

xˉ≈∑fm∑f=58030≈19.3 min\bar{x} \approx \frac{\sum f m}{\sum f} = \frac{580}{30} \approx 19.3 \text{ min}

With 3030 values, the median is between the 15th and 16th values. The first interval holds the first 66 values, and the second holds values 77 to 1717. So the median interval is 1010 to under 2020 minutes.

This is only an estimate of the mean, because we assumed every value is at the midpoint of its interval.

Two players on a school basketball team scored these points in their last 77 games:

  • Aisha: 12,14,15,15,16,17,1812, 14, 15, 15, 16, 17, 18
  • Maya: 4,6,8,10,11,30,364, 6, 8, 10, 11, 30, 36

Compare their scoring using the mean and median. Who is the more reliable scorer?

Solution.

MeanMedian
Aisha1077≈15.3\dfrac{107}{7} \approx 15.31515
Maya1057=15\dfrac{105}{7} = 151010

Their means are almost the same, so the means alone suggest they’re equally good scorers. The medians tell a different story: in a typical game, Aisha scores about 1515 points but Maya scores about 1010. Maya’s mean is pulled up by two big games (3030 and 3636).

Aisha is the more reliable scorer: her scores are close together, while Maya’s swing widely. When you compare data sets, compare both a measure of centre and a measure of spread.

Finding the median without sorting first. The median is the middle of the ordered list. Always put the data in order before you look for the middle.

Taking the middle position as the median value. For n=9n = 9, the median is the 5th value, not 55. The formula n+12\dfrac{n + 1}{2} gives a position, not the answer.

Taking a simple mean when the parts have different weights. Course marks, GPA, and averages of groups of different sizes all need a weighted mean.

Using interval endpoints instead of midpoints for grouped data. Use the midpoint of each interval, and remember that the result is an estimate, not an exact mean.

Reporting the mean when there are outliers. One very large salary can make an “average” salary misleading. For skewed data or data with outliers, the median usually describes a typical value better.

Comparing data sets with only one number. Two sets with the same mean can be very different. Look at the median and the spread too.

1. (Warm-up) Find the mean, median, and mode: 4,9,6,4,7,10,24, 9, 6, 4, 7, 10, 2.

Solution

In order: 2,4,4,6,7,9,102, 4, 4, 6, 7, 9, 10. The sum is 4242, so the mean is 427=6\dfrac{42}{7} = 6. The median is the 4th value, 66. The mode is 44.

2. (Warm-up) Find the mean and median: 13,18,11,22,15,1613, 18, 11, 22, 15, 16.

Solution

The sum is 9595, so the mean is 956≈15.8\dfrac{95}{6} \approx 15.8.

In order: 11,13,15,16,18,2211, 13, 15, 16, 18, 22. The median is 15+162=15.5\dfrac{15 + 16}{2} = 15.5.

3. (Warm-up) Which measure of central tendency (mean, median, or mode) is best in each situation? Explain briefly.

  • (a) A store decides which shoe size to order the most of.
  • (b) A survey asks students for their favourite season.
  • (c) A real estate agent describes a typical house price in a town with a few mansions.
Solution

(a) Mode: the store wants the size that sells most often.

(b) Mode: the data are categorical, so the mean and median don’t make sense.

(c) Median: the mansions are outliers that would pull the mean up.

4. (Core) In a course, term work is worth 70%70\% and the final exam is worth 30%30\%. Jordan’s term mark is 84%84\%. What mark does Jordan need on the exam to finish with at least 80%80\%?

Solution

Let xx be the exam mark:

0.70(84)+0.30x≥8058.8+0.30x≥800.30x≥21.2x≥70.6‾\begin{aligned} 0.70(84) + 0.30x &\ge 80 \\ 58.8 + 0.30x &\ge 80 \\ 0.30x &\ge 21.2 \\ x &\ge 70.\overline{6} \end{aligned}

Jordan needs about 71%71\% on the exam.

5. (Core) Forty students recorded how long they spent on homework last night.

Time (min)00 to under 30303030 to under 60606060 to under 90909090 to under 120120
Frequency551212151588
  • (a) Estimate the mean time.
  • (b) Which interval contains the median?
Solution

(a) The midpoints are 15,45,75,10515, 45, 75, 105:

xˉ≈5(15)+12(45)+15(75)+8(105)40=258040=64.5 min\bar{x} \approx \frac{5(15) + 12(45) + 15(75) + 8(105)}{40} = \frac{2580}{40} = 64.5 \text{ min}

(b) The median is between the 20th and 21st values. The first two intervals hold 5+12=175 + 12 = 17 values, and the third holds values 1818 to 3232. The median interval is 6060 to under 9090 minutes.

6. (Core) A runner’s times (in seconds) for six 100100 m sprints are 13.2,13.5,13.8,14.0,14.1,21.613.2, 13.5, 13.8, 14.0, 14.1, 21.6. In the last sprint, she tripped.

  • (a) Find the mean and median of all six times.
  • (b) Find the mean and median without the 21.621.6.
  • (c) Which measure better describes her typical time? Why?
Solution

(a) Mean: 90.26≈15.03\dfrac{90.2}{6} \approx 15.03 s. Median: 13.8+14.02=13.9\dfrac{13.8 + 14.0}{2} = 13.9 s.

(b) Mean: 68.65=13.72\dfrac{68.6}{5} = 13.72 s. Median: 13.813.8 s.

(c) The median. The outlier raised the mean by more than 1.31.3 s, but changed the median by only 0.10.1 s. A mean of 15.0315.03 s is slower than five of her six runs.

7. (Core) Two brands of batteries were tested in the same flashlight. Lifetimes in hours:

  • Brand A: 48,50,51,52,52,53,5448, 50, 51, 52, 52, 53, 54
  • Brand B: 40,45,52,55,58,60,6140, 45, 52, 55, 58, 60, 61

Find the mean and median for each brand. Which brand would you buy, and why?

Solution

Brand A: mean 3607≈51.4\dfrac{360}{7} \approx 51.4 h, median 5252 h.

Brand B: mean 3717=53\dfrac{371}{7} = 53 h, median 5555 h.

Brand B lasts longer on average, but its lifetimes are much more spread out (4040 to 6161 h, compared with 4848 to 5454 h for A). If you want the longest life, choose B. If you need batteries you can count on (for example, for an emergency kit), A is more predictable. Either answer is fine with a good reason.

8. (Challenge) One class of 2424 students has a mean test mark of 72%72\%. Another class of 3030 students has a mean of 81%81\%. Find the mean mark of all 5454 students. Why isn’t it 76.5%76.5\%?

Solution

Find each class’s total, then divide by the total number of students:

μ=24(72)+30(81)54=1728+243054=415854=77%\mu = \frac{24(72) + 30(81)}{54} = \frac{1728 + 2430}{54} = \frac{4158}{54} = 77\%

This is a weighted mean with the class sizes as weights. The answer isn’t 72+812=76.5%\dfrac{72 + 81}{2} = 76.5\% because the second class has more students, so it pulls the combined mean toward 81%81\%.

9. (Challenge) Make up a set of five positive whole numbers with a mean of 1010, a median of 88, and a mode of 66.

Solution

The median is 88, so the middle (3rd) value is 88. The mode is 66, so 66 must appear at least twice, and it has to be below the median: the list starts 6,6,86, 6, 8. The mean is 1010, so the sum is 5050, and the last two values add to 50−20=3050 - 20 = 30. They must both be greater than 88 and different from each other (so 66 stays the only mode).

One answer: 6,6,8,12,186, 6, 8, 12, 18. Check: mean 505=10\dfrac{50}{5} = 10, median 88, mode 66. ✓ (Other answers, such as 6,6,8,14,166, 6, 8, 14, 16, also work.)