Spearman's Rank Correlation
Pearson’s correlation coefficient tells you how close points are to a straight line. But lots of real relationships aren’t straight: they just keep going up (or keep going down). Spearman’s rank correlation coefficient measures exactly that. It works with the ranks of the data instead of the values, so it also works for data that only come as rankings, like judges’ placings.
Key ideas
Section titled “Key ideas”Ranking data
Section titled “Ranking data”To rank a set of values, put them in order and number them . You can give rank to the largest value or to the smallest, as long as you rank both variables the same way.
For example, ranking the test marks from the highest down gives:
| Mark | ||||
|---|---|---|---|---|
| Rank |
Tied ranks
Section titled “Tied ranks”If two or more values are equal, give each of them the average of the ranks they would have taken. For the marks (highest first), the two s would take ranks and , so each gets
and the next value, , still gets rank . Three equal values that would take ranks , and each get .
What Spearman’s coefficient is
Section titled “What Spearman’s coefficient is”Spearman’s rank correlation coefficient is Pearson’s correlation coefficient calculated on the ranks instead of on the original values. Like , it is always between and :
- : as one variable goes up, the other always goes up (the ranks match exactly).
- : as one variable goes up, the other always goes down (the ranks are exactly reversed).
- near : no consistent increasing or decreasing pattern.
A relationship where one variable always goes up (or always goes down) as the other goes up is called monotonic. So measures the strength of a monotonic relationship, while measures the strength of a linear one. Use the same words to describe the size of as for (strong, moderate, weak; positive or negative).
Finding Spearman’s coefficient with technology
Section titled “Finding Spearman’s coefficient with technology”In IB exams you find with your GDC:
- Rank each variable (averaging tied ranks).
- Enter the two lists of ranks.
- Use your GDC’s two-variable statistics (or linear regression) to find the correlation coefficient of the ranks. That number is .
Some GDCs and spreadsheets can rank the data or find for you, but it’s worth knowing how to rank by hand, especially with ties. You don’t need to know how either coefficient is derived.
Pearson or Spearman?
Section titled “Pearson or Spearman?”| Pearson’s | Spearman’s | |
|---|---|---|
| Measures | how close the points are to a straight line | how consistently one variable rises (or falls) as the other rises |
| Use it when | you want to test for a linear relationship | the relationship might be any monotonic curve, or the data are ranks |
| Data needed | numerical values | values or ranks |
| Effect of outliers | can change a lot | less sensitive |
doesn’t mean the points lie on a line. It means they’re in perfect increasing order. The figure shows data that curve steeply upward: is only , but the ranks match exactly, so .
The effect of an outlier
Section titled “The effect of an outlier”Seven points have a strong linear pattern. Then an eighth point, , is added:
| First seven points | ||
| All eight points |
The outlier pulls down a lot, because it’s far from the line through the other points. But is simply the largest -value, so it gets rank , the same as if it had been . Ranks don’t care how far a value is from the others, only where it sits in the order. That’s why is less sensitive to outliers. (Less sensitive isn’t immune: an outlier that’s out of order, like , changes too.)
As always, a strong correlation of either kind doesn’t prove that one variable causes the other (see correlation and causation).
Worked examples
Section titled “Worked examples”Example 1: Ranking with ties
Section titled “Example 1: Ranking with ties”Seven students scored on a quiz. Rank the scores, giving rank to the highest score.
Solution. Put the scores in order from highest to lowest and number them:
| In order | |||||||
|---|---|---|---|---|---|---|---|
| Position | |||||||
| Rank |
The two s share positions and , so each gets . The two s share positions and , so each gets . In the original order:
| Score | |||||||
|---|---|---|---|---|---|---|---|
| Rank |
Check: the ranks of items always add up to , and . ✓
Example 2: Two judges
Section titled “Example 2: Two judges”Two judges ranked eight dancers, A to H, in a competition.
| Dancer | A | B | C | D | E | F | G | H |
|---|---|---|---|---|---|---|---|---|
| Judge 1 | ||||||||
| Judge 2 |
Find and interpret it.
Solution. The data are already ranks, so enter the two rows as lists and find their correlation coefficient on your GDC:
There is a strong positive agreement between the judges: dancers ranked highly by Judge 1 also tend to be ranked highly by Judge 2. Pearson’s on the original scores isn’t possible here, since only the rankings are given. That’s a situation where Spearman’s coefficient is the natural choice.
Example 3: Ranking raw data first
Section titled “Example 3: Ranking raw data first”The table shows how many hours eight runners trained per week and their times (in minutes) for a km race.
| Runner | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Hours, | ||||||||
| Time, |
- (a) Rank both variables, giving rank to the largest value.
- (b) Find and interpret it in context.
Solution.
(a) For hours, is rank , then , , and the two s share ranks and , so each gets . For times, is rank , then , and the two s share ranks and , so each gets .
| Runner | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Rank of | ||||||||
| Rank of |
(b) Enter the two rows of ranks and find their correlation coefficient:
There is a very strong negative monotonic relationship: runners who train more hours tend to have faster (lower) race times.
Example 4: Choosing a coefficient
Section titled “Example 4: Choosing a coefficient”A new social media account had these follower counts (in thousands) at the end of each of its first eight weeks:
| Week, | ||||||||
|---|---|---|---|---|---|---|---|---|
| Followers, |
- (a) Find Pearson’s and Spearman’s .
- (b) Explain why the two values are different, and say which better describes the relationship.
Solution.
(a) On your GDC, the correlation coefficient of and is (to 3 s.f.).
The follower counts increase every week, so their ranks are , exactly the same as the ranks of the weeks. So .
(b) The number of followers always increases, so the relationship is perfectly monotonic, giving . But it isn’t linear: the counts grow faster and faster (the scatter plot curves upward, as in the figure above), so the points don’t lie close to a straight line and is smaller. Spearman’s describes this relationship better. Pearson’s is only the right tool if you’re asking whether the relationship is linear, and here it clearly isn’t.
Common mistakes
Section titled “Common mistakes”Giving tied values different ranks. If two values are equal, they must get the same rank: the average of the positions they share. Ranking two s as and (in whichever order) gives a value of that depends on an arbitrary choice.
Ranking the two variables in opposite directions. If you give rank to the largest but the smallest , the sign of flips. Rank both variables the same way.
Typing the raw data into the GDC instead of the ranks. The correlation coefficient of the raw values is Pearson’s . You only get when you enter the ranks.
Forgetting to skip ranks after a tie. After two values share ranks and (both ), the next value gets rank , not . Check that your ranks add up to .
Reading as “the points lie on a line”. means the points are in perfect increasing order. They could lie on any increasing curve. Only means a perfect straight line.
Saying Spearman’s coefficient ignores outliers. It’s less sensitive than , not immune. An outlier that changes the order of the data still changes .
Practice
Section titled “Practice”1. (Warm-up) Rank these values, giving rank to the largest: .
Solution
In order: . The three s share positions , and , so each gets .
| Value | |||||||
|---|---|---|---|---|---|---|---|
| Rank |
Check: . ✓
2. (Warm-up) Describe what each value tells you about the relationship between two variables.
- (a)
- (b)
- (c)
Solution
(a) A very strong positive monotonic relationship: as one variable increases, the other almost always increases too.
(b) A perfect negative monotonic relationship: as one variable increases, the other always decreases (the ranks are exactly reversed).
(c) Very weak or no monotonic relationship: there’s no consistent increasing or decreasing pattern.
3. (Core) Two food critics ranked six restaurants.
| Restaurant | P | Q | R | S | T | U |
|---|---|---|---|---|---|---|
| Critic 1 | ||||||
| Critic 2 |
Find and comment on how well the critics agree.
Solution
Enter the two rows of ranks into your GDC and find their correlation coefficient:
There is fairly strong positive agreement: restaurants that one critic ranks highly tend to be ranked highly by the other, though not exactly (restaurant S is the biggest disagreement).
4. (Core) A café recorded the average daily temperature and the number of hot chocolates sold on seven winter days.
| Temperature (°C) | |||||||
|---|---|---|---|---|---|---|---|
| Hot chocolates |
- (a) Rank both variables, giving rank to the largest value.
- (b) Find and interpret it in context.
Solution
(a) Temperatures in order: . The two s share ranks and , so each gets . Sales in order: . The two s share ranks and , so each gets .
| Rank of temperature | |||||||
|---|---|---|---|---|---|---|---|
| Rank of sales |
(b) Using the ranks on a GDC:
There is a very strong negative monotonic relationship: the colder the day, the more hot chocolates the café tends to sell.
5. (Core) For a set of data, and .
- (a) What does tell you about the data?
- (b) Why might be smaller than here?
- (c) Would a straight-line model be a good choice? Explain.
Solution
(a) The data are perfectly monotonic increasing: whenever increases, increases too, so the ranks of and match exactly.
(b) The points follow an increasing curve rather than a straight line, so they aren’t as close to a line as they could be. only measures closeness to a straight line.
(c) Probably not. Since but is noticeably less than , the relationship is increasing but curved, so a curved (non-linear) model would likely fit better. Looking at the scatter plot would confirm this.
6. (Core) The table shows six points.
- (a) Find and .
- (b) The point is added. Find and for all seven points.
- (c) Comment on the effect of the new point on each coefficient.
Solution
(a) On a GDC, (to 3 s.f.). The -values increase whenever increases, so the ranks match and .
(b) With included, and (to 3 s.f.).
(c) The outlier nearly destroys Pearson’s , which drops from to . Spearman’s also drops, from to , because the new point is out of order (a small with the largest ). But it drops by less: is less sensitive to outliers than , though not immune to them.
7. (Core) A website shows each hotel’s star rating ( to ) and its average nightly price.
| Hotel | A | B | C | D | E | F | G | H |
|---|---|---|---|---|---|---|---|---|
| Stars | ||||||||
| Price ($) |
- (a) Explain why Spearman’s coefficient is more appropriate than Pearson’s here.
- (b) Find , giving rank to the largest value of each variable.
Solution
(a) Star ratings are ordinal: a -star hotel is better than a -star hotel, but not necessarily “twice as good”, and the gaps between ratings aren’t equal amounts of anything. Spearman’s coefficient only uses the order, so it suits this kind of data. (There are also many tied values.)
(b) Stars in order: , so the s get , the s get , the s get , then and . Prices in order: .
| Hotel | A | B | C | D | E | F | G | H |
|---|---|---|---|---|---|---|---|---|
| Rank of stars | ||||||||
| Rank of price |
Using the ranks on a GDC:
There is a very strong positive monotonic relationship: hotels with more stars tend to cost more.
8. (Challenge) The table shows a population of bacteria (in thousands) after hours.
- (a) Find and for and .
- (b) Find and for and .
- (c) Explain why didn’t change but did.
Solution
(a) (to 3 s.f.). The -values always increase, so .
(b) For and : (to 4 s.f.), which is to 3 s.f., and .
(c) Taking keeps the values in the same order (a bigger always gives a bigger ), so the ranks don’t change and neither does . But it changes the shape of the scatter plot: grows roughly exponentially, so against is almost a straight line, and gets much closer to .
9. (Challenge) Five data points are , , , and .
- (a) Find correctly, using average ranks for the tie.
- (b) Amir ranks the two s as and instead (largest first), giving rank to the point . Bea does the same but gives rank to . Find the value of that each of them gets.
- (c) Explain why averaging tied ranks is the fair method.
Solution
(a) Ranking from the largest: ranks are and ranks are . On a GDC:
(b) Amir’s ranks are , which give . Bea’s are , which give .
(c) The two s are equal, so there’s no reason to put one ahead of the other. Breaking the tie by choice gives different answers ( or ) from the same data. Averaging treats the tied values equally and gives one answer, , which happens to lie between the two.