Skip to content
Family Table Math
Auto

Spearman's Rank Correlation

Pearson’s correlation coefficient rr tells you how close points are to a straight line. But lots of real relationships aren’t straight: they just keep going up (or keep going down). Spearman’s rank correlation coefficient rsr_s measures exactly that. It works with the ranks of the data instead of the values, so it also works for data that only come as rankings, like judges’ placings.

To rank a set of values, put them in order and number them 1,2,3,…1, 2, 3, \ldots. You can give rank 11 to the largest value or to the smallest, as long as you rank both variables the same way.

For example, ranking the test marks 72,91,64,8572, 91, 64, 85 from the highest down gives:

Mark7272919164648585
Rank33114422

If two or more values are equal, give each of them the average of the ranks they would have taken. For the marks 80,75,75,6080, 75, 75, 60 (highest first), the two 7575s would take ranks 22 and 33, so each gets

2+32=2.5\frac{2 + 3}{2} = 2.5

and the next value, 6060, still gets rank 44. Three equal values that would take ranks 44, 55 and 66 each get 4+5+63=5\dfrac{4 + 5 + 6}{3} = 5.

Spearman’s rank correlation coefficient rsr_s is Pearson’s correlation coefficient calculated on the ranks instead of on the original values. Like rr, it is always between −1-1 and 11:

  • rs=1r_s = 1: as one variable goes up, the other always goes up (the ranks match exactly).
  • rs=−1r_s = -1: as one variable goes up, the other always goes down (the ranks are exactly reversed).
  • rsr_s near 00: no consistent increasing or decreasing pattern.

A relationship where one variable always goes up (or always goes down) as the other goes up is called monotonic. So rsr_s measures the strength of a monotonic relationship, while rr measures the strength of a linear one. Use the same words to describe the size of rsr_s as for rr (strong, moderate, weak; positive or negative).

Finding Spearman’s coefficient with technology

Section titled “Finding Spearman’s coefficient with technology”

In IB exams you find rsr_s with your GDC:

  1. Rank each variable (averaging tied ranks).
  2. Enter the two lists of ranks.
  3. Use your GDC’s two-variable statistics (or linear regression) to find the correlation coefficient rr of the ranks. That number is rsr_s.

Some GDCs and spreadsheets can rank the data or find rsr_s for you, but it’s worth knowing how to rank by hand, especially with ties. You don’t need to know how either coefficient is derived.

Pearson’s rrSpearman’s rsr_s
Measureshow close the points are to a straight linehow consistently one variable rises (or falls) as the other rises
Use it whenyou want to test for a linear relationshipthe relationship might be any monotonic curve, or the data are ranks
Data needednumerical valuesvalues or ranks
Effect of outlierscan change a lotless sensitive

rs=1r_s = 1 doesn’t mean the points lie on a line. It means they’re in perfect increasing order. The figure shows data that curve steeply upward: rr is only 0.8730.873, but the ranks match exactly, so rs=1r_s = 1.

Left: followers against week, curving upward, r = 0.873. Right: the ranks of the same data lie exactly on a straight line, so Spearman's rs = 1. 1 2 3 4 5 6 7 8 0 25 50 75 100 week followers (thousands) The data r = 0.873 1 1 2 2 3 3 4 4 5 5 6 6 7 7 8 8 rank of week rank of followers The ranks rs = 1
The data curve upward (left), but their ranks lie on a straight line (right), so rs=1r_s = 1 while r=0.873r = 0.873.

Seven points have a strong linear pattern. Then an eighth point, (8,60)(8, 60), is added:

xx1122334455667788
yy12121515141418182121202024246060
rrrsr_s
First seven points0.9570.9570.9290.929
All eight points0.7580.7580.9520.952

The outlier pulls rr down a lot, because it’s far from the line through the other points. But 6060 is simply the largest yy-value, so it gets rank 88, the same as if it had been 2525. Ranks don’t care how far a value is from the others, only where it sits in the order. That’s why rsr_s is less sensitive to outliers. (Less sensitive isn’t immune: an outlier that’s out of order, like (8,2)(8, 2), changes rsr_s too.)

As always, a strong correlation of either kind doesn’t prove that one variable causes the other (see correlation and causation).

Seven students scored 84,71,90,71,65,84,7784, 71, 90, 71, 65, 84, 77 on a quiz. Rank the scores, giving rank 11 to the highest score.

Solution. Put the scores in order from highest to lowest and number them:

In order9090848484847777717171716565
Position11223344556677
Rank112.52.52.52.5445.55.55.55.577

The two 8484s share positions 22 and 33, so each gets 2.52.5. The two 7171s share positions 55 and 66, so each gets 5.55.5. In the original order:

Score8484717190907171656584847777
Rank2.52.55.55.5115.55.5772.52.544

Check: the ranks of 77 items always add up to 1+2+⋯+7=281 + 2 + \cdots + 7 = 28, and 2.5+5.5+1+5.5+7+2.5+4=282.5 + 5.5 + 1 + 5.5 + 7 + 2.5 + 4 = 28. ✓

Two judges ranked eight dancers, A to H, in a competition.

DancerABCDEFGH
Judge 11122334455667788
Judge 23311225544886677

Find rsr_s and interpret it.

Solution. The data are already ranks, so enter the two rows as lists and find their correlation coefficient on your GDC:

rs=0.833 (to 3 s.f.)r_s = 0.833 \text{ (to 3 s.f.)}

There is a strong positive agreement between the judges: dancers ranked highly by Judge 1 also tend to be ranked highly by Judge 2. Pearson’s rr on the original scores isn’t possible here, since only the rankings are given. That’s a situation where Spearman’s coefficient is the natural choice.

The table shows how many hours eight runners trained per week and their times (in minutes) for a 55 km race.

Runner12345678
Hours, hh225533886655101044
Time, tt31.531.526.026.029.829.824.124.125.225.227.327.323.923.927.327.3
  • (a) Rank both variables, giving rank 11 to the largest value.
  • (b) Find rsr_s and interpret it in context.

Solution.

(a) For hours, 1010 is rank 11, then 88, 66, and the two 55s share ranks 44 and 55, so each gets 4.54.5. For times, 31.531.5 is rank 11, then 29.829.8, and the two 27.327.3s share ranks 33 and 44, so each gets 3.53.5.

Runner12345678
Rank of hh884.54.57722334.54.51166
Rank of tt11552277663.53.5883.53.5

(b) Enter the two rows of ranks and find their correlation coefficient:

rs=−0.982 (to 3 s.f.)r_s = -0.982 \text{ (to 3 s.f.)}

There is a very strong negative monotonic relationship: runners who train more hours tend to have faster (lower) race times.

A new social media account had these follower counts (in thousands) at the end of each of its first eight weeks:

Week, ww1122334455667788
Followers, FF22335599161630305555100100
  • (a) Find Pearson’s rr and Spearman’s rsr_s.
  • (b) Explain why the two values are different, and say which better describes the relationship.

Solution.

(a) On your GDC, the correlation coefficient of ww and FF is r=0.873r = 0.873 (to 3 s.f.).

The follower counts increase every week, so their ranks are 1,2,…,81, 2, \ldots, 8, exactly the same as the ranks of the weeks. So rs=1r_s = 1.

(b) The number of followers always increases, so the relationship is perfectly monotonic, giving rs=1r_s = 1. But it isn’t linear: the counts grow faster and faster (the scatter plot curves upward, as in the figure above), so the points don’t lie close to a straight line and rr is smaller. Spearman’s rsr_s describes this relationship better. Pearson’s rr is only the right tool if you’re asking whether the relationship is linear, and here it clearly isn’t.

Giving tied values different ranks. If two values are equal, they must get the same rank: the average of the positions they share. Ranking two 77s as 33 and 44 (in whichever order) gives a value of rsr_s that depends on an arbitrary choice.

Ranking the two variables in opposite directions. If you give rank 11 to the largest xx but the smallest yy, the sign of rsr_s flips. Rank both variables the same way.

Typing the raw data into the GDC instead of the ranks. The correlation coefficient of the raw values is Pearson’s rr. You only get rsr_s when you enter the ranks.

Forgetting to skip ranks after a tie. After two values share ranks 22 and 33 (both 2.52.5), the next value gets rank 44, not 33. Check that your ranks add up to 1+2+⋯+n1 + 2 + \cdots + n.

Reading rs=1r_s = 1 as “the points lie on a line”. rs=1r_s = 1 means the points are in perfect increasing order. They could lie on any increasing curve. Only r=1r = 1 means a perfect straight line.

Saying Spearman’s coefficient ignores outliers. It’s less sensitive than rr, not immune. An outlier that changes the order of the data still changes rsr_s.

1. (Warm-up) Rank these values, giving rank 11 to the largest: 12,15,9,15,20,11,1512, 15, 9, 15, 20, 11, 15.

Solution

In order: 20,15,15,15,12,11,920, 15, 15, 15, 12, 11, 9. The three 1515s share positions 22, 33 and 44, so each gets 2+3+43=3\dfrac{2 + 3 + 4}{3} = 3.

Value12121515991515202011111515
Rank55337733116633

Check: 5+3+7+3+1+6+3=28=1+2+⋯+75 + 3 + 7 + 3 + 1 + 6 + 3 = 28 = 1 + 2 + \cdots + 7. ✓

2. (Warm-up) Describe what each value tells you about the relationship between two variables.

  • (a) rs=0.95r_s = 0.95
  • (b) rs=−1r_s = -1
  • (c) rs=0.08r_s = 0.08
Solution

(a) A very strong positive monotonic relationship: as one variable increases, the other almost always increases too.

(b) A perfect negative monotonic relationship: as one variable increases, the other always decreases (the ranks are exactly reversed).

(c) Very weak or no monotonic relationship: there’s no consistent increasing or decreasing pattern.

3. (Core) Two food critics ranked six restaurants.

RestaurantPQRSTU
Critic 1112233445566
Critic 2221133664455

Find rsr_s and comment on how well the critics agree.

Solution

Enter the two rows of ranks into your GDC and find their correlation coefficient:

rs=0.771 (to 3 s.f.)r_s = 0.771 \text{ (to 3 s.f.)}

There is fairly strong positive agreement: restaurants that one critic ranks highly tend to be ranked highly by the other, though not exactly (restaurant S is the biggest disagreement).

4. (Core) A café recorded the average daily temperature and the number of hot chocolates sold on seven winter days.

Temperature (°C)22−5-588221212−1-155
Hot chocolates14014019019095951501506060150150120120
  • (a) Rank both variables, giving rank 11 to the largest value.
  • (b) Find rsr_s and interpret it in context.
Solution

(a) Temperatures in order: 12,8,5,2,2,−1,−512, 8, 5, 2, 2, -1, -5. The two 22s share ranks 44 and 55, so each gets 4.54.5. Sales in order: 190,150,150,140,120,95,60190, 150, 150, 140, 120, 95, 60. The two 150150s share ranks 22 and 33, so each gets 2.52.5.

Rank of temperature4.54.577224.54.5116633
Rank of sales4411662.52.5772.52.555

(b) Using the ranks on a GDC:

rs=−0.973 (to 3 s.f.)r_s = -0.973 \text{ (to 3 s.f.)}

There is a very strong negative monotonic relationship: the colder the day, the more hot chocolates the café tends to sell.

5. (Core) For a set of data, r=0.81r = 0.81 and rs=1r_s = 1.

  • (a) What does rs=1r_s = 1 tell you about the data?
  • (b) Why might rr be smaller than rsr_s here?
  • (c) Would a straight-line model be a good choice? Explain.
Solution

(a) The data are perfectly monotonic increasing: whenever xx increases, yy increases too, so the ranks of xx and yy match exactly.

(b) The points follow an increasing curve rather than a straight line, so they aren’t as close to a line as they could be. rr only measures closeness to a straight line.

(c) Probably not. Since rs=1r_s = 1 but rr is noticeably less than 11, the relationship is increasing but curved, so a curved (non-linear) model would likely fit better. Looking at the scatter plot would confirm this.

6. (Core) The table shows six points.

xx22445577991010
yy66991010131315151818
  • (a) Find rr and rsr_s.
  • (b) The point (3,30)(3, 30) is added. Find rr and rsr_s for all seven points.
  • (c) Comment on the effect of the new point on each coefficient.
Solution

(a) On a GDC, r=0.993r = 0.993 (to 3 s.f.). The yy-values increase whenever xx increases, so the ranks match and rs=1r_s = 1.

(b) With (3,30)(3, 30) included, r=0.117r = 0.117 and rs=0.464r_s = 0.464 (to 3 s.f.).

(c) The outlier nearly destroys Pearson’s rr, which drops from 0.9930.993 to 0.1170.117. Spearman’s rsr_s also drops, from 11 to 0.4640.464, because the new point is out of order (a small xx with the largest yy). But it drops by less: rsr_s is less sensitive to outliers than rr, though not immune to them.

7. (Core) A website shows each hotel’s star rating (11 to 55) and its average nightly price.

HotelABCDEFGH
Stars3355224433114455
Price ($)1281282652651221221401401351359595130130252252
  • (a) Explain why Spearman’s coefficient is more appropriate than Pearson’s here.
  • (b) Find rsr_s, giving rank 11 to the largest value of each variable.
Solution

(a) Star ratings are ordinal: a 44-star hotel is better than a 22-star hotel, but not necessarily “twice as good”, and the gaps between ratings aren’t equal amounts of anything. Spearman’s coefficient only uses the order, so it suits this kind of data. (There are also many tied values.)

(b) Stars in order: 5,5,4,4,3,3,2,15, 5, 4, 4, 3, 3, 2, 1, so the 55s get 1.51.5, the 44s get 3.53.5, the 33s get 5.55.5, then 77 and 88. Prices in order: 265,252,140,135,130,128,122,95265, 252, 140, 135, 130, 128, 122, 95.

HotelABCDEFGH
Rank of stars5.55.51.51.5773.53.55.55.5883.53.51.51.5
Rank of price6611773344885522

Using the ranks on a GDC:

rs=0.933 (to 3 s.f.)r_s = 0.933 \text{ (to 3 s.f.)}

There is a very strong positive monotonic relationship: hotels with more stars tend to cost more.

8. (Challenge) The table shows a population of bacteria (in thousands) after xx hours.

xx112233445566
yy337720205555150150400400
  • (a) Find rr and rsr_s for xx and yy.
  • (b) Find rr and rsr_s for xx and ln⁡y\ln y.
  • (c) Explain why rsr_s didn’t change but rr did.
Solution

(a) r=0.849r = 0.849 (to 3 s.f.). The yy-values always increase, so rs=1r_s = 1.

(b) For xx and ln⁡y\ln y: r=0.9997r = 0.9997 (to 4 s.f.), which is 1.001.00 to 3 s.f., and rs=1r_s = 1.

(c) Taking ln⁡\ln keeps the values in the same order (a bigger yy always gives a bigger ln⁡y\ln y), so the ranks don’t change and neither does rsr_s. But it changes the shape of the scatter plot: yy grows roughly exponentially, so ln⁡y\ln y against xx is almost a straight line, and rr gets much closer to 11.

9. (Challenge) Five data points are (4,10)(4, 10), (7,14)(7, 14), (7,11)(7, 11), (9,13)(9, 13) and (12,20)(12, 20).

  • (a) Find rsr_s correctly, using average ranks for the tie.
  • (b) Amir ranks the two 77s as 33 and 44 instead (largest first), giving rank 33 to the point (7,14)(7, 14). Bea does the same but gives rank 33 to (7,11)(7, 11). Find the value of rsr_s that each of them gets.
  • (c) Explain why averaging tied ranks is the fair method.
Solution

(a) Ranking from the largest: xx ranks are 5,3.5,3.5,2,15, 3.5, 3.5, 2, 1 and yy ranks are 5,2,4,3,15, 2, 4, 3, 1. On a GDC:

rs=0.821 (to 3 s.f.)r_s = 0.821 \text{ (to 3 s.f.)}

(b) Amir’s xx ranks are 5,3,4,2,15, 3, 4, 2, 1, which give rs=0.9r_s = 0.9. Bea’s are 5,4,3,2,15, 4, 3, 2, 1, which give rs=0.7r_s = 0.7.

(c) The two 77s are equal, so there’s no reason to put one ahead of the other. Breaking the tie by choice gives different answers (0.90.9 or 0.70.7) from the same data. Averaging treats the tied values equally and gives one answer, 0.8210.821, which happens to lie between the two.