Skip to content
Family Table Math
Auto

Linearizing Data with Logarithms

Curved data is hard to judge by eye: is it exponential, or a power law like y=ax2y = ax^2? Logarithms solve this problem. Taking logs turns both kinds of curve into straight lines, and straight lines are easy to recognize and easy to fit. From the gradient and intercept of the best-fit line you can read off the parameters of the original model. Logs also let you show data that spans many powers of 1010 on one sensible scale.

When values range from tiny to huge (masses of animals, populations of towns and countries, website visits), a normal axis squashes most of the data into a corner. Plotting log⁡x\log x instead of xx fixes this: each power of 1010 becomes one equal step, as on the logarithmic scales for pH and decibels. A logarithmic scale is also useful when what matters is the rate of growth rather than the actual size.

Suppose y=kaxy = ka^x. Take natural logs and use the laws of logarithms:

ln⁡y=ln⁡(kax)=ln⁡k+xln⁡a\ln y = \ln\left(ka^x\right) = \ln k + x\ln a

This has the form Y=mX+cY = mX + c with Y=ln⁡yY = \ln y and X=xX = x. So:

  • if y=kaxy = ka^x, the graph of ln⁡y\ln y against xx is a straight line;
  • its gradient is m=ln⁡am = \ln a, so a=ema = e^m;
  • its intercept is c=ln⁡kc = \ln k, so k=eck = e^c.

A graph with ln⁡y\ln y (or log⁡y\log y) on the vertical axis and plain xx on the horizontal axis is called a semi-log graph.

Suppose y=axny = ax^n. Taking natural logs:

ln⁡y=ln⁡(axn)=ln⁡a+nln⁡x\ln y = \ln\left(ax^n\right) = \ln a + n\ln x

Now Y=ln⁡yY = \ln y and X=ln⁡xX = \ln x. So:

  • if y=axny = ax^n, the graph of ln⁡y\ln y against ln⁡x\ln x is a straight line;
  • its gradient is nn (the power itself);
  • its intercept is c=ln⁡ac = \ln a, so a=eca = e^c.

A graph with logs on both axes is called a log-log graph.

ModelPlotGradientIntercept
y=kaxy = ka^x (exponential)ln⁡y\ln y against xx (semi-log)ln⁡a\ln aln⁡k\ln k
y=axny = ax^n (power)ln⁡y\ln y against ln⁡x\ln x (log-log)nnln⁡a\ln a

You can use base-1010 logs instead of ln⁡\ln; then undo them with 10(…)10^{(\ldots)} instead of e(…)e^{(\ldots)}. Just be consistent.

Transform the data both ways and see which plot is closer to a straight line. A quick numerical check is Pearson’s correlation coefficient rr for each transformed data set: the one with ∣r∣|r| closer to 11 is more linear. Then use your GDC’s linear regression on the transformed data to find the best-fit line, and convert its gradient and intercept back into the parameters of the model.

Two practical points:

  • A log-log plot needs x>0x \gt 0 and y>0y \gt 0, and a semi-log plot needs y>0y \gt 0, because ln⁡0\ln 0 and the logs of negative numbers are undefined.
  • Keep full calculator values for the gradient and intercept, and round only the final parameters, because ece^c magnifies small rounding errors in cc.

Published graphs often use axes whose labels go 1,10,100,10001, 10, 100, 1000 at equal spacing (a logarithmic axis) instead of showing ln⁡y\ln y directly. A straight line on such a graph means the same thing: straight on a semi-log graph means exponential, straight on a log-log graph means a power law. To find the model, read two points off the line in their original units and solve for the parameters. In IB examinations you’ll be asked to interpret these graphs, but not to draw or sketch them.

The typical masses of five animals are: mouse 0.020.02 kg, cat 44 kg, human 7070 kg, elephant 50005000 kg, blue whale 150 000150\,000 kg. Explain why a logarithmic scale is better for displaying these masses, and find the value of log⁡10m\log_{10} m for each.

Solution. On an ordinary scale from 00 to 150 000150\,000 kg, the mouse, cat and human would all sit on top of each other at the far left; even the elephant would be squeezed near 00. Taking log⁡10\log_{10}:

Animalmousecathumanelephantblue whale
mm (kg)0.020.0244707050005000150 000150\,000
log⁡10m\log_{10} m−1.70-1.700.6020.6021.851.853.703.705.185.18

All five values now fit comfortably between −2-2 and 66, and each step of 11 means ”1010 times heavier”. The whale has a log⁡10\log_{10} mass about 6.96.9 greater than the mouse’s, so it’s about 106.8810^{6.88} times heavier: 150 0000.02=7 500 000\dfrac{150\,000}{0.02} = 7\,500\,000, or 7.57.5 million times.

Example 2: Exponential or power? A semi-log fit

Section titled “Example 2: Exponential or power? A semi-log fit”

A biologist counts the bacteria NN in a culture every hour.

tt (h)001122334455
NN52052079079012301230188018802900290044504450
  • (a) Use a semi-log transformation to show that the data is very close to exponential.
  • (b) Find a model of the form N=katN = ka^t.
  • (c) Interpret aa.

Solution.

(a) Find ln⁡N\ln N for each value:

tt001122334455
ln⁡N\ln N6.2546.2546.6726.6727.1157.1157.5397.5397.9727.9728.4018.401

The values of ln⁡N\ln N go up by almost exactly the same amount (about 0.430.43) each hour, so the points (t,ln⁡N)(t, \ln N) lie on a straight line. On a GDC, linear regression of ln⁡N\ln N on tt gives r=1.00r = 1.00 (to 3 s.f.; more precisely 0.999980.99998). For comparison, a log-log fit is not possible with t=0t = 0, and without that point it gives r≈0.974r \approx 0.974, clearly less linear. So the data is exponential.

Semi-log plot of ln N against t for six bacteria counts. The points lie almost exactly on the straight best-fit line ln N = 0.430t + 6.25. 0 1 2 3 4 5 6 6.5 7 7.5 8 8.5 ln N = 0.430t + 6.25 time t (h) ln N
On a semi-log graph the exponential data becomes a straight line.

(b) The best-fit line is ln⁡N=0.43028t+6.24977\ln N = 0.43028t + 6.24977 (gradient m≈0.430m \approx 0.430, intercept c≈6.25c \approx 6.25). Convert back:

a=em=e0.43028≈1.54,k=ec=e6.24977≈518a = e^{m} = e^{0.43028} \approx 1.54, \qquad k = e^{c} = e^{6.24977} \approx 518

So N=518(1.54)tN = 518(1.54)^t (parameters to 3 s.f.).

(c) a≈1.54a \approx 1.54 means the number of bacteria is multiplied by about 1.541.54 every hour: it grows by about 54%54\% per hour.

Example 3: Kepler’s third law from a log-log fit

Section titled “Example 3: Kepler’s third law from a log-log fit”

The table gives each planet’s average distance from the Sun, rr (in astronomical units, where Earth’s distance is 11), and the time TT (in years) it takes to orbit the Sun.

PlanetMercuryVenusEarthMarsJupiterSaturn
rr (AU)0.3870.3870.7230.723111.5241.5245.2035.2039.5379.537
TT (years)0.2410.2410.6150.615111.8811.88111.8611.8629.4629.46

Find a model of the form T=arnT = ar^n.

Solution. For a power model, plot ln⁡T\ln T against ln⁡r\ln r:

PlanetMercuryVenusEarthMarsJupiterSaturn
ln⁡r\ln r−0.949-0.949−0.324-0.324000.4210.4211.6491.6492.2552.255
ln⁡T\ln T−1.423-1.423−0.486-0.486000.6320.6322.4732.4733.3833.383

Linear regression of ln⁡T\ln T on ln⁡r\ln r gives gradient 1.49971.4997 and intercept 0.0002800.000280, with r=1.00r = 1.00 (to 3 s.f.). The points lie on a straight line, so a power model fits.

Log-log plot of ln T against ln r for six planets, Mercury to Saturn. The points lie on the straight best-fit line ln T = 1.50 ln r + 0.000280, which has gradient 1.5. −2 −1 1 2 −1 1 2 3 ln T = 1.50 ln r + 0.000280 Mercury Venus Earth Mars Jupiter Saturn ln r ln T
The planets lie on a straight line of gradient 1.51.5 on a log-log graph, so TT is a power of rr.

The gradient is the power: n≈1.50n \approx 1.50. The intercept gives a=e0.000280≈1.00a = e^{0.000280} \approx 1.00. So

T=1.00 r1.50(to 3 s.f.), that is, T≈r3/2T = 1.00\,r^{1.50} \qquad\text{(to 3 s.f.), that is, } T \approx r^{3/2}

This is Kepler’s third law: the square of the orbital period is proportional to the cube of the distance (T2=r3T^2 = r^3 in these units).

Example 4: Interpreting given best-fit lines

Section titled “Example 4: Interpreting given best-fit lines”
  • (a) For one data set, the best-fit line of ln⁡y\ln y against xx is ln⁡y=1.6−0.25x\ln y = 1.6 - 0.25x. Find the model for yy and interpret it.
  • (b) For another data set, the best-fit line of ln⁡y\ln y against ln⁡x\ln x is ln⁡y=3.5−2ln⁡x\ln y = 3.5 - 2\ln x. Find the model for yy.

Solution.

(a) A straight line on a semi-log graph means y=kaxy = ka^x, with ln⁡k=1.6\ln k = 1.6 and ln⁡a=−0.25\ln a = -0.25:

k=e1.6≈4.95,a=e−0.25≈0.779k = e^{1.6} \approx 4.95, \qquad a = e^{-0.25} \approx 0.779

So y=4.95(0.779)xy = 4.95(0.779)^x. Since a<1a \lt 1, yy is decaying: it decreases by about 22.1%22.1\% (because 1−0.7788=0.22121 - 0.7788 = 0.2212) for each increase of 11 in xx.

(b) A straight line on a log-log graph means y=axny = ax^n, with n=−2n = -2 and ln⁡a=3.5\ln a = 3.5:

a=e3.5≈33.1⇒y=33.1x−2=33.1x2a = e^{3.5} \approx 33.1 \quad\Rightarrow\quad y = 33.1x^{-2} = \frac{33.1}{x^2}

This is an inverse-square law, like those on the direct and inverse variation page.

Mixing up which graph goes with which model. Exponential y=kaxy = ka^x is straight on a semi-log graph (ln⁡y\ln y against xx). Power y=axny = ax^n is straight on a log-log graph (ln⁡y\ln y against ln⁡x\ln x). The variable in the exponent (xx in axa^x) stays unlogged.

Using the intercept directly as a parameter. The intercept of the line is ln⁡k\ln k (or ln⁡a\ln a), not kk itself. You must undo the log: k=eck = e^c. For a semi-log graph, the gradient is ln⁡a\ln a, so a=ema = e^m, but for a log-log graph the gradient is already the power nn.

Mixing bases. If you took log⁡10\log_{10} of the data, the parameters are 10c10^c and 10m10^m, not ece^c and eme^m. Use the same base throughout.

Rounding the intercept too early. e6.25≈518.0e^{6.25} \approx 518.0 but e6.2≈492.7e^{6.2} \approx 492.7: a small change in cc makes a big change in kk. Keep full values from your GDC and round only at the end.

Trying to log zero or negative values. ln⁡0\ln 0 is undefined, so a point with x=0x = 0 can’t go on a log-log graph. Leave such points out (and say so), or use a semi-log graph if the model is exponential.

Reading a logarithmic axis as if it were ordinary. Halfway between 1010 and 100100 on a log axis is about 31.631.6 (since 101.5≈31.610^{1.5} \approx 31.6), not 5555. Read values at the labelled powers of 1010, or convert using logs.

1. (Warm-up) The best-fit line of ln⁡y\ln y against xx is ln⁡y=0.4x+1.2\ln y = 0.4x + 1.2. Find kk and aa in the model y=kaxy = ka^x.

Solution

k=e1.2≈3.32k = e^{1.2} \approx 3.32 and a=e0.4≈1.49a = e^{0.4} \approx 1.49, so y=3.32(1.49)xy = 3.32(1.49)^x (3 s.f.).

2. (Warm-up) The best-fit line of ln⁡y\ln y against ln⁡x\ln x is ln⁡y=2.5ln⁡x−0.7\ln y = 2.5\ln x - 0.7. Find the model for yy.

Solution

This is a power model with n=2.5n = 2.5 and a=e−0.7≈0.497a = e^{-0.7} \approx 0.497, so y=0.497x2.5y = 0.497x^{2.5} (3 s.f.).

3. (Warm-up) Three quantities are 0.000 040.000\,04, 33 and 25 00025\,000. Find log⁡10\log_{10} of each to 3 s.f., and explain why these logs would be easier to show on one axis than the original numbers.

Solution

log⁡100.000 04≈−4.40\log_{10} 0.000\,04 \approx -4.40, log⁡103≈0.477\log_{10} 3 \approx 0.477, log⁡1025 000≈4.40\log_{10} 25\,000 \approx 4.40.

The original numbers differ by a factor of more than 600600 million, so on an ordinary axis the first two would be indistinguishable from 00. The logs all lie between −5-5 and 55, so every value is clearly visible.

4. (Core) The value VV dollars of a car tt years after it was bought is shown.

tt (years)001122334455
VV ($)32 00032\,00027 10027\,10023 20023\,20019 70019\,70016 80016\,80014 30014\,300
  • (a) Find the equation of the best-fit line of ln⁡V\ln V against tt, and the value of rr.
  • (b) Hence find a model V=katV = ka^t, and the annual percentage decrease.
  • (c) Use the model to estimate when the car will be worth $10 000.
Solution

(a) The values of ln⁡V\ln V are 10.37310.373, 10.20710.207, 10.05210.052, 9.8889.888, 9.7299.729, 9.5689.568. Linear regression gives

ln⁡V=−0.16073t+10.37151,r=−1.00 (to 3 s.f.; −0.99998)\ln V = -0.16073t + 10.37151, \qquad r = -1.00 \ (\text{to 3 s.f.; } -0.99998)

(b) a=e−0.16073≈0.852a = e^{-0.16073} \approx 0.852 and k=e10.37151≈31 900k = e^{10.37151} \approx 31\,900. So V=31 900(0.852)tV = 31\,900(0.852)^t. Since 1−0.8515=0.14851 - 0.8515 = 0.1485, the car loses about 14.8%14.8\% of its value each year.

(c) Solve ln⁡V=ln⁡10 000\ln V = \ln 10\,000 using the line: −0.16073t+10.37151=ln⁡10 000≈9.21034-0.16073t + 10.37151 = \ln 10\,000 \approx 9.21034, so

t=10.37151−9.210340.16073≈7.22 yearst = \frac{10.37151 - 9.21034}{0.16073} \approx 7.22 \text{ years}

5. (Core) The period TT (s) of a pendulum is measured for different lengths LL (m).

LL (m)0.10.10.250.250.50.51.01.02.02.0
TT (s)0.630.631.001.001.421.422.012.012.842.84
  • (a) Use a log-log transformation to find a model T=aLnT = aL^n.
  • (b) Find the length of a pendulum with a period of 33 s.
Solution

(a) Linear regression of ln⁡T\ln T on ln⁡L\ln L gives gradient 0.502890.50289 and intercept 0.697130.69713, with r=1.00r = 1.00 (to 3 s.f.). So n≈0.503n \approx 0.503 and a=e0.69713≈2.01a = e^{0.69713} \approx 2.01:

T=2.01L0.503T = 2.01L^{0.503}

(This matches the physics formula T=2πL/g≈2.01L0.5T = 2\pi\sqrt{L/g} \approx 2.01L^{0.5}.)

(b) Solve 2.00799L0.50289=32.00799L^{0.50289} = 3 (using full values): L=(32.00799)1/0.50289≈2.22L = \left(\dfrac{3}{2.00799}\right)^{1/0.50289} \approx 2.22 m.

6. (Core) A best-fit line of log⁡10y\log_{10} y against xx has gradient −0.15-0.15 and intercept 2.82.8. Find the model y=kaxy = ka^x and describe what aa tells you.

Solution

Base 1010, so undo with powers of 1010: k=102.8≈631k = 10^{2.8} \approx 631 and a=10−0.15≈0.708a = 10^{-0.15} \approx 0.708. So y=631(0.708)xy = 631(0.708)^x.

Since 1−0.7079=0.29211 - 0.7079 = 0.2921, yy decreases by about 29.2%29.2\% each time xx increases by 11.

7. (Core) A graph has an ordinary horizontal axis for xx and a logarithmic vertical axis for yy (labelled 1010, 100100, 10001000, 10 00010\,000). The data lies on a straight line through the points (0,200)(0, 200) and (10,3200)(10, 3200), read in original units. Find the model and interpret it.

Solution

A straight line on a semi-log graph means y=kaxy = ka^x. At x=0x = 0, y=k=200y = k = 200. Then

200a10=3200⇒a10=16⇒a=161/10≈1.32200a^{10} = 3200 \quad\Rightarrow\quad a^{10} = 16 \quad\Rightarrow\quad a = 16^{1/10} \approx 1.32

So y=200(1.32)xy = 200(1.32)^x (3 s.f.): yy grows by about 32%32\% for each increase of 11 in xx (it doubles every 2.52.5 units, since 16=2416 = 2^4 over 1010 units).

8. (Challenge) For the data below, decide whether an exponential or a power model fits better by comparing rr for a semi-log and a log-log transformation. Find the better model and use it to predict yy when x=10x = 10.

xx112233446688
yy3.03.010.610.621.521.536.636.674.874.8127127
Solution

Semi-log (ln⁡y\ln y against xx): r≈0.955r \approx 0.955. Log-log (ln⁡y\ln y against ln⁡x\ln x): r≈1.00r \approx 1.00 (0.999980.99998). The log-log plot is far more linear, so use a power model.

The log-log best-fit line is ln⁡y=1.79710ln⁡x+1.10306\ln y = 1.79710\ln x + 1.10306, so n≈1.80n \approx 1.80 and a=e1.10306≈3.01a = e^{1.10306} \approx 3.01:

y=3.01x1.80y = 3.01x^{1.80}

At x=10x = 10: y=3.01339×101.79710≈189y = 3.01339 \times 10^{1.79710} \approx 189 (3 s.f.). This is a small extrapolation beyond x=8x = 8, so treat it with some caution.

9. (Challenge) On a log-log graph (both axes logarithmic, in original units), a straight line passes through (2,50)(2, 50) and (20,8)(20, 8). Find the model y=axny = ax^n.

Solution

Both points satisfy y=axny = ax^n, so divide:

850=a⋅20na⋅2n=10n⇒n=log⁡100.16≈−0.796\frac{8}{50} = \frac{a \cdot 20^n}{a \cdot 2^n} = 10^n \quad\Rightarrow\quad n = \log_{10} 0.16 \approx -0.796

Then a=502n=50×20.79588≈86.8a = \dfrac{50}{2^n} = 50 \times 2^{0.79588} \approx 86.8. So y=86.8x−0.796y = 86.8x^{-0.796} (3 s.f.).

Check with the other point: 86.8068×20−0.79588≈8.0086.8068 \times 20^{-0.79588} \approx 8.00. ✓