Linearizing Data with Logarithms
Curved data is hard to judge by eye: is it exponential, or a power law like ? Logarithms solve this problem. Taking logs turns both kinds of curve into straight lines, and straight lines are easy to recognize and easy to fit. From the gradient and intercept of the best-fit line you can read off the parameters of the original model. Logs also let you show data that spans many powers of on one sensible scale.
Key ideas
Section titled “Key ideas”Scaling with logarithms
Section titled “Scaling with logarithms”When values range from tiny to huge (masses of animals, populations of towns and countries, website visits), a normal axis squashes most of the data into a corner. Plotting instead of fixes this: each power of becomes one equal step, as on the logarithmic scales for pH and decibels. A logarithmic scale is also useful when what matters is the rate of growth rather than the actual size.
Exponential data: the semi-log graph
Section titled “Exponential data: the semi-log graph”Suppose . Take natural logs and use the laws of logarithms:
This has the form with and . So:
- if , the graph of against is a straight line;
- its gradient is , so ;
- its intercept is , so .
A graph with (or ) on the vertical axis and plain on the horizontal axis is called a semi-log graph.
Power data: the log-log graph
Section titled “Power data: the log-log graph”Suppose . Taking natural logs:
Now and . So:
- if , the graph of against is a straight line;
- its gradient is (the power itself);
- its intercept is , so .
A graph with logs on both axes is called a log-log graph.
| Model | Plot | Gradient | Intercept |
|---|---|---|---|
| (exponential) | against (semi-log) | ||
| (power) | against (log-log) |
You can use base- logs instead of ; then undo them with instead of . Just be consistent.
Deciding which model fits
Section titled “Deciding which model fits”Transform the data both ways and see which plot is closer to a straight line. A quick numerical check is Pearson’s correlation coefficient for each transformed data set: the one with closer to is more linear. Then use your GDC’s linear regression on the transformed data to find the best-fit line, and convert its gradient and intercept back into the parameters of the model.
Two practical points:
- A log-log plot needs and , and a semi-log plot needs , because and the logs of negative numbers are undefined.
- Keep full calculator values for the gradient and intercept, and round only the final parameters, because magnifies small rounding errors in .
Reading graphs with logarithmic axes
Section titled “Reading graphs with logarithmic axes”Published graphs often use axes whose labels go at equal spacing (a logarithmic axis) instead of showing directly. A straight line on such a graph means the same thing: straight on a semi-log graph means exponential, straight on a log-log graph means a power law. To find the model, read two points off the line in their original units and solve for the parameters. In IB examinations you’ll be asked to interpret these graphs, but not to draw or sketch them.
Worked examples
Section titled “Worked examples”Example 1: Choosing a scale
Section titled “Example 1: Choosing a scale”The typical masses of five animals are: mouse kg, cat kg, human kg, elephant kg, blue whale kg. Explain why a logarithmic scale is better for displaying these masses, and find the value of for each.
Solution. On an ordinary scale from to kg, the mouse, cat and human would all sit on top of each other at the far left; even the elephant would be squeezed near . Taking :
| Animal | mouse | cat | human | elephant | blue whale |
|---|---|---|---|---|---|
| (kg) | |||||
All five values now fit comfortably between and , and each step of means ” times heavier”. The whale has a mass about greater than the mouse’s, so it’s about times heavier: , or million times.
Example 2: Exponential or power? A semi-log fit
Section titled “Example 2: Exponential or power? A semi-log fit”A biologist counts the bacteria in a culture every hour.
| (h) | ||||||
|---|---|---|---|---|---|---|
- (a) Use a semi-log transformation to show that the data is very close to exponential.
- (b) Find a model of the form .
- (c) Interpret .
Solution.
(a) Find for each value:
The values of go up by almost exactly the same amount (about ) each hour, so the points lie on a straight line. On a GDC, linear regression of on gives (to 3 s.f.; more precisely ). For comparison, a log-log fit is not possible with , and without that point it gives , clearly less linear. So the data is exponential.
(b) The best-fit line is (gradient , intercept ). Convert back:
So (parameters to 3 s.f.).
(c) means the number of bacteria is multiplied by about every hour: it grows by about per hour.
Example 3: Kepler’s third law from a log-log fit
Section titled “Example 3: Kepler’s third law from a log-log fit”The table gives each planet’s average distance from the Sun, (in astronomical units, where Earth’s distance is ), and the time (in years) it takes to orbit the Sun.
| Planet | Mercury | Venus | Earth | Mars | Jupiter | Saturn |
|---|---|---|---|---|---|---|
| (AU) | ||||||
| (years) |
Find a model of the form .
Solution. For a power model, plot against :
| Planet | Mercury | Venus | Earth | Mars | Jupiter | Saturn |
|---|---|---|---|---|---|---|
Linear regression of on gives gradient and intercept , with (to 3 s.f.). The points lie on a straight line, so a power model fits.
The gradient is the power: . The intercept gives . So
This is Kepler’s third law: the square of the orbital period is proportional to the cube of the distance ( in these units).
Example 4: Interpreting given best-fit lines
Section titled “Example 4: Interpreting given best-fit lines”- (a) For one data set, the best-fit line of against is . Find the model for and interpret it.
- (b) For another data set, the best-fit line of against is . Find the model for .
Solution.
(a) A straight line on a semi-log graph means , with and :
So . Since , is decaying: it decreases by about (because ) for each increase of in .
(b) A straight line on a log-log graph means , with and :
This is an inverse-square law, like those on the direct and inverse variation page.
Common mistakes
Section titled “Common mistakes”Mixing up which graph goes with which model. Exponential is straight on a semi-log graph ( against ). Power is straight on a log-log graph ( against ). The variable in the exponent ( in ) stays unlogged.
Using the intercept directly as a parameter. The intercept of the line is (or ), not itself. You must undo the log: . For a semi-log graph, the gradient is , so , but for a log-log graph the gradient is already the power .
Mixing bases. If you took of the data, the parameters are and , not and . Use the same base throughout.
Rounding the intercept too early. but : a small change in makes a big change in . Keep full values from your GDC and round only at the end.
Trying to log zero or negative values. is undefined, so a point with can’t go on a log-log graph. Leave such points out (and say so), or use a semi-log graph if the model is exponential.
Reading a logarithmic axis as if it were ordinary. Halfway between and on a log axis is about (since ), not . Read values at the labelled powers of , or convert using logs.
Practice
Section titled “Practice”1. (Warm-up) The best-fit line of against is . Find and in the model .
Solution
and , so (3 s.f.).
2. (Warm-up) The best-fit line of against is . Find the model for .
Solution
This is a power model with and , so (3 s.f.).
3. (Warm-up) Three quantities are , and . Find of each to 3 s.f., and explain why these logs would be easier to show on one axis than the original numbers.
Solution
, , .
The original numbers differ by a factor of more than million, so on an ordinary axis the first two would be indistinguishable from . The logs all lie between and , so every value is clearly visible.
4. (Core) The value dollars of a car years after it was bought is shown.
| (years) | ||||||
|---|---|---|---|---|---|---|
| ($) |
- (a) Find the equation of the best-fit line of against , and the value of .
- (b) Hence find a model , and the annual percentage decrease.
- (c) Use the model to estimate when the car will be worth $10 000.
Solution
(a) The values of are , , , , , . Linear regression gives
(b) and . So . Since , the car loses about of its value each year.
(c) Solve using the line: , so
5. (Core) The period (s) of a pendulum is measured for different lengths (m).
| (m) | |||||
|---|---|---|---|---|---|
| (s) |
- (a) Use a log-log transformation to find a model .
- (b) Find the length of a pendulum with a period of s.
Solution
(a) Linear regression of on gives gradient and intercept , with (to 3 s.f.). So and :
(This matches the physics formula .)
(b) Solve (using full values): m.
6. (Core) A best-fit line of against has gradient and intercept . Find the model and describe what tells you.
Solution
Base , so undo with powers of : and . So .
Since , decreases by about each time increases by .
7. (Core) A graph has an ordinary horizontal axis for and a logarithmic vertical axis for (labelled , , , ). The data lies on a straight line through the points and , read in original units. Find the model and interpret it.
Solution
A straight line on a semi-log graph means . At , . Then
So (3 s.f.): grows by about for each increase of in (it doubles every units, since over units).
8. (Challenge) For the data below, decide whether an exponential or a power model fits better by comparing for a semi-log and a log-log transformation. Find the better model and use it to predict when .
Solution
Semi-log ( against ): . Log-log ( against ): (). The log-log plot is far more linear, so use a power model.
The log-log best-fit line is , so and :
At : (3 s.f.). This is a small extrapolation beyond , so treat it with some caution.
9. (Challenge) On a log-log graph (both axes logarithmic, in original units), a straight line passes through and . Find the model .
Solution
Both points satisfy , so divide:
Then . So (3 s.f.).
Check with the other point: . ✓