Non-linear Regression
Not every relationship is a straight line. A thrown ball follows a parabola, a viral video’s views grow exponentially, and hours of daylight rise and fall with the seasons. Non-linear regression lets your GDC find the curve of a chosen type that fits the data best, just as linear regression finds the best line. On this page you’ll learn how the “best” curve is defined, how to measure how well it fits using and , and why the best number isn’t always the best model.
Key ideas
Section titled “Key ideas”Least squares regression curves
Section titled “Least squares regression curves”For each data point , the residual is the vertical distance from the point to the model:
The sum of square residuals is
A least squares regression curve is the curve of the chosen type with the smallest possible . You choose the type; technology finds the parameters. The types examined in IB AI are:
| Regression | Model | GDC hint |
|---|---|---|
| linear | linear regression | |
| quadratic | quadratic regression | |
| cubic | cubic regression | |
| exponential | (or ) | exponential regression |
| power | power regression | |
| sine | sine regression (radian mode) |
On most GDCs: enter the data in two lists, open the statistics calculation menu, choose the regression type, and read off the parameters (store the equation as a function so you can graph it and use it).
Two practical notes:
- Many GDCs find exponential and power regressions by fitting a straight line to the logged data, as on the linearizing data with logarithms page. Different software can therefore give slightly different parameters. The data on this page is chosen so that the methods agree to 3 s.f.
- Sine regression needs radian mode, and the answer can come out in different but equivalent forms (for example with a different ). On this page, and .
SS_res as a measure of fit
Section titled “SS_res as a measure of fit”For the same data, a smaller means the model’s predictions are closer to the observed values. only if the curve passes through every point. Because is in squared units of and grows with the number of points, compare values only between models of the same data.
The coefficient of determination
Section titled “The coefficient of determination”The coefficient of determination is a number between and that your GDC reports with the regression (you may need to turn on “diagnostics” or “stat diagnostics”). Interpretation:
is the proportion of the variability in that is accounted for by the model.
For example, means of the variation in is explained by the model, and is not.
- For a linear model, , the square of Pearson’s correlation coefficient.
- It can help to know that , where measures the total variation in . So exactly when . (This formula helps understanding but isn’t examined.)
Choosing between models
Section titled “Choosing between models”is useful, but by itself it is not a good way to decide between models. Also think about:
- Context. Does the shape make sense? A drug concentration should decrease toward , not rise again; a population can’t be negative.
- Behaviour beyond the data. Polynomials eventually head to . If you need to predict, the model must behave sensibly there.
- Simplicity. A cubic always has an at least as high as a quadratic on the same data (it has an extra parameter), but that doesn’t make it a better model.
- Residuals. A good model’s residuals are small and show no pattern.
Worked examples
Section titled “Worked examples”Example 1: Quadratic regression for a thrown ball
Section titled “Example 1: Quadratic regression for a thrown ball”The height metres of a ball is recorded seconds after it is thrown.
| (s) | ||||||
|---|---|---|---|---|---|---|
| (m) |
- (a) Find the quadratic regression model.
- (b) Find and , and interpret .
- (c) Use the model to find the maximum height and when the ball lands.
Solution.
(a) Quadratic regression on the GDC gives
(b) The residuals (observed minus predicted, using full values) are
| residual |
(using the unrounded residuals; the GDC can also compute this sum from a list of residuals)
The GDC gives (to 3 s.f.; more precisely ). So of the variability in the height is accounted for by the quadratic model: an excellent fit.
(c) Graph the model on the GDC. The maximum is at s, with m. The ball lands when : the positive zero is s. (The other zero, , is outside the domain .)
The leading coefficient is close to , half the acceleration due to gravity, which supports the model.
Example 2: Exponential regression for video views
Section titled “Example 2: Exponential regression for video views”A new video’s daily views are recorded.
| Day | ||||||
|---|---|---|---|---|---|---|
| Views |
- (a) Find an exponential regression model and interpret .
- (b) Predict the views on day .
- (c) On which day will the daily views first exceed , if the trend continues?
Solution.
(a) Exponential regression gives and , with (to 3 s.f.):
means the daily views are multiplied by about each day: an increase of about per day.
(b) Using the stored model with full values, views (3 s.f.).
(c) Solve with the GDC (using full values): . Since counts whole days, the views first exceed on day 9. This is an extrapolation: real viral growth always slows down eventually, so this prediction should be treated with caution.
Example 3: Two models with almost the same R²
Section titled “Example 3: Two models with almost the same R²”A patient’s blood concentration (mg/L) of a drug is measured every hour after an injection.
| (h) | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| (mg/L) |
- (a) Find a quadratic regression model and an exponential regression model, with for each.
- (b) Which model should be used to estimate the concentration at ? Justify your answer.
Solution.
(a) Quadratic regression: , with and .
Exponential regression: , with and (to 3 s.f.; a GDC’s exponential regression shows about ).
(b) Both models fit the data extremely well: their values differ by only about . The deciding factors are the context and the behaviour beyond the data:
- The quadratic has its minimum at h and then increases, predicting the concentration will rise again with no new dose. That’s not reasonable.
- The exponential keeps decreasing toward , as the body steadily clears the drug. A constant ratio of means is removed each hour, which is a standard model for drug clearance.
So use the exponential: with the full values from a GDC, mg/L. (Software that minimizes directly gives , so about mg/L either way.) Even then, is an extrapolation, so it’s an estimate.
Example 4: Sine regression for daylight hours
Section titled “Example 4: Sine regression for daylight hours”The average number of daylight hours in a Canadian city in month (with for January) is shown.
| (h) |
- (a) Find a sine regression model .
- (b) Find the period of the model and comment.
- (c) Estimate the daylight hours in mid-April ().
Solution.
(a) With the GDC in radian mode, sine regression gives
with and .
(b) The period is months. That’s very close to months, the length of a year, so the model makes sense. The amplitude h is close to half the difference between the longest and shortest days, and h is the average daylight over the year.
(c) hours (using full values).
Common mistakes
Section titled “Common mistakes”Choosing the type from alone. A higher doesn’t guarantee a better model. A cubic always fits at least as well as a quadratic on the same data, and a model can fit the data well but behave absurdly just outside it. Use context, the shape of the data and the behaviour of the curve as well.
Comparing SS_res across different data sets. depends on the units and on how many points there are. Only compare for different models of the same data.
Misinterpreting R². means of the variability in is accounted for by the model. It does not mean of the points lie on the curve, or that the predictions are accurate.
Using degree mode for sine regression. GDC sine regression works in radians. In degree mode you’ll get nonsense (or an error).
Rounding the parameters before predicting. Store the regression equation in your GDC and use the full values. Rounding to and then raising it to a large power changes the answer.
Trusting extrapolation. Regression describes the data you have. Predictions far outside the range of (especially with polynomial or exponential models) can be wildly wrong.
Practice
Section titled “Practice”1. (Warm-up) The model is proposed for the data below. Find .
Solution
Predicted values: , , , . Residuals: , , , .
2. (Warm-up) A power regression of a bird’s wingspan on its mass gives . Interpret this value.
Solution
of the variability in wingspan is accounted for by the power model of wingspan against mass; the other is due to other factors (or random variation).
3. (Warm-up) Which type of regression would you try first for each situation?
- (a) Daily maximum temperature over two years.
- (b) The number of bacteria in a dish over the first few hours.
- (c) The height of a basketball during a shot.
- (d) The time for a ball to fall from different heights, where appears to be proportional to a power of .
Solution
(a) Sine: the temperature repeats every year.
(b) Exponential: bacteria multiply at a constant rate at first.
(c) Quadratic: projectile motion.
(d) Power: .
4. (Core) A café tests different prices (dollars) for a lunch special and records its weekly profit (dollars).
| ($) | ||||||
|---|---|---|---|---|---|---|
| ($) |
- (a) Find the quadratic regression model and .
- (b) Use the model to find the price that gives the maximum profit, and the maximum profit.
Solution
(a) (full values , , ), with (and ).
(b) The vertex (using the GDC maximum, or with full values) is at , so a price of about $23.70 (3 s.f.), giving a maximum weekly profit of about $1870.
5. (Core) The number of users (thousands) of a new app months after launch is shown.
| (thousands) |
- (a) Find an exponential regression model , and the monthly percentage growth.
- (b) Use the model to find when the app will reach thousand users.
Solution
(a) (both parameters to 3 s.f.; ). Since , the number of users grows by about per month.
(b) Solve with full values: months after launch (during the tenth month).
6. (Core) The time (s) for a ball to fall from height (m) is measured.
| (m) | |||||
|---|---|---|---|---|---|
| (s) |
- (a) Find a power regression model .
- (b) Predict the time to fall m.
Solution
(a) (3 s.f.), with . So is proportional to .
(b) s (using full values).
7. (Core)
- (a) A linear model has Pearson’s correlation coefficient . Find and interpret it.
- (b) Two models are fitted to the same data. Model A has and model B has . Which fits the data better?
- (c) For a data set, and a model has . Use to find .
Solution
(a) For a linear model . About of the variability in is accounted for by the linear model.
(b) Model B: its smaller means its predictions are closer to the data. (Whether it’s the better model also depends on context.)
(c) .
8. (Challenge) A sunflower’s height (cm) is measured weekly.
| Week | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| (cm) |
- (a) Find quadratic and cubic regression models, with for each.
- (b) Use each model to predict the height at week .
- (c) Which model would you choose for predicting the height at week ? Explain.
Solution
(a) Quadratic: , .
Cubic: , .
(b) Using full values: quadratic cm; cubic cm.
(c) Neither. The cubic has the higher but predicts the sunflower shrinks from cm to about cm, which is impossible. The quadratic predicts growth speeding up forever, but the data shows growth already slowing down (the weekly increases fall from cm to cm). The shape suggests the height is levelling off, so a model with a horizontal asymptote, such as a logistic model, would be more suitable. This is a clear example of why alone is not enough.
9. (Challenge) The height of the tide metres at a harbour is recorded every hours, starting at midnight ().
| (h) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| (m) |
- (a) Find a sine regression model with and , and its .
- (b) Find the period of the tide according to the model.
- (c) Predict the height of the tide at 1 a.m. the next day.
Solution
(a) In radian mode: , with .
(b) Period hours (a typical period for ocean tides).
(c) 1 a.m. the next day is : m (using full values). This is a short extrapolation of a regular periodic pattern, so it’s reasonably reliable.