Skip to content
Family Table Math
Auto

Non-linear Regression

Not every relationship is a straight line. A thrown ball follows a parabola, a viral video’s views grow exponentially, and hours of daylight rise and fall with the seasons. Non-linear regression lets your GDC find the curve of a chosen type that fits the data best, just as linear regression finds the best line. On this page you’ll learn how the “best” curve is defined, how to measure how well it fits using SSresSS_{res} and R2R^2, and why the best number isn’t always the best model.

For each data point (xi,yi)(x_i, y_i), the residual is the vertical distance from the point to the model:

residual=yi−f(xi)(observed−predicted)\text{residual} = y_i - f(x_i) \qquad (\text{observed} - \text{predicted})

The sum of square residuals is

SSres=∑i=1n(yi−f(xi))2SS_{res} = \sum_{i=1}^{n} \big(y_i - f(x_i)\big)^2

A least squares regression curve is the curve of the chosen type with the smallest possible SSresSS_{res}. You choose the type; technology finds the parameters. The types examined in IB AI are:

RegressionModelGDC hint
lineary=ax+by = ax + blinear regression
quadraticy=ax2+bx+cy = ax^2 + bx + cquadratic regression
cubicy=ax3+bx2+cx+dy = ax^3 + bx^2 + cx + dcubic regression
exponentialy=abxy = ab^x (or y=aekxy = ae^{kx})exponential regression
powery=axby = ax^bpower regression
siney=asin⁡(bx+c)+dy = a\sin(bx + c) + dsine regression (radian mode)

On most GDCs: enter the data in two lists, open the statistics calculation menu, choose the regression type, and read off the parameters (store the equation as a function so you can graph it and use it).

Two practical notes:

  • Many GDCs find exponential and power regressions by fitting a straight line to the logged data, as on the linearizing data with logarithms page. Different software can therefore give slightly different parameters. The data on this page is chosen so that the methods agree to 3 s.f.
  • Sine regression needs radian mode, and the answer can come out in different but equivalent forms (for example with a different cc). On this page, a>0a \gt 0 and b>0b \gt 0.

For the same data, a smaller SSresSS_{res} means the model’s predictions are closer to the observed values. SSres=0SS_{res} = 0 only if the curve passes through every point. Because SSresSS_{res} is in squared units of yy and grows with the number of points, compare SSresSS_{res} values only between models of the same data.

The coefficient of determination R2R^2 is a number between 00 and 11 that your GDC reports with the regression (you may need to turn on “diagnostics” or “stat diagnostics”). Interpretation:

R2R^2 is the proportion of the variability in yy that is accounted for by the model.

For example, R2=0.94R^2 = 0.94 means 94%94\% of the variation in yy is explained by the model, and 6%6\% is not.

  • For a linear model, R2=r2R^2 = r^2, the square of Pearson’s correlation coefficient.
  • It can help to know that R2=1−SSresSStotR^2 = 1 - \dfrac{SS_{res}}{SS_{tot}}, where SStot=∑(yi−yˉ)2SS_{tot} = \sum (y_i - \bar{y})^2 measures the total variation in yy. So R2=1R^2 = 1 exactly when SSres=0SS_{res} = 0. (This formula helps understanding but isn’t examined.)

R2R^2 is useful, but by itself it is not a good way to decide between models. Also think about:

  • Context. Does the shape make sense? A drug concentration should decrease toward 00, not rise again; a population can’t be negative.
  • Behaviour beyond the data. Polynomials eventually head to ±∞\pm\infty. If you need to predict, the model must behave sensibly there.
  • Simplicity. A cubic always has an R2R^2 at least as high as a quadratic on the same data (it has an extra parameter), but that doesn’t make it a better model.
  • Residuals. A good model’s residuals are small and show no pattern.

Example 1: Quadratic regression for a thrown ball

Section titled “Example 1: Quadratic regression for a thrown ball”

The height hh metres of a ball is recorded tt seconds after it is thrown.

tt (s)000.50.5111.51.5222.52.5
hh (m)1.61.67.77.711.511.513.113.111.811.88.48.4
  • (a) Find the quadratic regression model.
  • (b) Find SSresSS_{res} and R2R^2, and interpret R2R^2.
  • (c) Use the model to find the maximum height and when the ball lands.

Solution.

(a) Quadratic regression on the GDC gives

h=−4.85t2+14.9t+1.55(full values −4.85, 14.86214, 1.55357)h = -4.85t^2 + 14.9t + 1.55 \qquad (\text{full values } {-4.85},\ 14.86214,\ 1.55357)

(b) The residuals (observed minus predicted, using full values) are

tt000.50.5111.51.5222.52.5
residual0.0460.046−0.072-0.072−0.066-0.0660.1660.166−0.078-0.0780.0040.004
SSres=0.0462+(−0.072)2+⋯+0.0042≈0.0452SS_{res} = 0.046^2 + (-0.072)^2 + \cdots + 0.004^2 \approx 0.0452

(using the unrounded residuals; the GDC can also compute this sum from a list of residuals)

The GDC gives R2=0.999R^2 = 0.999 (to 3 s.f.; more precisely 0.999480.99948). So 99.9%99.9\% of the variability in the height is accounted for by the quadratic model: an excellent fit.

(c) Graph the model on the GDC. The maximum is at t≈1.53t \approx 1.53 s, with h≈12.9h \approx 12.9 m. The ball lands when h=0h = 0: the positive zero is t≈3.17t \approx 3.17 s. (The other zero, t≈−0.101t \approx -0.101, is outside the domain t≥0t \ge 0.)

The leading coefficient −4.85-4.85 is close to −4.9-4.9, half the acceleration due to gravity, which supports the model.

Example 2: Exponential regression for video views

Section titled “Example 2: Exponential regression for video views”

A new video’s daily views are recorded.

Day xx112233445566
Views yy2162163313315125127867861214121418671867
  • (a) Find an exponential regression model y=abxy = ab^x and interpret bb.
  • (b) Predict the views on day 88.
  • (c) On which day will the daily views first exceed 50005000, if the trend continues?

Solution.

(a) Exponential regression gives a≈140a \approx 140 and b≈1.54b \approx 1.54, with R2=1.00R^2 = 1.00 (to 3 s.f.):

y=140(1.54)xy = 140(1.54)^x

b=1.54b = 1.54 means the daily views are multiplied by about 1.541.54 each day: an increase of about 54%54\% per day.

(b) Using the stored model with full values, y(8)≈4430y(8) \approx 4430 views (3 s.f.).

(c) Solve 140(1.54)x=5000140(1.54)^x = 5000 with the GDC (using full values): x≈8.28x \approx 8.28. Since xx counts whole days, the views first exceed 50005000 on day 9. This is an extrapolation: real viral growth always slows down eventually, so this prediction should be treated with caution.

Example 3: Two models with almost the same R²

Section titled “Example 3: Two models with almost the same R²”

A patient’s blood concentration CC (mg/L) of a drug is measured every hour after an injection.

tt (h)001122334455667788
CC (mg/L)40.140.131.931.925.725.720.420.416.516.513.013.010.610.68.38.36.86.8
  • (a) Find a quadratic regression model and an exponential regression model, with R2R^2 for each.
  • (b) Which model should be used to estimate the concentration at t=12t = 12? Justify your answer.

Solution.

(a) Quadratic regression: C=0.440t2−7.55t+39.5C = 0.440t^2 - 7.55t + 39.5, with SSres≈1.60SS_{res} \approx 1.60 and R2=0.998R^2 = 0.998.

Exponential regression: C=40.0(0.800)tC = 40.0(0.800)^t, with SSres<0.1SS_{res} \lt 0.1 and R2=1.00R^2 = 1.00 (to 3 s.f.; a GDC’s exponential regression shows about 0.99980.9998).

(b) Both models fit the data extremely well: their R2R^2 values differ by only about 0.00150.0015. The deciding factors are the context and the behaviour beyond the data:

  • The quadratic has its minimum at t≈8.58t \approx 8.58 h and then increases, predicting the concentration will rise again with no new dose. That’s not reasonable.
  • The exponential keeps decreasing toward 00, as the body steadily clears the drug. A constant ratio of 0.8000.800 means 20%20\% is removed each hour, which is a standard model for drug clearance.

So use the exponential: with the full values from a GDC, C(12)≈2.76C(12) \approx 2.76 mg/L. (Software that minimizes SSresSS_{res} directly gives 2.752.75, so about 2.82.8 mg/L either way.) Even then, t=12t = 12 is an extrapolation, so it’s an estimate.

Nine data points of drug concentration from 40.1 mg/L at t = 0 to 6.8 mg/L at t = 8 h, with an exponential regression curve and a quadratic regression curve. Both fit the data closely, but after t = 8 the quadratic turns and rises while the exponential keeps decreasing toward 0. 2 4 6 8 10 12 14 5 10 15 20 25 30 35 40 quadratic exponential end of data time t (h) concentration C (mg/L)
Inside the data the two curves are almost identical; beyond t=8t = 8 only the exponential behaves sensibly.

Example 4: Sine regression for daylight hours

Section titled “Example 4: Sine regression for daylight hours”

The average number of daylight hours yy in a Canadian city in month xx (with x=1x = 1 for January) is shown.

xx112233445566778899101011111212
yy (h)9.19.110.410.411.911.913.513.514.914.915.615.615.315.314.114.112.512.510.910.99.59.58.88.8
  • (a) Find a sine regression model y=asin⁡(bx+c)+dy = a\sin(bx + c) + d.
  • (b) Find the period of the model and comment.
  • (c) Estimate the daylight hours in mid-April (x=4.5x = 4.5).

Solution.

(a) With the GDC in radian mode, sine regression gives

y=3.37sin⁡(0.516x−1.62)+12.2y = 3.37\sin(0.516x - 1.62) + 12.2

with SSres≈0.0470SS_{res} \approx 0.0470 and R2=0.999R^2 = 0.999.

(b) The period is 2πb=2π0.51607≈12.2\dfrac{2\pi}{b} = \dfrac{2\pi}{0.51607} \approx 12.2 months. That’s very close to 1212 months, the length of a year, so the model makes sense. The amplitude 3.373.37 h is close to half the difference between the longest and shortest days, and d=12.2d = 12.2 h is the average daylight over the year.

(c) y(4.5)≈14.3y(4.5) \approx 14.3 hours (using full values).

Choosing the type from R2R^2 alone. A higher R2R^2 doesn’t guarantee a better model. A cubic always fits at least as well as a quadratic on the same data, and a model can fit the data well but behave absurdly just outside it. Use context, the shape of the data and the behaviour of the curve as well.

Comparing SS_res across different data sets. SSresSS_{res} depends on the units and on how many points there are. Only compare SSresSS_{res} for different models of the same data.

Misinterpreting R². R2=0.85R^2 = 0.85 means 85%85\% of the variability in yy is accounted for by the model. It does not mean 85%85\% of the points lie on the curve, or that the predictions are 85%85\% accurate.

Using degree mode for sine regression. GDC sine regression works in radians. In degree mode you’ll get nonsense (or an error).

Rounding the parameters before predicting. Store the regression equation in your GDC and use the full values. Rounding b=1.53995b = 1.53995 to 1.541.54 and then raising it to a large power changes the answer.

Trusting extrapolation. Regression describes the data you have. Predictions far outside the range of xx (especially with polynomial or exponential models) can be wildly wrong.

1. (Warm-up) The model y=x2+2y = x^2 + 2 is proposed for the data below. Find SSresSS_{res}.

xx11223344
yy2.82.86.36.311.211.217.917.9
Solution

Predicted values: 33, 66, 1111, 1818. Residuals: −0.2-0.2, 0.30.3, 0.20.2, −0.1-0.1.

SSres=0.04+0.09+0.04+0.01=0.18SS_{res} = 0.04 + 0.09 + 0.04 + 0.01 = 0.18

2. (Warm-up) A power regression of a bird’s wingspan on its mass gives R2=0.91R^2 = 0.91. Interpret this value.

Solution

91%91\% of the variability in wingspan is accounted for by the power model of wingspan against mass; the other 9%9\% is due to other factors (or random variation).

3. (Warm-up) Which type of regression would you try first for each situation?

  • (a) Daily maximum temperature over two years.
  • (b) The number of bacteria in a dish over the first few hours.
  • (c) The height of a basketball during a shot.
  • (d) The time for a ball to fall from different heights, where tt appears to be proportional to a power of hh.
Solution

(a) Sine: the temperature repeats every year.

(b) Exponential: bacteria multiply at a constant rate at first.

(c) Quadratic: projectile motion.

(d) Power: t=ahbt = ah^b.

4. (Core) A café tests different prices pp (dollars) for a lunch special and records its weekly profit PP (dollars).

pp ($)101015152020252530303535
PP ($)8208201490149017501750190019001610161011701170
  • (a) Find the quadratic regression model and R2R^2.
  • (b) Use the model to find the price that gives the maximum profit, and the maximum profit.
Solution

(a) P=−5.54p2+262p−1230P = -5.54p^2 + 262p - 1230 (full values −5.53571-5.53571, 262.021262.021, −1232.71-1232.71), with R2=0.992R^2 = 0.992 (and SSres≈6740SS_{res} \approx 6740).

(b) The vertex (using the GDC maximum, or p=−b2ap = -\dfrac{b}{2a} with full values) is at p≈23.67p \approx 23.67, so a price of about $23.70 (3 s.f.), giving a maximum weekly profit of about $1870.

5. (Core) The number of users UU (thousands) of a new app mm months after launch is shown.

mm001122334455
UU (thousands)12.012.015.115.119.119.124.024.030.330.338.138.1
  • (a) Find an exponential regression model U=abmU = ab^m, and the monthly percentage growth.
  • (b) Use the model to find when the app will reach 100100 thousand users.
Solution

(a) U=12.0(1.26)mU = 12.0(1.26)^m (both parameters to 3 s.f.; R2=1.00R^2 = 1.00). Since b≈1.26b \approx 1.26, the number of users grows by about 26%26\% per month.

(b) Solve 12.0(1.26)m=10012.0(1.26)^m = 100 with full values: m≈9.17m \approx 9.17 months after launch (during the tenth month).

6. (Core) The time tt (s) for a ball to fall from height hh (m) is measured.

hh (m)11225510102020
tt (s)0.4520.4520.6390.6391.0101.0101.4281.4282.0192.019
  • (a) Find a power regression model t=ahbt = ah^b.
  • (b) Predict the time to fall 4545 m.
Solution

(a) t=0.452h0.500t = 0.452h^{0.500} (3 s.f.), with R2=1.00R^2 = 1.00. So tt is proportional to h\sqrt{h}.

(b) t(45)≈0.452×450.500≈3.03t(45) \approx 0.452 \times 45^{0.500} \approx 3.03 s (using full values).

7. (Core)

  • (a) A linear model has Pearson’s correlation coefficient r=−0.92r = -0.92. Find R2R^2 and interpret it.
  • (b) Two models are fitted to the same data. Model A has SSres=12.4SS_{res} = 12.4 and model B has SSres=3.1SS_{res} = 3.1. Which fits the data better?
  • (c) For a data set, SStot=250SS_{tot} = 250 and a model has SSres=20SS_{res} = 20. Use R2=1−SSresSStotR^2 = 1 - \dfrac{SS_{res}}{SS_{tot}} to find R2R^2.
Solution

(a) For a linear model R2=r2=(−0.92)2=0.8464≈0.846R^2 = r^2 = (-0.92)^2 = 0.8464 \approx 0.846. About 84.6%84.6\% of the variability in yy is accounted for by the linear model.

(b) Model B: its smaller SSresSS_{res} means its predictions are closer to the data. (Whether it’s the better model also depends on context.)

(c) R2=1−20250=1−0.08=0.92R^2 = 1 - \dfrac{20}{250} = 1 - 0.08 = 0.92.

8. (Challenge) A sunflower’s height HH (cm) is measured weekly.

Week ww001122334455667788
HH (cm)10101818333358589292130130162162183183195195
  • (a) Find quadratic and cubic regression models, with R2R^2 for each.
  • (b) Use each model to predict the height at week 1212.
  • (c) Which model would you choose for predicting the height at week 1212? Explain.
Solution

(a) Quadratic: H=0.596w2+21.3w−0.879H = 0.596w^2 + 21.3w - 0.879, R2=0.980R^2 = 0.980.

Cubic: H=−0.749w3+9.59w2−5.81w+11.7H = -0.749w^3 + 9.59w^2 - 5.81w + 11.7, R2=0.999R^2 = 0.999.

(b) Using full values: quadratic H(12)≈341H(12) \approx 341 cm; cubic H(12)≈27.9H(12) \approx 27.9 cm.

(c) Neither. The cubic has the higher R2R^2 but predicts the sunflower shrinks from 195195 cm to about 2828 cm, which is impossible. The quadratic predicts growth speeding up forever, but the data shows growth already slowing down (the weekly increases fall from 3838 cm to 1212 cm). The shape suggests the height is levelling off, so a model with a horizontal asymptote, such as a logistic model, would be more suitable. This is a clear example of why R2R^2 alone is not enough.

9. (Challenge) The height of the tide yy metres at a harbour is recorded every 22 hours, starting at midnight (t=0t = 0).

tt (h)00224466881010121214141616181820202222
yy (m)3.93.93.13.11.41.40.60.61.21.22.82.83.93.93.33.31.61.60.60.60.90.92.52.5
  • (a) Find a sine regression model y=asin⁡(bt+c)+dy = a\sin(bt + c) + d with a>0a \gt 0 and b>0b \gt 0, and its R2R^2.
  • (b) Find the period of the tide according to the model.
  • (c) Predict the height of the tide at 1 a.m. the next day.
Solution

(a) In radian mode: y=1.69sin⁡(0.508t+1.58)+2.19y = 1.69\sin(0.508t + 1.58) + 2.19, with R2=0.998R^2 = 0.998.

(b) Period =2π0.50837≈12.4= \dfrac{2\pi}{0.50837} \approx 12.4 hours (a typical period for ocean tides).

(c) 1 a.m. the next day is t=25t = 25: y(25)≈3.86y(25) \approx 3.86 m (using full values). This is a short extrapolation of a regular periodic pattern, so it’s reasonably reliable.