skip to content

In a model with advertising and advertising squared, how do you read diminishing returns?

level: seniorimportance: nice to knowfreq 34%

answer

  1. the slope is no longer a constant
  2. differentiate the fitted equation
  3. the sign of the squared term sets the shape
  4. set the derivative to zero
  5. watch the factor of 2 and the minus

basics

~20 s

Neither coefficient stands alone. The marginal effect of spend is b1 + 2b2spend, so a negative squared term means each extra unit buys less than the last, and the fitted curve turns downward at spend equal to -b1 / (2*b2).

solid answer

~50 s

In `sales = b0 + b1*ad + b2*ad^2`, the derivative with respect to spend is `b1 + 2*b2*ad`, so the effect of another unit of advertising depends on how much you already spend. With `b1 > 0` and `b2 < 0` the curve is concave: returns stay positive but shrink, which is the shape people mean by diminishing returns. Setting the derivative to zero gives the turning point `ad* = -b1 / (2*b2)`, the spend beyond which the fitted curve declines. Before that number leaves my desk I check whether `ad*` falls inside the range of spend actually observed - if it sits far above anything ever spent, the model says returns are flattening, not where the peak is. I report marginal effects at a few representative spend levels rather than quoting `b1`, which is only the slope at zero spend.

go deeper

for a junior

Know that adding a squared term makes the effect depend on the level of the predictor, and that a negative squared coefficient means each extra unit adds less than the previous one.

for a middle

Derive the marginal effect b1 + 2b2x and the turning point at minus b1 over 2*b2, and be able to say why the linear coefficient alone means the slope at zero.

for a senior

Demonstrate the range check: whether the turning point sits inside observed spend, whether a symmetric parabola can represent saturation, and how you would present marginal effects instead of raw coefficients.

for a principal

Own the decision risk: how much evidence you require before a fitted curve sets a budget, and whether the honest answer is to run spend experiments rather than extrapolate an observational fit.

## Two coefficients, one curve Adding a squared term makes the relationship a parabola: `sales = b0 + b1*ad + b2*ad^2 + error` The model is still linear in its parameters, so ordinary least squares fits it without any change of machinery - what changes is the interpretation. The effect of one more unit of spend is the derivative `marginal effect = b1 + 2*b2*ad` which depends on the current level of spend. There is no single "effect of advertising" to report, exactly as with an interaction; in fact a squared term is a variable interacted with itself, and the same reading rules apply. ## Reading the two coefficients - `b1` is the marginal effect at `ad = 0` - the slope at zero spend. Alone it is rarely a quantity anyone cares about, and quoting it as "the effect of advertising" is the standard error. - `b2` sets the curvature. Negative means concave, the diminishing-returns shape; positive means convex, accelerating returns. Its size controls how fast the slope changes: each extra unit of spend changes the marginal effect by `2*b2`. ## The turning point Setting `b1 + 2*b2*ad = 0` gives `ad* = -b1 / (2*b2)` With `b1 > 0` and `b2 < 0` this is a maximum: the fitted curve rises, flattens, then falls. For example `b1 = 8` and `b2 = -0.05` give `ad* = -8 / (-0.1) = 80`. Getting the factor of 2 right matters - it comes from differentiating the square - and so does the sign, since forgetting the minus sign flips the answer. ## Where honest analysis starts The formula is easy; the judgment is the interview. Three things to check before treating `ad*` as an optimum: 1. **Is it inside the data?** If observed spend runs from 0 to 40 and `ad*` is 80, the quadratic never actually turns within the data, and the estimate of the peak is pure extrapolation. The defensible claim is "returns are still positive but flattening across the observed range". 2. **Is the downturn supported?** A parabola is forced to be symmetric, so a fitted decline on the right can be an artefact of a few high-spend points or of the functional form rather than a real saturation effect. Real response curves usually saturate toward a ceiling rather than reversing; a quadratic cannot represent that shape and will invent a decline to approximate flatness. 3. **Is this a causal quantity?** Spend levels are usually chosen, not randomised, so the curve describes how sales and spend moved together in the past. Calling `ad*` an optimal budget imports a causal claim the fit alone does not license. ## Presenting it Rather than a coefficient table, report the marginal effect at a few representative spend levels - say the 25th, 50th and 90th percentiles of observed spend - each with an interval. That communicates the shape without inviting anyone to read `b1` as the effect. A plot of fitted sales against spend over the observed range does the same job, clipped at the range boundary so no one extrapolates off the end. ## Model-building details - Keep the linear term whenever the square is in the model. Dropping it forces the vertex to sit at zero spend, a constraint you almost certainly do not believe. - The linear and squared terms are strongly correlated on raw data; centering spend before squaring reduces that artificial correlation and makes the linear coefficient the slope at the mean spend rather than at zero. The curvature coefficient and the fit are unchanged. - Judge the pair jointly rather than only by the squared term's individual t-statistic, and be wary of adding higher powers - cubic and beyond wave around at the edges of the data and rarely reflect anything real. ## What interviewers listen for The derivative `b1 + 2*b2*ad`, the sign of `b2` as the source of the shape, the turning point `-b1 / (2*b2)` with the factor of 2 correct, and - most of all - the discipline to check that the turning point falls inside the observed range before anyone budgets against it.

  • The turning point lands far above any spend you have observed. What do you report?
    That returns are positive and flattening across the observed range, and nothing about a peak. A turning point outside the data is an extrapolation of a functional form, not an estimate the data supports, and presenting it as an optimal budget invites a decision the model cannot back. I would report marginal effects at spend levels the business has actually run, and note that finding the saturation point would need experimentation at higher spend.
  • Why can't you interpret the linear advertising coefficient on its own?
    Because with a squared term present the marginal effect is `b1 + 2*b2*ad`, so `b1` is just the slope at zero spend - a level the business never operates at. Quoting it as "the effect of advertising" overstates returns wherever spend is meaningful, since the concave term pulls the slope down as spend rises. The interpretable quantities are marginal effects evaluated at representative spend levels.
  • Is the squared term's t-test enough to justify keeping the curvature?
    It is evidence, not a verdict. I would look at whether the curvature is plausible in the domain, how much data sits at the high-spend end driving it, and whether the implied shape is credible - a parabola is symmetric and cannot represent saturation toward a ceiling, so it may fake a decline to approximate flatness. I would also keep the linear term regardless, since dropping it pins the vertex at zero spend.

It is like fertiliser on a field: the first bags raise the yield a lot, later bags add less, and past some point more fertiliser actively hurts - one number cannot describe all three regions.

saying these in an interview costs you the question

  • Reports the linear coefficient as the advertising effect
  • Treats a negative squared coefficient as a defect in the model
  • Calls the turning point optimal spend without checking the data range
  • Extrapolates the parabola well beyond observed spend
  • Forgets the factor of 2 when differentiating the squared term
  • Keeps the squared term after dropping the linear one

context