skip to content

Interactions and Log Terms

An interaction lets one predictor's slope depend on another, and logging a variable turns its coefficient into a percentage or an elasticity. Interviewers love asking you to read these aloud.

on this pageshow

questions

5

In a log(wage) regression, how do you interpret a coefficient of 0.07 on years of schooling?

level: juniorimportance: must knowfreq 66%

answer

  1. the outcome is not in its original units
  2. think proportional, not additive
  3. the coefficient is not currency
  4. exponentiate to get back to wage
  5. exp(b) - 1 is the exact percent

basics

~20 s

One more year of schooling is associated with roughly a 7% higher wage, holding the other predictors fixed. Because the outcome is logged, the coefficient reads as an approximate percent change; the exact figure is exp(0.07) - 1, about 7.25%.

solid answer

~50 s

With the outcome on a log scale, the coefficient is an absolute change in `log(wage)` per extra year of schooling, which translates into a *relative* change in wage. Adding one year multiplies wage by `exp(0.07) = 1.0725`, so about a 7.25% increase, and since the coefficient is small the shortcut "about 7% per year" is close enough for a first read. That percentage is percent of the worker's own wage, so the same coefficient means a bigger absolute gain for a high earner than for a low one. The rule of thumb frays as the coefficient grows: at 0.5, the honest number is `exp(0.5) - 1 = 64.9%`, not 50%. And this is an association within the fitted model, holding the other regressors fixed, not a causal return to schooling unless the design supports that claim.

go deeper

for a junior

Be ready to say immediately that a logged outcome turns the coefficient into a percent change, and to give the exact conversion exp(b) - 1. Do not report the number in the outcome's original units.

for a middle

Explain the mechanics: the difference in predicted log values exponentiates into a ratio, so the model is multiplicative in the original scale, and show where the small-coefficient approximation starts to mislead.

for a senior

Demonstrate that you know exponentiating fitted values does not return the mean, and that a currency answer requires naming a baseline. Interviewers look for the retransformation point here.

for a principal

Own the choice of scale itself: whether the business question is about proportional or absolute effects, what reporting a percent does to how stakeholders compare segments, and when the interpretability cost of logging the outcome is not worth it.

## Why a logged outcome changes the reading In an ordinary linear regression, a slope answers "how many units of y per one unit of x". When the outcome has been replaced by its natural logarithm, the model is `log(wage) = b0 + b1 * schooling + ... + error` and the slope `b1` answers a different question: how many units of *log wage* per one extra year. Log units are not interesting on their own, so the standard move is to convert them back into a statement about wage itself. ## The exact conversion Compare two workers who differ by one year of schooling and are identical on every other predictor in the model. Their predicted log wages differ by `b1`, so `log(wage_2) - log(wage_1) = 0.07` Exponentiating both sides, `wage_2 / wage_1 = exp(0.07) = 1.0725`. The predicted wage is multiplied by 1.0725, which is a 7.25% increase. In general the exact percent change per unit is `100 * (exp(b) - 1)` A logged outcome therefore makes the model *multiplicative* in the original units: each unit of a predictor scales the outcome by a constant factor rather than adding a constant amount. ## The small-coefficient shortcut For small values, `exp(b) - 1` is very close to `b`, because `exp(b) = 1 + b + b^2/2 + ...`. That is why practitioners read 0.07 as "about 7%" without apology. The size of the error grows quickly, though: - b = 0.02 gives 2.02% (shortcut 2%) - b = 0.07 gives 7.25% (shortcut 7%) - b = 0.20 gives 22.1% (shortcut 20%) - b = 0.50 gives 64.9% (shortcut 50%) A useful discipline: quote the approximation for coefficients under roughly 0.1, and compute `exp(b) - 1` for anything larger, especially for dummy predictors whose coefficients are often big. ## Percent of what The percent is always relative to that observation's own predicted baseline. If one worker is predicted at 20,000 and another at 200,000, an extra year of schooling adds about 1,450 to the first and about 14,500 to the second. This is exactly why a log outcome is chosen for quantities like wages, prices, revenue or counts of visits: the belief is that effects are proportional rather than additive. Stakeholders who want an answer in currency need a stated baseline; "7% of whatever you currently earn" is the honest phrasing, and any single dollar figure is a percent evaluated at one chosen point. ## Predicting back on the original scale A subtlety that separates careful analysts from careless ones: exponentiating a fitted log value does not give the predicted *mean* of the outcome. `exp(E[log y])` is a median-like quantity, not `E[y]`, because the exponential is a convex function and Jensen's inequality applies. Under the common assumption that the errors on the log scale are normal with variance s^2, the mean on the original scale is `exp(fitted) * exp(s^2 / 2)`. Ignoring that factor produces predictions that are systematically low. For interpreting a coefficient as a percent change this does not matter, because the correction factor is the same for both workers and cancels in the ratio; it matters as soon as someone asks for a forecast in currency. ## Related model forms The same logic gives the whole family of readings: - level-level (`y` on `x`): b units of y per unit of x. - log-level (`log y` on `x`): roughly `100 * b` percent change in y per unit of x, exactly `100 * (exp(b) - 1)`. - level-log (`y` on `log x`): roughly `b / 100` units of y per 1% change in x. - log-log (`log y` on `log x`): b percent change in y per 1% change in x. ## What interviewers listen for They want to hear that the coefficient is not in currency, that the effect is proportional, that `exp(b) - 1` is the exact conversion, and ideally that the approximation has a range of validity. Saying "wage goes up by 0.07" is the answer that fails the screen, because it reports a log-scale quantity as if it were money.

  • At what point does the "multiply by 100" shortcut break down, and what do you report instead?
    It breaks down as the coefficient grows, because it is a first-order approximation to `exp(b) - 1`. At 0.07 the gap is trivial (7% versus 7.25%), at 0.20 it is 20% versus 22.1%, and at 0.50 it is 50% versus 64.9% - a serious misstatement. My rule is to quote the shortcut below about 0.1 and report `100 * (exp(b) - 1)` above that, which matters most for dummy predictors whose coefficients are often large.
  • If you exponentiate the fitted values, what quantity are you actually predicting?
    Not the mean of the outcome. Exponentiating a fitted log value gives a median-like prediction, because the exponential is convex and the expectation does not pass through it. Under normal errors on the log scale with variance s^2, the mean on the original scale is `exp(fitted) * exp(s^2 / 2)`. Skipping that retransformation factor makes currency forecasts systematically too low, even though it cancels out when you interpret a coefficient as a ratio.
  • A stakeholder wants the effect of schooling in dollars, not percent. What do you tell them?
    That the model does not carry a single dollar answer, because a log outcome makes the effect proportional. I would say it is about 7% of that person's own wage, then translate at one or two explicit baselines: roughly 1,450 a year at a 20,000 baseline and roughly 14,500 at a 200,000 baseline. Quoting one dollar figure without naming the baseline hides the fact that the number varies across the sample.

A log-outcome coefficient behaves like an interest rate: it tells you the percentage your balance grows per period, not a fixed number of coins, so the same rate means more money to whoever already has more.

saying these in an interview costs you the question

  • Says wage rises by 0.07 currency units per year of schooling
  • Reports the coefficient in the outcome's raw units
  • Uses exp(b) rather than exp(b) - 1 as the percent change
  • Claims the 100 times b shortcut is exact at any size
  • Treats exponentiated fitted values as the predicted mean

context

open as a page

In a sales model with a promotion x weekend interaction, what does the promotion coefficient mean?

level: middleimportance: must knowfreq 72%

basics

~20 s

It is the promotion effect when the weekend indicator equals zero, that is on weekdays only. The interaction coefficient is the extra promotion effect on weekends, so the weekend promotion effect is the sum of the two.

open as a page

In a log-log demand model, what does a price coefficient of -1.4 tell you?

level: middleimportance: must knowfreq 58%

basics

~20 s

It is a price elasticity: a 1% price increase is associated with roughly a 1.4% fall in quantity, holding the other predictors fixed. Because the size exceeds 1, demand is elastic, so raising price lowers revenue.

open as a page

Why center age and income before adding their interaction to a regression?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Centering makes each main effect the effect at the other predictor's mean rather than at zero, which is often outside the data, and it removes the artificial correlation between main effects and their product. Fit is unchanged.

open as a page

In a model with advertising and advertising squared, how do you read diminishing returns?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Neither coefficient stands alone. The marginal effect of spend is b1 + 2b2spend, so a negative squared term means each extra unit buys less than the last, and the fitted curve turns downward at spend equal to -b1 / (2*b2).

open as a page