skip to content

Your churn model reports an odds ratio of 2.0 — how do you explain that to a product manager?

level: seniorimportance: should knowfreq 48%

answer

  1. odds are not probabilities
  2. the baseline rate changes everything
  3. rare events behave differently from common ones
  4. convert at two starting probabilities
  5. report percentage points, not ratios

basics

~20 s

An odds ratio of 2 doubles the odds, not the probability. From a 1% baseline that lands near 2%, but from a 40% baseline it lands near 57%, not 80%. Report an average marginal effect in percentage points instead.

solid answer

~50 s

I would hand a product manager percentage points, not the odds ratio. "Twice as likely" is the wrong translation: an odds ratio of 2 multiplies the *odds* by 2, and what that does to the probability depends on the baseline. Starting at 1%, the odds go from 0.0101 to 0.0202 and the probability lands near 2.0% — close to a doubling, because for rare events odds and probability nearly coincide. Starting at 40%, the odds go from 0.667 to 1.333 and the probability lands near 57%, a 17-point rise, not 80%. So the same coefficient means very different things for a high-risk and a low-risk segment. I report the average marginal effect — "about 4 percentage points more churn on average" — plus predicted probabilities for two or three concrete customer profiles. The odds ratio and its interval stay in the appendix.

go deeper

for a junior

Recall that odds ratios multiply odds rather than probabilities, and be able to convert a probability to odds and back so you can check any claim about doubling.

for a middle

Explain why the mapping is baseline-dependent: odds and probability nearly agree for rare events and diverge as the baseline rises toward the ceiling at 1. Work through a worked conversion.

for a senior

Demonstrate that you choose the reporting quantity for the audience — an average marginal effect in percentage points, plus predicted probabilities at concrete profiles, with the ratio kept in the appendix.

for a principal

Own the standard for how model effects reach decision-makers: the scale used, whether an average or segment-level numbers are shown, and how firmly an association is separated from a causal claim.

## The translation that goes wrong An odds ratio of 2.0 says: multiply the odds by two. It does *not* say: multiply the probability by two. The phrase "twice as likely" describes a ratio of probabilities, and the two agree only in a special case. Work the arithmetic at two baselines. **Baseline probability 1%** ``` odds = 0.01 / 0.99 = 0.0101 new odds = 0.0101 * 2 = 0.0202 new p = 0.0202 / 1.0202 = 0.0198 -> about 2.0% ``` Here "twice as likely" is nearly right: 1.0% to 1.98%, a relative rise of 1.98x. **Baseline probability 40%** ``` odds = 0.40 / 0.60 = 0.667 new odds = 0.667 * 2 = 1.333 new p = 1.333 / 2.333 = 0.5714 -> about 57% ``` Here "twice as likely" is badly wrong: 40% to 57% is a rise of 17 percentage points and a relative rise of only 1.43x. Doubling the probability would have required 80%, which the model never claimed. The general reason: `odds = p/(1-p)` is close to `p` when `p` is small, so the odds scale and the probability scale nearly coincide for rare events and diverge sharply as the baseline grows. Push the baseline higher and the gap widens further — at a 70% baseline, an odds ratio of 2 gives about 82%, a 12-point move. The ceiling at 1 leaves less and less room. ## Consequence: one coefficient, many effects Because the mapping from odds to probability is nonlinear, a single odds ratio implies a *different* probability change for every starting point. Two customers with the same feature flag but different baseline risk experience different absolute effects, even though the model has exactly one coefficient for that flag. This is not a defect of the model; it is what a bounded outcome forces. But it means "the effect of the feature" is not a single number on the scale the business cares about. ## What to report instead: marginal effects For a logistic model, the derivative of the fitted probability with respect to a continuous predictor is ``` dp/dx = b * p * (1 - p) ``` which depends on where the row sits on the curve — largest at `p = 0.5` (where it equals `0.25 * b`) and near zero at either extreme. Two standard summaries follow: - **Average marginal effect (AME).** Compute `dp/dx` for every observation at that row's own fitted probability, then average across the sample. This answers: "if this predictor rose by one unit for everyone in our actual customer base, how much would the average churn probability move?" It is reported in percentage points and is usually the right number for a product decision. - **Marginal effect at the mean (MEM).** Evaluate the derivative once, at a hypothetical row holding every predictor at its average. It is cheaper but describes a customer who may not exist — an average of a binary region flag is not a real customer. For a binary predictor, the honest analogue is a discrete change rather than a derivative: predict each row twice, once with the flag off and once on, take the difference, and average. That gives "about 4 percentage points more churn on average" directly. A third option, often the most persuasive in a product review, is a small table of **predicted probabilities at named profiles**: a new monthly customer, an established annual customer, a high-usage enterprise account — each shown with and without the feature. It shows the variation across segments that a single average hides. ## Framing the conversation A workable script: "Accounts with this attribute churn more. For a typical account at our 12% baseline, having it moves the risk to about 21% — nine points higher. For our highest-risk segment, already near 40%, it moves them to about 57%. On average across the base it is about 4 points. The model's raw effect is an odds ratio of 2.0 with an interval of 1.4 to 2.9, which is in the appendix." That framing does four things: it names the scale (percentage points), it exposes the baseline-dependence rather than hiding it, it gives the uncertainty, and it keeps the technical quantity available without making a non-specialist parse it. ## Two caveats to keep in the answer - **The odds ratio is not a ratio of probabilities.** They coincide only when the event is rare in both groups, which is why the distinction matters most for common outcomes. - **These are associations, not causal effects**, unless the design earns the causal reading. A stakeholder who hears "raises churn by 4 points" will act as if intervening on the feature would move churn, so state whether that inference is licensed by how the data was collected.

  • When is an odds ratio a decent approximation to a relative change in probability?
    When the event is rare in both groups. For small p, p/(1-p) is close to p, so multiplying the odds is nearly the same as multiplying the probability — at a 1% baseline an odds ratio of 2 gives about 1.98%. As the baseline rises the odds ratio increasingly overstates the relative change, and above roughly 10% the gap is large enough to mislead.
  • How is an average marginal effect computed for a logistic model?
    For a continuous predictor, evaluate dp/dx = b times p times (1-p) at each observation's own fitted probability and average those values across the sample. For a binary predictor, predict each row with the flag off and on, take the per-row difference, and average. Either way the result is in percentage points.
  • What do you do when the marginal effect varies sharply across customer segments?
    Stop leading with the average. Report predicted probabilities for two or three named profiles that bracket the range, and say plainly that a single odds ratio maps to very different point changes at different baselines. An average that hides a 2-point effect in one segment and a 17-point effect in another will drive the wrong prioritisation.
  • Why avoid the phrase 'twice as likely' even when the outcome is rare?
    Because it commits you to a probability-ratio reading that the model did not estimate, and the same sentence becomes flatly wrong the moment the baseline rises or the audience applies it to a different segment. Naming the scale — odds, relative change, or percentage points — costs one extra clause and removes the ambiguity.

An odds ratio is like a currency exchange rate quoted on a scale nobody shops in. Doubling the price in that currency buys wildly different amounts depending on where you start.

saying these in an interview costs you the question

  • Says an odds ratio of 2 means twice the probability
  • Reports raw log-odds coefficients to a non-technical audience
  • Treats an odds ratio as a percentage-point change
  • Assumes the probability effect is identical at every baseline
  • Presents an association to stakeholders as a causal lever

context