skip to content

In a store-demand model, the partial dependence curve for discount depth is flat while its ICE curves fan out — what does that mean?

level: seniorimportance: should knowfreq 45%

answer

  1. the average is not the individuals
  2. opposite slopes cancel
  3. one line per row, averaged
  4. anchor every curve at zero first
  5. look for what splits the fan

basics

~20 s

Stores respond differently and the responses cancel. Partial dependence is the pointwise average of the per-row ICE curves, so steeply rising stores and falling or flat stores average to a flat line. The feature matters — just not uniformly.

solid answer

~50 s

A flat partial dependence curve with fanning individual conditional expectation curves is the classic signature of heterogeneous effects. Partial dependence is the pointwise mean of the ICE curves, so if half the stores' predicted weekly units climb steeply with discount depth while the rest barely move or drift down, the average is flat although no single store is. The wrong conclusion is that the model ignores discount depth; the right one is that its effect is conditional on something else. Next I centre the ICE curves — subtract each curve's value at the leftmost grid point so every line starts at zero — which strips out baseline level differences and leaves only slope differences. Then I colour by candidate moderators such as format, region or baseline traffic to find what splits the fan. A single global discount rule read off that flat line would be wrong for both halves of the estate.

go deeper

for a junior

Know that a partial dependence curve is an average and that per-row curves can differ from it. Being able to say that opposite responses cancel in an average already answers most of this.

for a middle

Explain the pointwise-average relationship between the two plot types, and describe what centring the per-row curves at the left edge of the grid removes and what it leaves visible.

for a senior

Show the diagnostic loop end to end: spot the cancellation, centre, colour by candidate moderators, confirm the split is stable rather than model noise, and stop reporting a single global curve.

for a principal

Decide what the organisation ships when effects are segment-dependent: whether a global policy is defensible at all, which segmentation is worth operating, and how much heterogeneity justifies the extra complexity.

## What each object is An **ICE curve** (individual conditional expectation) is one line per row. For a given store's row, you sweep the feature — discount depth — across a grid, hold every other value in that row fixed at what it actually is, and plot that store's prediction at each grid value. With a thousand stores you get a thousand lines. The **partial dependence curve** is the pointwise average of those lines. That identity is the whole answer to this question: an average is flat whenever the things being averaged cancel. ## Reading the fan A fan can mean two different things, and distinguishing them is the point of centring. **Vertical spread** — the lines sit at different heights but run roughly parallel. That is not heterogeneity of *effect*; it is just that a flagship store predicts 8,000 units and a small format predicts 900. Discount depth moves both by the same amount. **Slope spread** — the lines cross, splay apart, or run in opposite directions. That is heterogeneity of effect: the model believes the same discount does different things at different stores. Raw ICE plots mix the two, and the vertical spread is usually much larger, so it visually dominates. **Centred ICE** fixes this: pick an anchor, normally the leftmost grid value, and subtract each curve's value there from the whole curve. Every line now starts at zero and its height at any later point reads directly as "how much this store's prediction has moved since the smallest discount". Slope differences become the only thing left on the plot. With a flat average and centred curves that still splay, the diagnosis is unambiguous: real, opposing per-row responses. ## Why the average went flat Suppose 40% of stores have centred curves rising to +300 units at deep discounts, 40% sit near zero, and 20% drift to -200 because the model has learned that deep discounts there mostly pull forward demand or coincide with clearance of slow stock. The mean across all stores can easily land within noise of zero at every grid point. The feature is doing a great deal of work in the model; the marginal average simply is not the right summary of that work. This is also why "the partial dependence curve is flat, so the feature is inert" is a genuinely dangerous inference. It has the same shape as concluding a drug does nothing because the average of a group it helps and a group it harms is zero. ## What to do next **Find the moderator.** Colour the centred ICE curves by candidate features — store format, urban versus rural, baseline traffic, own-brand share, season. When one colouring cleanly separates the risers from the non-responders, you have found the interacting feature, and you can plot partial dependence for discount depth *within* each group and get two honest, non-flat curves. **Quantify rather than eyeball.** Friedman's H-statistic measures how much of the joint two-feature partial dependence is not explained by adding the two single-feature curves together, i.e. how non-additive the pair is. Near zero means the pair acts additively; larger values mean a substantial share of the joint behaviour is interaction. It is expensive and it is noisy when either main effect is weak, so treat it as a ranking device across candidate pairs rather than a precise number. **Keep the plot readable.** Thousands of lines become a black smear. Sample a few hundred rows, or draw the deciles of the curve family as a small number of bands with the average on top, so the reader sees the spread without the ink. **Change what you report.** If effects genuinely differ by segment, a single global curve is the wrong artefact for a decision-maker. Report per-segment curves, or report the average with the spread drawn around it, and say explicitly that the average describes no actual store. ## The trap to avoid All of this describes the *model*. The fan says the model's predictions respond differently across stores; it does not by itself establish that discounting would produce those different outcomes if you ran the promotion. Keep the claim at the level the artefact supports: this is what the model has learned, and it is enough to justify not shipping one global discount rule read off a flat line.

  • How would you quantify that interaction instead of eyeballing the fan?
    Friedman's H-statistic. It compares the two-feature partial dependence surface with the sum of the two single-feature curves; the statistic captures the share of the joint behaviour that the additive combination fails to explain, so near zero means additive and larger means genuinely interacting. It is expensive to compute and unstable when the main effects are weak, so I use it to rank candidate pairs, then confirm the winner visually.
  • Why centre the ICE curves rather than plot them raw?
    Raw curves are separated mostly by each store's baseline prediction level, and that vertical spread is usually far larger than the slope differences you care about. Anchoring every curve at zero at the left edge of the grid removes the level and leaves only how far each prediction moves, which is the actual question.
  • With 5,000 stores, how do you keep the plot usable?
    Plot a random sample of a few hundred curves, or summarise the family by drawing a handful of quantile bands with the average overlaid. Then colour by one suspected moderator at a time. The goal is to show the spread and its structure, not to render every line.
  • Could a flat average with fanning curves come from noise rather than a real interaction?
    Yes, especially with a high-variance model or few rows per store. I would check whether the fan is stable across refits or bootstrap resamples of the training data, and whether the splitting moderator is one the business finds plausible. A fan that reshuffles on every refit is model variance, not a discovered interaction.

Averaging a store whose sales climb with discounts and one whose sales sag gives a flat line — the same way averaging a rising and a falling line gives a horizontal one that describes neither.

saying these in an interview costs you the question

  • Concludes the model ignores discount depth because the average is flat
  • Treats vertical spread of raw curves as effect heterogeneity
  • Assumes the average shape describes every individual store
  • Reports the flat curve as a single global discount policy
  • Never asks which feature separates the responders

context