skip to content

Why prefer accumulated local effects over partial dependence when two input features are strongly correlated?

level: seniorimportance: nice to knowfreq 26%

answer

  1. correlation makes fake rows
  2. no light engine in a heavy car
  3. local differences, not global averages
  4. accumulate interval by interval, then centre
  5. relative y-axis, not a predicted value

basics

~20 s

Partial dependence sets one feature to each grid value in every row, manufacturing combinations the data never contains, such as a tiny engine in a very heavy car. Accumulated local effects compare predictions only within each row's own neighbourhood.

solid answer

~40 s

Partial dependence overwrites the feature in every row, so in a fuel-economy model where engine displacement and curb weight move together, the grid asks the model about 1.0-litre engines in two-tonne cars. Those rows never occurred, the model never learned that region, and its answers there still land in the average. Accumulated local effects avoid the extrapolation: split the feature's range into intervals, and for the rows falling inside an interval, compute the change in prediction when that feature moves from the interval's lower edge to its upper edge, keeping each row's other values. Average those differences within the interval, accumulate across intervals, then centre. Every evaluation stays near a value the row actually had. ALE does not make correlation disappear, but it stops the curve resting on impossible inputs.

go deeper

for a junior

Recall the core failure: forcing one feature to every grid value in every row can create combinations that never occur in the data, and the model's answers there are guesswork that still lands in the curve.

for a middle

Explain the mechanics — quantile intervals, prediction differences at the interval edges for rows inside it, averaged and then accumulated — and say why the resulting y-axis is relative rather than absolute.

for a senior

Show the operating judgment: inspect the correlation structure first, plot both curves and compare shapes, tune the interval count for stability, and state plainly what the curve still cannot separate.

for a principal

Set the house rule for which effect plot is published under which conditions, and weigh a curve that fewer stakeholders understand against a cheaper one that is quietly wrong when inputs move together.

## The problem with sweeping one column Partial dependence is computed by pinning a feature at a grid value in **every** row and averaging the model's predictions. That is fine when the feature moves roughly independently of the others. It stops being fine when it does not. Take a fuel-economy model with **engine displacement** and **curb weight** as inputs. In the real fleet these move together: large engines sit in large, heavy vehicles. To draw the displacement curve, partial dependence still forces displacement to 6 litres in the row describing a 900 kg city car, and to 1.0 litres in the row describing a 2,300 kg pickup. Those rows are not in the data, are not on the road, and are places the model never learned anything about. Whatever the model happens to output there — the arbitrary extrapolation of a tree ensemble past its last split, say — is averaged into the plotted curve with full weight. The practical symptom is a displacement curve with a shape nobody can explain, driven by a region of input space that does not exist. ## Why the obvious fix is worse The intuitive repair is to stop making things up: at each value of displacement, average the predictions of only the rows that *already* have roughly that displacement. That conditional average is sometimes called an M-plot, and it introduces a different failure. Cars with 6-litre engines are also heavy, less aerodynamic and geared differently. Averaging their predictions gives you the combined behaviour of the whole correlated bundle, then labels the resulting curve "displacement". You have traded impossible inputs for misattributed credit. ## How accumulated local effects work ALE takes a third route: stay local, and measure **differences** rather than levels. 1. **Partition** the feature's range into intervals, usually by quantiles so each interval holds a similar number of rows. 2. **For each interval**, take the rows whose feature value falls inside it. For each such row, score it twice — once with the feature set to the interval's **lower** edge, once at its **upper** edge — keeping that row's other features untouched. Take the difference. 3. **Average** those differences within the interval. This is the local effect: how much the model's prediction moves when the feature nudges across this small interval, estimated only on rows that genuinely live there. 4. **Accumulate** the interval averages from left to right, a running sum, producing a curve. 5. **Centre** the accumulated curve so its mean over the data is zero. Because a row is only ever evaluated at the edges of the interval it already belongs to, the model is never asked about a 1.0-litre pickup. And because the quantity averaged is a *difference* computed within one row, everything else about that row cancels — the curve is not credited with the effect of curb weight the way a conditional average would be. ## Reading the result The y-axis is **relative**, not absolute. Because the curve was built by accumulating differences and then centred, a value of +2 mpg at some displacement means "about 2 mpg above the average prediction, attributable to this feature's accumulated local effect" — not "the model predicts 2 mpg here". Partial dependence, by contrast, is on the prediction's own scale. Comparing the two curves' heights directly is a beginner's error; compare their shapes. ## What ALE does not fix **It does not recover an independent effect that the data cannot identify.** If displacement and weight were perfectly collinear, no procedure could tell you what the model would do to a heavy car with a small engine, because nothing in the data ever showed it. ALE reports the effect of moving displacement *among the cars that actually have that displacement*, which is a meaningful and honest quantity, but it is a local, conditional one. **Interval width is a real knob.** Too few intervals over-smooths and can flatten genuine structure; too many leaves each local difference estimated from a handful of rows and the accumulated curve becomes visibly jagged. Quantile-based intervals with a floor on the row count per interval are the usual compromise. **It is still a model artefact.** Like any of these curves it describes what the model has learned from historical data, and it does not license a claim about what happens if you intervene on engine size. ## When to reach for which Check the correlation structure of the inputs first. With weakly correlated features the two curves agree closely and partial dependence is cheaper to compute and far easier to explain to a non-technical audience, so use it. When a feature is strongly tied to another, either plot the accumulated local effects instead, or plot partial dependence within strata where the correlation is broken. A candidate who reports both and shows they agree has done the cheap check that most people skip.

  • Does that fully solve the correlation problem?
    No. ALE reports the effect of moving the feature among rows that already sit at those values, which is a local, conditional quantity. If two features move together almost perfectly, no method can separate their contributions from data that never showed them apart. ALE removes the extrapolation onto impossible rows; it does not manufacture identification the data lacks.
  • Why not just average predictions among rows already near each feature value?
    That conditional average absorbs everything correlated with the feature. Cars with large engines are also heavier and less aerodynamic, so the curve credits displacement with the whole bundle. ALE avoids this by averaging within-row differences, so whatever else the row carries cancels out of the quantity being averaged.
  • How do you read the y-axis of an accumulated local effects curve?
    As a centred, relative change: the height at a point is how far the model's prediction sits above or below its average because of this feature's accumulated local effect. It is not a predicted value on the target's scale, which is why you compare its shape with a partial dependence curve rather than its level.
  • How many intervals would you use?
    Quantile intervals, enough that each holds a reasonable number of rows — a few dozen at minimum — and typically ten to fifty overall. Too few over-smooths and can hide a real bend; too many makes each local difference noisy and the accumulated curve visibly jagged. I check that the shape is stable when I change the count.

Partial dependence asks what the model would say about a compact car fitted with a six-litre engine. Accumulated local effects only ask what happens when a car's engine is slightly bigger than the ones it actually sits beside.

saying these in an interview costs you the question

  • Claims the two curves always agree so the choice does not matter
  • Thinks ALE removes the need to check feature correlation
  • Reads the ALE y-axis as an absolute predicted value
  • Says correlated features just mean dropping one before plotting
  • Recommends averaging predictions among nearby rows as the fix

context