skip to content

How is a partial dependence curve for one feature computed from a trained model and a dataset?

level: middleimportance: must knowfreq 60%

answer

  1. a grid, then an average
  2. overwrite the column in every row
  3. other columns keep their real values
  4. rows times grid points, many predictions
  5. the mean of the per-row curves

basics

~20 s

Choose a grid of values for that feature. For each grid value, overwrite the feature with it in every row, score the whole dataset with the model, and average the predictions. Plot grid value against average prediction.

solid answer

~50 s

You pick a grid over the feature's observed range, usually its quantiles so points are not wasted in sparse tails. For each grid value you take a copy of the dataset, overwrite that one column with the grid value in every row, leave every other column at its real value, score all rows with the trained model, and average the predictions. That average is the plotted point. Formally `PD(v) = (1/n) * sum_i f(v, x_other_i)` — the model's prediction with the feature pinned at `v`, marginalised over the empirical distribution of the other features. Equivalently, it is the pointwise average of the per-row curves you would get by sweeping the feature for each row individually. The cost is rows times grid points model evaluations, so on large data people score a random sample of rows. The procedure implicitly assumes the swept feature is not strongly tied to the others, since it evaluates every row at every grid value.

code

python · 12 lines
python
# a stand-in trained model and three rows of real data
def predict(temperature, floor_area):
    return 0.4 * abs(temperature - 18.0) + 0.01 * floor_area

rows = [(5.0, 120.0), (20.0, 300.0), (31.0, 80.0)]

for grid_value in (0.0, 9.0, 18.0, 26.0, 34.0):
    # the feature of interest is forced to grid_value in EVERY row,
    # every other column keeps its own observed value
    predictions = [predict(grid_value, area) for _temperature, area in rows]
    partial_dependence = sum(predictions) / len(predictions)
    print(grid_value, round(partial_dependence, 2))

go deeper

for a junior

Recall the three moves in order: fix the feature at a value in every row, score the data, average the predictions. Knowing the y-axis is an averaged model output is most of the credit here.

for a middle

Be able to write the average out and say what it marginalises over, distinguish it from averaging only nearby rows, and quote the cost as rows times grid points.

for a senior

Demonstrate the operating habits: sampling rows on large data, quantile grids, probabilities rather than labels for classifiers, and drawing data density so nobody over-reads an extrapolated arm.

for a principal

Frame when this cheap global summary is the right artefact at all, and set the team convention for grid choice, sampling and centring so two people's plots of the same model are comparable.

## The procedure, step by step Partial dependence is a recipe, not a model property. Given a trained model `f`, a dataset of `n` rows, and one feature of interest: 1. **Build a grid.** Choose values to evaluate the feature at — commonly 20 to 50 points placed at quantiles of the observed values, or evenly spaced across the observed range. Quantile grids avoid spending half the plot on a sparse tail. 2. **For each grid value `v`:** copy the dataset, overwrite the feature's column with `v` in *every* row, and leave all other columns exactly as they were. 3. **Score.** Run the model over that whole modified dataset. 4. **Average.** Take the mean of those `n` predictions. That single number is the curve's height at `v`. 5. **Plot** grid value against averaged prediction, and repeat for the whole grid. Written out: `PD(v) = (1/n) * sum over i of f(v, x_other_i)`, where `x_other_i` are row `i`'s remaining feature values, untouched. ## What that average actually estimates The averaging step is a **marginalisation over the empirical distribution of the other features**. You are asking: if I force this feature to `v` but let the rest of the world look the way my data looks, what does the model say on average? That is why every row participates at every grid value — the other columns are the sample you are averaging over. This is also the sharpest contrast with the naive alternative. Averaging the model's predictions (or the observed targets) among *only* the rows whose feature is already near `v` gives a **conditional** average, a different quantity. Under correlation it absorbs the effect of everything that moves with the feature, and credits the feature with all of it. ## Relationship to per-row curves If instead of averaging you keep one line per row — sweep the grid for row `i` with row `i`'s other values fixed, and plot that row's predictions — you get the individual conditional expectation curves. Partial dependence is exactly their **pointwise mean**. That identity is worth stating in an interview because it explains what the averaging can hide: opposite-signed per-row slopes cancel and leave a flat average. ## Practical details that come up **Cost.** The curve costs `n x grid` model evaluations. Ten thousand rows and a 40-point grid is 400,000 predictions for one feature; for a slow model or a whole dashboard of features that matters. The standard remedy is to score a random sample of rows — a few hundred to a few thousand — which changes the estimate's variance but not what it estimates. **Categorical features.** The grid is simply the set of categories, and the result is one averaged prediction per level, usually drawn as bars rather than a line. Ordering the bars by value rather than alphabetically makes them readable. **Classification.** For a classifier you average a *probability* (or a log-odds score), not the hard label, otherwise the curve becomes a step function that hides most of the structure. **Two features at once.** The same recipe with a two-dimensional grid gives a surface or heatmap over a pair of features, which is how interactions are inspected visually. **Centring.** Curves are sometimes shifted so their mean over the grid is zero. That changes only the y-axis reference, never the shape; be explicit about which convention a plot uses before comparing two of them. **Show the data density.** Because every row is evaluated at every grid value, the curve is drawn just as confidently in regions with almost no data as in dense ones. A rug or decile ticks under the axis is the cheapest honesty device available. ## The assumption baked into step 2 Overwriting the column in every row means the model is scored on rows that combine the grid value with other feature values it may never actually co-occur with. When the swept feature is roughly independent of the rest, that is harmless. When it is strongly tied to another feature, some of those synthetic rows are combinations the data never contains, and the average is partly made of predictions in regions the model never learned. Recognising that limit — and reaching for a method that only compares predictions locally when it bites — is the mark of someone who has drawn these curves in anger rather than only read about them.

  • Why score the whole dataset at each grid value instead of only the rows already near it?
    Because those are two different quantities. Averaging over all rows marginalises over the other features and isolates the model's response to this one. Averaging only nearby rows gives a conditional average, which under correlation absorbs the effect of every feature that moves with the one you are plotting, and hands it to the wrong column.
  • How do you choose the grid, and how many points?
    Quantiles of the observed values, typically 20 to 50 points. Quantiles put resolution where the data is instead of stretching the plot across an empty tail, and they cap the influence of a single extreme value. I also draw decile ticks under the axis so a reader can see which parts of the curve are supported and which are extrapolation.
  • What changes for a categorical feature or for a classifier?
    For a categorical feature the grid is the category set, so you get one averaged prediction per level, drawn as bars. For a classifier, average the predicted probability or the score, never the hard label — averaging labels turns a smooth relationship into a step function and throws away the interesting part.

saying these in an interview costs you the question

  • Says the curve plots the average observed target per feature value
  • Scores only rows whose feature already sits near the grid value
  • Sets the other features to their means instead of keeping them
  • Forgets the cost is rows times grid points predictions
  • Thinks the curve is read out of the model's coefficients

context