How is a feature's Shapley value computed for a single model prediction?
answer
- fair split of a joint payout
- coalitions of features, not one feature
- add a feature, measure the jump
- average that jump over every ordering
- missing features come from a baseline
basics
~20 sA feature's Shapley value is its average marginal contribution: for every ordering of the features, measure how much adding that feature moves the prediction away from a baseline, then average those changes across all orderings.
solid answer
~50 sTreat the prediction as a payout that a team of features earned together, and the Shapley value as each feature's fair share. You define a value function `v(S)`: the model's output when the features in the coalition `S` take this row's actual values and every feature outside `S` takes values from a background or baseline row. A feature's marginal contribution to `S` is `v(S union {i}) - v(S)`. Its Shapley value is the average of that marginal contribution over all orderings in which the features could be added, which is the same as a weighted average over all subsets `S` that exclude the feature. Averaging over orderings is what handles interactions: a feature added first often looks different from the same feature added last, and averaging splits shared credit evenly. The resulting attributions add up to the prediction minus the baseline prediction.
code
python · 21 linesfrom itertools import permutations
def f(x): # toy model: x0 and x1 interact, x2 is never used
return 2*x[0] + 3*x[1] + 4*x[0]*x[1] + 0*x[2]
row, base = (1, 1, 1), (0, 0, 0)
def v(S): # coalition value: S takes row values, the rest the baseline
return f(tuple(row[i] if i in S else base[i] for i in range(3)))
phi = [0.0, 0.0, 0.0]
orders = list(permutations(range(3)))
for order in orders:
S = set()
for i in order:
before = v(S)
S.add(i)
phi[i] += (v(S) - before) / len(orders)
print([round(p, 3) for p in phi]) # [4.0, 5.0, 0.0]
print(round(sum(phi), 3), v({0, 1, 2}) - v(set())) # 9.0 9go deeper
Be ready to say in one sentence what an attribution is for: it splits one prediction into per-feature pieces measured against a baseline, and the pieces are averages of how much each feature moves the output.
You are expected to write the mechanics down: define the coalition value, the marginal contribution, and the average over orderings, and explain why interactions force the averaging rather than a single leave-one-out measurement.
Show you know what has to be pinned down before an attribution is trustworthy in production - which background the absent features come from, that the value is per-row, and that the same feature can attribute in opposite directions on different rows.
Own the framing question: what business question does 'why this prediction?' actually mean here, what is the honest reference point, and when is a per-row attribution the wrong instrument for the decision being made.
## The question a local attribution answers A trained model turns one row of features into one number - a predicted price, a probability, an ETA. A local attribution splits that number into per-feature pieces so a human can read *why this row scored the way it did*. Shapley attribution is one specific, principled way to do that split, borrowed from cooperative game theory and brought into machine learning as the basis of SHAP. ## The cooperative game underneath Imagine a marketing team runs one joint campaign across email, search and display, and the campaign brings in a fixed amount of revenue. How much of that revenue belongs to each channel? You cannot simply ask what each channel earned on its own, because the channels reinforce each other: display drives awareness that makes search convert. The cooperative-game answer is to consider every *coalition* of channels, ask what each coalition would have earned, and pay each channel its average marginal contribution when it joins. A model prediction has exactly the same shape. The features are the players. The prediction is the payout. The only extra piece you need is a definition of what a *coalition* earns. ## The value function For a coalition `S` (a subset of the features), define ``` v(S) = the model's output when the features in S are set to this row's values and every feature outside S is filled in from a background (baseline) ``` The background can be a single reference row - an all-zeros row, the training median row - or a set of rows that you average over. Two anchors follow immediately: - `v(empty set)` is the baseline prediction, the model's output when nothing about this row is known. - `v(all features)` is the actual prediction for this row. The whole attribution problem is now: split `v(all) - v(empty)` among the features. ## Marginal contribution and the average over orderings A feature `i`'s marginal contribution to a coalition `S` that does not already contain it is ``` v(S union {i}) - v(S) ``` This number depends on `S`. If a model uses `tenure` and `monthly_charge` in an interaction, then adding `tenure` to an empty coalition may barely move the output while adding it after `monthly_charge` moves it a lot. There is no single "the" contribution. Shapley's answer is to average over every order in which the features could arrive. Pick a random ordering of all `n` features, walk along it, and record how much the prediction jumps each time a feature is added. Repeat for every one of the `n!` orderings and average the jumps for each feature: ``` phi_i = average over all orderings of [ v(predecessors of i, plus i) - v(predecessors of i) ] ``` Written as a sum over subsets rather than orderings, this is ``` phi_i = sum over S not containing i of w(|S|) * ( v(S + i) - v(S) ) with w(|S|) = |S|! * (n - |S| - 1)! / n! ``` The combinatorial weight is just the count of orderings in which exactly the members of `S` come before `i`, divided by `n!`. Both forms give the same number. ## A worked micro-example Take a model `f(a, b, c) = 2a + 3b + 4ab`, with `c` never used. Explain the row `(1, 1, 1)` against the baseline `(0, 0, 0)`. The coalition values are `v({}) = 0`, `v({a}) = 2`, `v({b}) = 3`, `v({c}) = 0`, `v({a,b}) = 9`, `v({a,c}) = 2`, `v({b,c}) = 3`, `v({a,b,c}) = 9`. Averaging marginal contributions over the six orderings gives `phi_a = 4`, `phi_b = 5`, `phi_c = 0`. The two main effects (2 and 3) go where they belong, the interaction of 4 is split evenly, `c` gets nothing, and the three attributions sum to `9 = v(all) - v(empty)`. ## What the number does and does not mean A Shapley value is a statement about the model's output, relative to a chosen baseline. "`monthly_charge` contributed +0.12 to this churn probability" means: averaged over the orders in which features could be revealed, knowing this row's `monthly_charge` instead of the background's pushes the model's output up by 0.12. It is not a claim that raising the charge in the real world would raise churn by 0.12, and it is not a global statement about the feature - the same feature can attribute positively on one row and negatively on another. ## Cost The definition ranges over `2^n` coalitions (or `n!` orderings), so exact enumeration is only practical for a handful of features. In practice the average is estimated by sampling orderings, or computed exactly by exploiting structure in particular model families. The cost is the price of the axiomatic guarantees the value carries.
- What does it actually mean for a feature to be 'absent' from a coalition when the model needs every input?Absence is simulated, not real. The features outside the coalition are filled in from a background - a single reference row, or a sample of rows that you average the model's output over. That substitution is a modelling choice, and it is why the baseline has to be stated alongside any attribution.
- Why average over all orderings instead of measuring each feature's effect on its own?Because interactions make a feature's effect depend on what is already known. Adding a feature to an empty coalition can move the output far less than adding the same feature last. A single ordering would hand all shared credit to whichever feature happened to arrive first; averaging over every ordering splits interaction credit evenly.
- Does a positive Shapley value mean increasing that feature would increase the prediction?No. It says that this row's value of the feature, rather than the background's, pushes the model's output up by that amount, averaged over coalitions. It describes the model relative to a baseline, not the effect of intervening on the world, and the sign can flip on a different row.
Three marketing channels run one joint campaign and it earns a fixed revenue. Each channel's fair share is what it adds, on average, across every order in which the channels could have been switched on.
saying these in an interview costs you the question
- Calls the attribution the feature's coefficient in the model
- Says it is simply the model's output with the feature dropped
- Forgets that absent features are filled from a baseline
- Reads a local attribution as a global importance ranking
- Treats a positive attribution as proof of a causal effect