Why can independently perturbing correlated features produce a misleading local explanation?
answer
- real rows fill a thin region
- vary one, leave the region
- the model still returns a number
- duplicated information, divided credit
- the split is a method tie-break
basics
~20 sPerturbing features one at a time builds rows no real case resembles, so the model is queried where it never learned. Redundant columns also share the credit, and how it splits is a property of the method.
solid answer
~50 sTwo things go wrong. First, off-manifold queries: if you vary features independently you construct impossible combinations - a wearable-health risk model asked about a resting heart rate of 42 paired with a running pace of 12 km/h, a pairing its training data never contained. The model still returns a number, but that number is unconstrained extrapolation, so an explanation built from those responses describes behaviour on inputs that will never arrive. Second, arbitrary credit splitting: an e-commerce model carrying both `total spend` and `orders x average order value` has the same information twice, so whatever influence exists is divided between them, and the size of each share is a tie-break the method performs rather than a fact about the customer. Practical mitigations: draw replacement values from real rows instead of independently, group redundant columns and report one number per group, and check the story survives dropping a duplicate.
go deeper
Know that real feature values are correlated, so changing one while freezing the rest can create combinations that never occur in practice.
Explain both failures: the model extrapolating on impossible rows, and two columns encoding the same information having no well-defined individual credit.
Show the tradeoff you would take in production - on-manifold sampling spreads credit to proxies, grouping loses detail - and say which you chose for which audience.
Own the upstream call: duplicated encodings in a feature set buy little accuracy and cost explainability, so decide where the team pays for grouping versus pruning.
## Where the perturbation goes Most local explanation methods work by asking counterfactual questions of the model: what would it have predicted if this feature had a different value, or had been replaced by a reference value? To answer, they must construct new input rows and score them. The quality of the explanation therefore depends on where those constructed rows land. Real feature vectors do not fill the input space. They concentrate on a much thinner region - the data manifold - because features are correlated. Height and weight move together; tenure and lifetime value move together; a resting heart rate and a training load move together. If a method varies one feature while holding the rest fixed, it walks straight off that region. **A concrete case.** A wearable-health risk model uses resting heart rate, average running pace, weekly active minutes and age. An explainer sweeps resting heart rate from 40 to 100 while pinning the other columns at this user's values. Somewhere in that sweep it constructs a row with a resting heart rate of 42 and a running pace of 12 km/h - a combination the training data essentially never contains. The model has no idea what to do there. It will still emit a risk score, because models always emit a score; the score is whatever the fitted function happens to do in a region no training example constrained. Tree ensembles extrapolate as flat surfaces there, linear models extrapolate straight lines, boosted models can do something jagged. None of it is knowledge. The explanation is then an accurate summary of the model's behaviour on rows that will never be scored in production. It can be internally consistent and completely uninformative about the decision you were trying to understand. ## Where the credit goes The second problem survives even if every query is on-manifold. Suppose an e-commerce churn model is given `total spend` and, separately, `orders` and `average order value`. The product of the last two is the first. The information is present twice. Whatever influence that information has on the prediction now has to be expressed across two or three columns, and how it gets divided is decided by the method's internals, not by the customer's behaviour. Change the method, change the reference values, or refit the model with the columns in a different order, and the split can move - while the model's prediction and the group's total influence are unchanged. Read the individual numbers as a ranking and you will tell a stakeholder that spend matters less than order count, which is not a finding about anything. The general statement: **local attributions are identified only up to the redundancy in the feature set.** Correlated features do not have well-defined individual credit, and no amount of extra sampling produces one, because the ambiguity is not statistical noise - it is a property of the question. ## Mitigations, and what each one costs **Sample replacement values from real data rather than independently.** Instead of sweeping a feature freely, draw substitutes from rows that resemble the target row, or condition the replacements on the features you are holding fixed. This keeps queries on-manifold. The cost is that it changes the question being answered: when you respect correlation, a feature the model never reads still receives credit because its correlated partner carries it. You have traded a nonsense-input problem for a credit-flows-to-proxies problem. Which you want depends on the purpose - reasoning about what the model computes points one way, reasoning about what a change of profile would predict points the other. Say which you chose and why. **Group redundant features and report the group.** If three columns encode one construct, report one contribution for the construct. This removes the meaningless internal split and usually makes the explanation easier for a stakeholder to act on. Grouping requires domain knowledge about which columns belong together; a correlation matrix or a clustering of the columns is a starting point, not the answer. **Prune the redundancy at the source.** Often the cleanest fix is upstream: do not carry both `total spend` and its two factors. Fewer duplicated encodings, fewer meaningless splits, and usually no loss of predictive performance. **Test robustness before presenting.** Recompute the explanation with a duplicate column removed, or with a different reference set. If the narrative flips, do not present the narrative. ## The interview answer Name both failure modes separately - impossible input rows, and undefined credit between redundant columns - because candidates who name only one usually think the fix is 'more samples'. It is not; more samples make an off-manifold estimate more precise, and precision on the wrong question is still the wrong answer.
- Does drawing replacement values from real rows instead of independently fix the problem outright?It fixes the impossible-row half and creates a different effect. Once replacements respect the correlation structure, a feature the model never reads receives credit because its correlated partner carries the signal. So you swap unconstrained extrapolation for credit flowing to proxies. Neither is wrong in the abstract - pick based on whether you are describing the model's mechanism or the predictive value of a profile, and state the choice alongside the explanation.
- How would you present an attribution when two columns encode the same information?Report one contribution for the pair rather than two competing numbers, label it by the construct the columns share, and say explicitly that the internal split is not identifiable. If a stakeholder needs the split for a decision, that is a signal to change the feature set upstream rather than to argue about the split. Better still, drop the duplicate encoding entirely and re-explain.
- Would running the explainer with far more samples resolve either problem?It resolves neither. More samples shrink the estimator's sampling error, so you converge faster to whatever the method targets - but if the queries are off-manifold you converge precisely to a description of nonsense inputs, and if two columns are redundant the credit split has no true value to converge to. The remedies are on-manifold sampling, grouping and feature-set surgery, not budget.
Asking a chef what a dish tastes like without salt is fair; asking what it tastes like with the salt of one recipe and the cooking time of another is a question about a meal nobody ever served.
saying these in an interview costs you the question
- Reads the split between duplicate columns as a real ranking
- Perturbs every feature independently and calls it the model's logic
- Assumes the model behaves sensibly on inputs it never saw
- Thinks more samples fix off-manifold queries
- Claims an impossible row is fine because the model returns a score