skip to content

Model Interpretability

Explaining a fitted model: permutation importance and partial dependence globally, Shapley values and counterfactuals for one row, audits of who it fails. Interviewers make you defend a black box.

on this pageshow

explore

questions

page 1 of 2

Why does removing race and gender columns from the training data fail to make a model fair?

level: juniorimportance: must knowfreq 70%

answer

  1. the column leaves, the signal stays
  2. ZIP code, school, purchase history
  3. redundant encoding, not one feature
  4. you removed the audit, not the bias

basics

~20 s

Other features reconstruct the removed attribute. ZIP code, school, job history and purchase patterns correlate with it, so the model relearns it indirectly. Dropping the column removes your ability to measure the disparity, not the disparity itself.

solid answer

~50 s

Deleting the protected column is called *fairness through unawareness*, and it fails because real feature sets are redundant. In an auto-insurance pricing model, ZIP code and census tract carry so much residential-segregation signal that a model with no race column can still price by race in effect — the digital version of redlining. A resume screener that never sees gender can learn to downweight a women's college and a women's chess club, because those tokens predict the label the historical data encoded. Individually weak proxies also combine: three features that each correlate 0.2 with the attribute can jointly reconstruct it well. The practical consequence is worse than a no-op — once the column is gone you cannot compute selection rates or error rates by group, so you have removed the audit rather than the bias. Keep the attribute as an audit-only column, and measure.

go deeper

for a junior

Be ready to name the failure by its label, fairness through unawareness, and give one concrete proxy such as ZIP code or school. Knowing that correlated features re-encode the attribute is the whole screening bar here.

for a middle

Explain the mechanism, not just the slogan: several weakly correlated features can jointly reconstruct an attribute that none of them reveals alone, which is why a per-feature correlation screen passes a model that is still encoding it.

for a senior

Show you have run the audit. Describe predicting the attribute from the remaining features, keeping it as an access-restricted audit-only column, and measuring the disparity before and after any feature removal rather than assuming the removal worked.

for a principal

Own the policy call: whether to collect protected attributes at all, who may read them, and how you defend a predictive-but-correlated feature as business necessity versus retiring it. Frame it as a documented tradeoff, not a technical fix.

## The claim being tested "We don't collect race, so the model can't be racist." This position has a name in the fairness literature — **fairness through unawareness** — and it is the first thing an interviewer wants you to be able to take apart. A **protected attribute** is a characteristic that law or policy forbids you from using as a basis for a decision: race, sex, age, disability, religion, national origin, and in some jurisdictions pregnancy or marital status. Fairness through unawareness is the policy of simply not putting that column in the feature matrix. ## Why it fails: redundant encoding A feature is a **proxy** for a protected attribute when it carries information about that attribute. A model does not need the attribute itself; it needs anything correlated with it plus a label whose historical pattern differs by group. Three mechanisms produce proxies: 1. **Geography.** Residence is strongly segregated in many countries, so ZIP code, census tract, neighbourhood, or even distance-to-office is a partial encoding of race and often of income. An auto-insurance pricing model built on ZIP-level features can produce systematically different prices by race with no race column anywhere in its training table — the mechanism people call digital redlining. 2. **Institutions and affiliations.** A women's college, a women's chess club, a religious school, a national-origin-linked professional body: each is close to a deterministic marker of the attribute. A resume screener trained on which past applicants were hired can learn to downweight exactly those tokens if the historical hiring pattern did. 3. **Behaviour.** Purchase categories, browsing, device type, name spelling, and language of use all shift by group. The subtle version is **joint reconstruction**. No single feature has to look suspicious. Several features that are individually weakly correlated with the attribute can, in combination, reconstruct it with high accuracy — this is ordinary supervised learning applied to the attribute rather than to your label. That is why a per-feature correlation screen is not a sufficient audit. ## How to actually check The direct test is to treat the protected attribute as a target: hold the attribute out of the model's feature set, then fit a separate model that tries to *predict the attribute* from the remaining features. If that model beats the base rate substantially — say it separates the groups far better than chance — your feature set encodes the attribute, and any downstream model is free to use that encoding. Supporting checks: compare each feature's distribution across groups, and look at how much a feature's predictive contribution changes when you condition on group. Running this test requires you to *have* the attribute. That is the second lesson of unawareness: not collecting the attribute at all makes the audit impossible while leaving the mechanism intact. The workable arrangement is to collect it, exclude it from the feature set, and store it as an audit-only column with restricted access, used for measurement and nothing else. ## What to do once you find a proxy Dropping the proxy is not an automatic win. Three considerations: - **The disparity may survive.** Remove ZIP and the remaining features often re-encode much of it. You measure the disparity before and after rather than assuming. - **You may destroy legitimate signal.** ZIP code genuinely predicts theft and collision risk for reasons that are not race. Removing it costs accuracy for everyone, including members of the group you meant to protect. - **The question is what the feature is doing.** A feature that predicts the outcome only through group membership is a pure proxy. A feature with a defensible causal path to the outcome is a business-necessity argument you can make out loud. Interviewers care that you can tell these apart and can say which one you think you have. ## What good sounds like in an interview Name the failure mode ("fairness through unawareness"), give one concrete proxy mechanism, make the joint-reconstruction point so it is clear you know a single-feature correlation screen is not enough, and finish on measurement: the reason to keep the attribute is to compute outcomes by group, not to feed it to the model. That last move is what separates a candidate who has read about fairness from one who has had to run an audit.

  • How would you test whether the remaining features can reconstruct a protected attribute you excluded?
    Treat the attribute as a label: fit a model predicting it from the feature set you actually train on, and compare its discrimination against the base rate. If it separates the groups well, the attribute is encoded. A per-feature correlation screen is weaker, because several individually weak features can reconstruct the attribute jointly.
  • If a proxy feature is genuinely predictive of the outcome, is dropping it the right move?
    Not automatically. Ask whether the feature predicts the outcome only through group membership or through a defensible causal path. Dropping a pure proxy is cheap; dropping a genuinely predictive one costs accuracy for every group and often leaves the disparity in place because other features re-encode it. Measure the disparity before and after rather than assuming.
  • Does collecting the protected attribute contradict the rule against using it in the model?
    No. Collection and use are separate decisions. The usual arrangement is to collect the attribute, keep it out of the feature set, and hold it as an audit-only column with restricted access so you can compute selection and error rates by group. Without it you cannot show anyone, including a regulator, that the model is or is not disparate.

Taking the name off an application does not hide who wrote it when the letterhead, the postcode and the club membership are all still on the page.

saying these in an interview costs you the question

  • Says the model cannot be biased if the column is absent
  • Screens features one at a time and calls the set clean
  • Assumes any correlated feature must simply be deleted
  • Refuses to collect the attribute, so no audit is possible
  • Treats fairness as a training-time setting with no measurement

context

open as a page

A partial dependence curve of predicted electricity demand against outdoor temperature is U-shaped — what does that tell you?

level: juniorimportance: must knowfreq 42%

basics

~20 s

The model has learned a non-monotone relationship: average predicted demand is highest at cold temperatures and at hot temperatures, and lowest in between. The curve summarises the model's average behaviour over the whole dataset, not any single building.

open as a page

What makes a model interpretable by design rather than explained after the fact?

level: juniorimportance: must knowfreq 62%

basics

~20 s

An interpretable-by-design model is one whose fitted structure is itself the explanation: a sparse weighted sum, a depth-3 tree, a short rule list. Post-hoc methods leave the model opaque and build a separate, approximate account of it.

open as a page

How is permutation importance computed for a single feature, and what does the number mean?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Permutation importance shuffles one column's values across held-out rows, breaking that column's link to the label, then re-scores the already-trained model. The drop from the baseline score, averaged over several shuffles, is the feature's importance.

open as a page

How do demographic parity and equalized odds differ as group fairness criteria?

level: middleimportance: must knowfreq 60%

basics

~20 s

Demographic parity requires the same positive-prediction rate in every group, ignoring the true outcome. Equalized odds requires equal true-positive and false-positive rates in every group, comparing errors only among people who share a true label.

open as a page

Where can you intervene to shrink a model's between-group disparity: before, during, or after fitting?

level: middleimportance: must knowfreq 58%

basics

~20 s

Three points. Pre-processing reweights or repairs the training data, in-processing adds a fairness penalty or constraint to the training objective, and post-processing changes the decision rule per group after fitting. Each needs different access and pays a different accuracy cost.

open as a page

How do you compute the four-fifths disparate-impact ratio for a hiring screen?

level: middleimportance: must knowfreq 55%

basics

~20 s

Compute each group's selection rate — the share of that group's applicants the screen passes — then divide the lowest by the highest. A ratio under 0.8 is the conventional flag for adverse impact and calls for justification.

open as a page

Why is an adverse-action notice saying 'your score was below the cutoff' inadequate?

level: middleimportance: must knowfreq 58%

basics

~20 s

An adverse-action notice must name the specific facts about the applicant that drove the decline, such as a revolving balance at 78% of the limit. A score and a cutoff describe the mechanism, not the reason.

open as a page

How is a partial dependence curve for one feature computed from a trained model and a dataset?

level: middleimportance: must knowfreq 60%

basics

~20 s

Choose a grid of values for that feature. For each grid value, overwrite the feature with it in every row, score the whole dataset with the model, and average the predictions. Plot grid value against average prediction.

open as a page

Why can two highly correlated features both look unimportant under permutation importance?

level: middleimportance: must knowfreq 62%

basics

~20 s

Because shuffling one of them leaves its twin intact, and the model recovers the same information from the twin, so the score barely moves. Each correlated column masks the other's credit, and both end up ranked near zero.

open as a page

What does it mean for a local explanation of one prediction to be faithful rather than plausible?

level: middleimportance: must knowfreq 60%

basics

~10 s

Faithful means the explanation reflects what the model actually computed for that row. Plausible means it matches what a human expects. An explanation can read convincingly while the model keys on something else.

open as a page

How is a feature's Shapley value computed for a single model prediction?

level: middleimportance: must knowfreq 62%

basics

~20 s

A feature's Shapley value is its average marginal contribution: for every ordering of the features, measure how much adding that feature moves the prediction away from a baseline, then average those changes across all orderings.

open as a page

How does a local surrogate model explain one prediction of a black-box classifier?

level: middleimportance: must knowfreq 60%

basics

~10 s

A local surrogate perturbs the row being explained, labels each perturbation with the black box's prediction, weights it by closeness to that row, then fits a small interpretable model whose coefficients are the explanation.

open as a page

What does a model card record about a model's intended use and training population?

level: juniorimportance: should knowfreq 44%

basics

~20 s

A model card is a short document shipped with a model stating what it is for, what it must not be used for, who and what the training data represented, how it performs on relevant subgroups, and its known limits.

open as a page

As a counterfactual explanation for a demoted marketplace seller, why prefer '3 fewer late shipments' to '400 more orders'?

level: juniorimportance: should knowfreq 38%

basics

~20 s

Both changes may flip the trust model's decision, but a counterfactual is only useful if the seller can act on it. Three fewer late shipments is a small change on one controllable feature; 400 extra lifetime orders may take years.

open as a page

Why can permutation importance rank features differently from a tree ensemble's built-in importance?

level: middleimportance: should knowfreq 55%

basics

~20 s

They measure different things. Built-in importance is read off the training-time structure of the fit in split-criterion units; permutation importance measures the drop in an evaluation metric on data you choose. Different quantity, different data, different ranking.

open as a page

Why can independently perturbing correlated features produce a misleading local explanation?

level: middleimportance: should knowfreq 45%

basics

~20 s

Perturbing features one at a time builds rows no real case resembles, so the model is queried where it never learned. Redundant columns also share the credit, and how it splits is a property of the method.

open as a page

What do the Shapley axioms - efficiency, symmetry, dummy and additivity - guarantee?

level: middleimportance: should knowfreq 44%

basics

~20 s

Efficiency makes the attributions sum exactly to the prediction minus the baseline; symmetry gives interchangeable features equal credit; dummy gives an unused feature zero; additivity makes attributions of summed models add. Together they pick out one rule.

open as a page

How can a risk score be calibrated within every group yet still show unequal false-positive rates?

level: seniorimportance: should knowfreq 44%

basics

~20 s

When two groups have different base rates, a score meaning the same thing in both must produce different error rates. Equal predictive value and equal false-positive rates cannot both hold, so the two criteria genuinely conflict.

open as a page

When is a different score threshold per protected group defensible in a deployed model?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Rarely, and never as an engineering decision alone. Per-group thresholds act directly on the measured gap but read the protected attribute at decision time, which is explicit differential treatment. Ship one only with counsel's sign-off, a recorded group attribute, and monitoring.

open as a page

A protected-group slice of 38 applicants shows a 12-point pass-rate gap. Is the gap real?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Not established. With 38 people, a 95% interval on that group's pass rate spans roughly 30 percentage points, and the interval on the 12-point gap itself runs from about minus 4 to plus 29 points. Report intervals, not point estimates.

open as a page

In a store-demand model, the partial dependence curve for discount depth is flat while its ICE curves fan out — what does that mean?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Stores respond differently and the responses cancel. Partial dependence is the pointwise average of the per-row ICE curves, so steeply rising stores and falling or flat stores average to a flat line. The feature matters — just not uniformly.

open as a page

When may a shallow tree fitted to a black-box model's predictions be quoted as its explanation?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A global surrogate is trained on the black box's own outputs, so it may be quoted only when it reproduces that model closely on held-out inputs from the deployment population, and closely within each segment you discuss.

open as a page

A readable pneumonia-risk rule list learned that asthma lowers risk — how do you respond?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The rule is a true pattern in the data and a lethal policy. Asthmatic pneumonia patients were routed straight to intensive care, so they survived more often. The label reflects the treatment they received, not their underlying risk.

open as a page

Agent ID tops a routing model's training permutation importance but is near zero on held-out data — why?

level: seniorimportance: should knowfreq 47%

basics

~20 s

The model memorised agent ID on the rows it was fitted to. Shuffling it on training data destroys that memorised fit, so the score collapses; on unseen rows the memorisation was never worth anything, so scrambling it costs nothing.

open as a page

A sampling-based explainer returns different top-3 features on two runs for one prediction. What do you do?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Separate estimator noise from real ambiguity: rerun across many seeds, measure rank agreement, and raise the sample budget until each spread is small next to the gaps you rank.

open as a page

How do you estimate Shapley values when exact enumeration over 2^40 coalitions is impossible?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Sample random feature orderings instead of enumerating all coalitions: average each feature's marginal contribution over the sampled orderings. The estimate is unbiased, and its error falls roughly as one over the square root of the number of draws.

open as a page

How does the neighbourhood kernel width change a local surrogate's explanation of the same prediction?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The kernel width sets how local the explanation is. Too wide, and the surrogate drifts toward the model's average behaviour; too narrow, and nearly all weight falls on a few near-identical points, so a handful of samples decide the answer.

open as a page

A fairness fix closes a 9-point between-group false positive rate gap but costs 1.5 AUC points. Do you ship it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Not from those two numbers alone. Convert both sides into decision consequences and money, check the tradeoff curve for a cheaper knee, confirm the gap is real and not an artefact of biased labels, and put the call to a named accountable group rather than deciding it as an engineer.

open as a page

Your interpretable scorecard scores 2 AUC points below a boosted model — how do you decide which to ship?

level: principalimportance: should knowfreq 45%

basics

~20 s

Decide from what the decision requires, not the metric gap: ship the readable scorecard when the logic must be signed off, contested or hand-edited, after confirming the 0.02 AUC difference is real out of time.

open as a page

showing 1–30 of 40