skip to content

Interpretable by Design

Some models are their own explanation: a linear coefficient, a shallow tree's path, a rule list, a per-feature shape function. Interviewers ask what accuracy you would trade to keep that.

on this pageshow

questions

5

What makes a model interpretable by design rather than explained after the fact?

level: juniorimportance: must knowfreq 62%

answer

  1. the model is the explanation
  2. same object deployed and read
  3. no second, approximating description
  4. sparse weights, shallow tree, short rules

basics

~20 s

An interpretable-by-design model is one whose fitted structure is itself the explanation: a sparse weighted sum, a depth-3 tree, a short rule list. Post-hoc methods leave the model opaque and build a separate, approximate account of it.

solid answer

~50 s

Interpretable by design means the thing you deploy and the thing you explain are the same object. If the model is a handful of weighted terms, a three-level tree, or an ordered list of `if ... then ...` rules, a reader can trace any prediction through the fitted parameters and see the entire decision logic at once, with no approximation in between. Post-hoc explanation runs the other way: you keep an unreadable model and fit a second description on top of it, so that description can be wrong about the model without anything visibly breaking. The important nuance is that interpretability is not a property of the model family. A 400-term linear model over engineered features, or a tree grown to depth 20, is as unreadable as anything else. What actually buys readability is sparsity, a small number of terms, and features a domain expert can name.

go deeper

for a junior

Be ready to define the two categories and name one model that falls in each. Have a concrete fitted form ready to describe out loud, such as a three-level tree or a five-term weighted score.

for a middle

Explain what makes an intrinsic explanation faithful by construction and why a post-hoc one has to have its faithfulness measured. Show that size and feature quality, not the algorithm name, decide readability.

for a senior

Demonstrate that you have hit the limit in practice: the linear model that became unreadable once the feature engineering grew, or the tree you had to cap in depth to keep reviewable. Say how you kept a model readable under pressure to add features.

for a principal

Own the policy question: which decisions in your organisation require an intrinsically readable model, who signs off on that classification, and what you do when a team wants to replace a readable production model with an opaque one.

## The distinction There are two ways to end up with an explanation of a model. **Interpretable by design (intrinsic).** You choose a model class whose fitted form a human can read directly, and you keep it small enough to stay readable. The explanation is not produced by a separate tool; it *is* the model. Examples: - A **sparse linear model**: `score = -2.1 + 0.4*years_at_address + 1.3*prior_default - 0.02*age`. Reading the fitted weights is reading the model. - A **shallow decision tree**: a depth-3 tree has at most eight leaves, so every possible decision path fits on one page. A warehouse supervisor can route parcels for manual inspection off a laminated card carrying the whole tree, and can tell you afterwards exactly which three questions decided each parcel. - A **rule list**: an ordered sequence such as `if A then class 1; else if B then class 0; else class 1`. The reader walks it top to bottom and stops at the first match. - A **generalized additive model**: one fitted curve per feature, summed. Non-linear, still readable one feature at a time. **Post-hoc (extrinsic).** You train whatever is most accurate, accept that its internals are not human-readable, and afterwards run a separate procedure that produces a description of it — an attribution method such as SHAP or LIME, a perturbation study, or a simpler model fitted to imitate it. The description is an approximation, and its quality is an empirical question, not a guarantee. ## Why the difference matters An intrinsic explanation is **faithful by construction**. If the weight on `prior_default` is 1.3, then that is what the model does; there is no gap between the account and the mechanism. A post-hoc explanation has a gap, and the gap is invisible in normal operation: a plausible-looking bar chart of feature contributions can be produced for a model whose actual behaviour is quite different, and nothing errors out. Two different post-hoc methods can even disagree about the same model and prediction, and you have no referee. A second, practical difference: an interpretable model can be **edited**. If a reviewer says a particular rule is unacceptable, you can delete or rewrite that rule and retrain around the constraint. You cannot hand-edit a bar chart of contributions. ## Interpretability is not a model-class checkbox The common mistake is to answer this question with a list — "linear and trees are interpretable, ensembles and networks are not." That is only true at small sizes. Readability degrades along axes that have nothing to do with the algorithm: - **Number of terms.** A linear model with 400 non-zero weights is not something a human reads and reasons about; sparsity is what makes the weight vector legible, which is why practitioners deliberately drive weights to zero or cap the model at a dozen terms. - **Depth and path count.** A depth-3 tree has 8 leaves; a depth-20 tree has up to a million, and no reader holds that. - **Feature comprehensibility.** A model over `feature_317` from an unsupervised transform is not interpretable even if the equation is short, because the terms mean nothing to a domain expert. Interpretable by design is a joint property of the model *and* its feature space. - **Correlated inputs.** Two nearly duplicated columns split the credit for one effect between two weights, so each weight reads smaller than the effect really is. ## What to say in an interview Define intrinsic versus post-hoc in one sentence, give one concrete intrinsic example with its fitted form, and then immediately volunteer the nuance about size and feature quality — that last part is what separates a memorised definition from someone who has actually tried to get a model past a reviewer. If asked which you would choose, say it depends on whether anyone has to defend, contest or hand-edit the logic, not on which family sounds more modern.

  • Is a linear model always interpretable?
    No. A linear model is interpretable when it is sparse and its features are named quantities a domain expert recognises. With hundreds of non-zero weights, heavy interaction terms, or opaque transformed inputs, nobody can read it, and correlated columns split one real effect across several weights so each reads smaller than it is.
  • If an intrinsic explanation is faithful by construction, why does anyone use post-hoc methods?
    Because the model in production is often not one you chose for readability — an ensemble that a team already ships, or a model whose accuracy is worth keeping. Post-hoc methods are the only option for describing something you are not allowed to replace. They buy a description of an existing model, not a guarantee about it.
  • What does it mean for an explanation to be faithful?
    Faithful means the explanation describes what the model actually computes, not what would be plausible or desirable. For an intrinsic model faithfulness is automatic. For a post-hoc description it has to be measured, usually by checking that the description reproduces the model's outputs on data you did not use to build the description.

An intrinsic explanation is the recipe the cook actually followed; a post-hoc one is a food critic's reconstruction of the recipe from tasting the dish.

saying these in an interview costs you the question

  • Says linear models are interpretable and trees are not
  • Treats interpretability as a yes/no property of the algorithm
  • Assumes any post-hoc explanation is faithful to the model
  • Calls a 400-term linear model readable because it is linear
  • Confuses interpretability with the model being accurate

context

open as a page

When may a shallow tree fitted to a black-box model's predictions be quoted as its explanation?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A global surrogate is trained on the black box's own outputs, so it may be quoted only when it reproduces that model closely on held-out inputs from the deployment population, and closely within each segment you discuss.

open as a page

A readable pneumonia-risk rule list learned that asthma lowers risk — how do you respond?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The rule is a true pattern in the data and a lethal policy. Asthmatic pneumonia patients were routed straight to intensive care, so they survived more often. The label reflects the treatment they received, not their underlying risk.

open as a page

Your interpretable scorecard scores 2 AUC points below a boosted model — how do you decide which to ship?

level: principalimportance: should knowfreq 45%

basics

~20 s

Decide from what the decision requires, not the metric gap: ship the readable scorecard when the logic must be signed off, contested or hand-edited, after confirming the 0.02 AUC difference is real out of time.

open as a page

In a generalized additive model, how does a shape function keep a non-linear effect readable?

level: middleimportance: nice to knowfreq 26%

basics

~10 s

A generalized additive model fits one curve per feature and sums them, so a feature's whole effect is a single readable plot: non-linear in shape, never tangled with the other features.

open as a page