In a generalized additive model, how does a shape function keep a non-linear effect readable?
answer
- one function per feature
- the terms are summed, not nested
- link function on the left-hand side
- curve replaces a single slope
- no interactions unless added explicitly
basics
~10 sA generalized additive model fits one curve per feature and sums them, so a feature's whole effect is a single readable plot: non-linear in shape, never tangled with the other features.
solid answer
~50 sA generalized additive model has the form `g(E[y]) = b0 + f1(x1) + f2(x2) + ... + fp(xp)`, where each `fj` is a flexible one-dimensional function fitted from the data and `g` is a link, such as the log-odds link for a binary outcome. That per-feature function is the shape function, and because the terms are added rather than nested, moving one feature changes the score by exactly `fj(new) - fj(old)` whatever the other features are, which is what makes the curve readable alone. Curves are centred to average zero, so the intercept carries the baseline. The payoff is shapes a straight line cannot hold: in a hospital 30-day readmission model, the age curve can show a distinct bump in the early sixties that one linear coefficient averages away. The cost is no interactions unless you add explicit pairwise terms, each of which is a surface rather than a curve.
go deeper
Recall the shape: one fitted curve per feature, added together, so each feature's effect is a plot rather than a single number. Knowing that much and that it is more flexible than a straight line is enough at this level.
Write the additive form with its link and explain why summing the terms is what makes each curve readable alone. Be ready to describe the smoothing parameter as a bias-variance dial and to name the missing-interactions limitation.
Show judgment about reading curves in production: uncertainty bands, thin tails, spikes that turn out to be imputation artefacts, and the decision of which pairwise interaction terms are worth their readability cost.
Argue when an additive model is the right house style for a regulated or safety-critical domain, what review process the curves feed into, and how you stop the model drifting into an unreadable pile of interaction terms over time.
## The form A linear model constrains every feature's contribution to a straight line: `score = b0 + b1*x1 + ... + bp*xp`. A generalized additive model (GAM) relaxes the *shape* while keeping the *additivity*: ``` g(E[y]) = b0 + f1(x1) + f2(x2) + ... + fp(xp) ``` Each `fj` is an arbitrary smooth one-dimensional function of a single feature, estimated from data. `g` is a link function chosen for the outcome type — the identity link for a continuous target, the log-odds link for a binary target, so that the additive score is on the log-odds scale and the predicted probability is its logistic transform. The `fj` are the **shape functions**. Fitting them is done with smoothers — splines with a roughness penalty, or many small one-feature learners summed — with a smoothing parameter controlling how wiggly each curve may get. That parameter is a bias-variance dial: too smooth and the curve degenerates towards a straight line and misses real structure; too wiggly and the curve chases noise, which shows up as implausible zig-zags a domain expert will reject on sight. ## Why additivity is what buys readability The curve is readable *because* the terms are summed. Under additivity, the effect of changing a feature from value `a` to value `b` is `fj(b) - fj(a)`, and that difference does not depend on any other feature's value. So a single two-dimensional plot — feature on the horizontal axis, contribution to the score on the vertical — is a complete and exact statement of what the model does with that feature. No caveats, no "holding everything else at its mean", no averaging over a distribution. Because the intercept and the curves are only identified up to a constant, shape functions are conventionally **centred** to average zero over the training data. Then `b0` is the baseline score for an average case and each curve reads as "how much this feature's value pushes the score above or below that baseline". ## What it buys over a single coefficient Consider a hospital model for 30-day readmission with age as an input. A linear term reports one number — say, risk rises steadily with age. The shape function can instead reveal risk climbing through middle age, a distinct bump in the early sixties, and a flattening or even a dip afterwards. That non-monotone structure is exactly what a single slope averages away. Two consequences follow: 1. **Better fit without losing the audit.** You get the non-linearity that a black box would have found, in a form a clinician can look at and either endorse or challenge. 2. **Detection of data problems.** Sharp steps or spikes in a shape function often mark an artefact rather than biology — an imputed placeholder value, a coding change, an age recorded as 0 for missing. A flexible curve makes those visible as an isolated spike; a linear coefficient absorbs them silently. ## The limits, and how to state them - **No interactions by default.** A pure GAM assumes the features act additively. If risk depends on age *only for patients with a particular comorbidity*, no set of independent curves can express that. The standard remedy is to add named pairwise terms, `f(xj, xk)`, which is honest but costs readability: a surface has to be read as a heat map rather than a curve, and adding many of them turns the model back into something nobody reads. - **Correlated features still split credit.** If two features carry the same information, the fit can distribute the effect between their two curves in ways that are jointly correct but individually misleading. Check pairwise correlation before reading any single curve as "the effect". - **Shape is not cause.** A curve tells you how the model's score moves with the feature in the data it was fitted on. It is not evidence that intervening on the feature moves the outcome. - **Uncertainty matters at the edges.** Regions with few training rows have poorly determined curves; a dramatic swing in the tail of an age curve at ages 95+ may rest on a handful of rows. Plot the curve with an uncertainty band and refuse to read the sparse ends. ## Interview framing Write the additive form, name the link, say the words "one function per feature, summed", and then make the readability argument explicitly: additivity is what lets a single curve be a complete statement of the feature's effect. Finish with the interaction limitation, unprompted — candidates who state the cost of their preferred model are the ones interviewers believe.
- How do you extend a generalized additive model when two features genuinely interact?Add an explicit pairwise term `f(xj, xk)` for that named pair, chosen from domain knowledge or from evidence that the additive fit is failing in that region. It stays auditable because the interaction is a listed term rather than an emergent property, but it must be read as a two-dimensional surface, so add few of them and only where they earn their keep.
- What does the smoothing parameter on a shape function control?How wiggly the curve is allowed to be. Heavy smoothing pushes the curve towards a straight line, raising bias and hiding real structure. Light smoothing lets it chase noise, raising variance and producing zig-zags that a domain expert will dismiss. It is usually chosen by cross-validation, then sanity-checked against whether the resulting shape is plausible.
- A shape function swings sharply at the extreme high end of a feature's range — do you trust it?Usually not. The tails typically contain few training rows, so the curve there is poorly determined. Plot the uncertainty band, count the rows in that region, and if it is thin either refuse to read the tail or constrain the fit to be flat beyond the last well-supported value.
A linear coefficient is a single average grade for a student; a shape function is the full transcript, showing where they climbed and where they dipped.
saying these in an interview costs you the question
- Says a GAM is just a linear model with more features
- Claims shape functions capture interactions automatically
- Reads a shape function as a causal effect of the feature
- Trusts a wiggly curve in a data-sparse tail
- Forgets the link function on a binary outcome