Why must a linear model be handed an explicit product column to capture a discount-by-loyalty-tier interaction?
answer
- the form is additive by construction
- one effect per feature, summed
- the product must be a column
- categorical partner means several products
- pairwise count grows quadratically
basics
~20 sA linear model is additive: every feature contributes its own amount, and nothing lets one feature change another's contribution. If discount depth works differently for each loyalty tier, that column has to be built by hand.
solid answer
~50 sThe model `demand = b0 + b1*discount_depth + b2*tier_indicator` is additive by construction — the two terms are simply summed, so the model has no way to express that deep discounts move gold-tier customers differently from bronze-tier ones. It applies one discount response to everybody. Adding a product column, discount depth multiplied by the tier indicator, gives the fit a new basis function whose value depends on both inputs at once, and the fitted surface can then respond to discount differently within each tier. Because tier is categorical, you need one product column per tier indicator, not a single one. It is still a linear model — the products are computed before fitting — but the interactions have to be chosen deliberately, and each one needs enough observations in its combination to be estimated at all.
go deeper
Recall that a plain linear model adds each feature's contribution up, so it cannot let one feature change another's effect unless you build the product column and give it to the model.
Explain the additive form explicitly, what a product column adds to the design matrix, and why a categorical partner with several levels needs one product column per indicator.
Demonstrate judgment about which interactions to build: driven by domain mechanism, checked for observations in every combination, and weighed against moving to a model class that finds them itself.
Own the standard for how interactions enter production models — domain-justified, hierarchy respected, data support verified, documented — and the decision about when hand-built terms stop being worth the effort.
## Additivity is an assumption, not a technicality Write the model out: `demand = b0 + b1*discount_depth + b2*is_gold + error` Whatever numbers the fit chooses, the structure says something strong about the world: the contribution of discount depth is added on, and the contribution of tier is added on, and neither term knows anything about the other. The model can say gold customers buy more overall, and it can say deeper discounts sell more units, but it cannot say that deep discounts barely move gold customers who would have bought anyway while transforming bronze-tier behaviour. That statement requires a term whose value depends on both inputs simultaneously. That is exactly what an **interaction** is, and in a linear model it is created as a **product column**: multiply the two feature columns together, row by row, and add the result to the design matrix. Because the multiplication happens before fitting, the model remains linear in its parameters and the ordinary least-squares machinery is unchanged — this is basis expansion again, with the basis function being a product of two features rather than a power of one. ## Categorical partners multiply the columns If tier has four levels, it is already represented by three indicator columns (one level absorbed into the intercept). Interacting a continuous discount depth with that tier means three product columns, one per indicator. The result is a model that can trace a different discount response inside each tier while sharing everything else. Interacting two continuous features gives a single product column and a fitted surface that twists rather than staying flat in the way a plane does. ## The cost side: columns and cells Two costs decide how far you take this. **Column count grows quadratically.** With 30 features there are 30 x 29 / 2 = 435 distinct pairwise products, before considering three-way terms or products with the polynomial expansions of the same features. Adding them all is rarely defensible: most encode no real mechanism, and each is a parameter estimated from the same finite data. **Every product column needs data in its cells.** A discount-by-tier interaction is estimated from customers who actually appear at various discount depths *within* each tier. If platinum customers were never given a deep discount, the corresponding product column is being fitted from a handful of rows, and the estimate will be unstable no matter how sensible the term looks on paper. Before adding an interaction, cross-tabulate and check the cells are populated. There is also a modelling convention worth knowing: the **hierarchy principle**, which says that if a product term is in the model you keep both of its parent terms too. Dropping a main effect forces the fitted surface through a constrained shape that the data rarely supports, and makes the fit depend on arbitrary details of how the inputs are scaled. ## Which interactions to build The interactions worth spending parameters on come from the domain, not from a search over all pairs. Promotions are the classic source: discount depth by customer tier, discount depth by channel, promotion by day-of-week. Ask the business owner where they believe one factor changes another's effect, build those, and leave the rest. ## Why other model families do not need this This is a limitation of the additive form specifically. A decision tree represents an interaction without being told: it splits on loyalty tier near the top, and then splits on discount depth separately inside each branch, so the discount response can differ by branch by construction. An ensemble of trees does this routinely across many feature pairs, which is a large part of why boosted-tree models often beat a hand-built linear model on tabular data with many interactions. The counter-argument for staying with the linear model and its hand-built terms is that the resulting model is small, stable, auditable and easy to reason about term by term — which matters when the model must be explained to a regulator, a pricing committee or an auditor. That is a deliberate trade: you pay in feature-engineering effort and in the interactions you failed to think of, and you are paid in transparency and stability. ## The one-line answer Additive means additive. A linear model can only express what its columns encode, so an effect that depends on two features at once exists only if you multiply those features into a column yourself.
- How does a decision tree represent an interaction without a product column?Through nested splits. It can split on loyalty tier near the root and then split on discount depth separately within each branch, so the discount response is allowed to differ per branch by construction. An ensemble does this across many feature pairs automatically, which is why boosted-tree models often outperform a hand-built additive model on tabular data with rich interactions.
- What limits how many interaction columns you can add?Two things. The number of pairwise products grows quadratically — 435 for 30 features — so exhaustive expansion spends parameters on terms with no mechanism behind them. And each product column must be supported by observations in its combinations: if platinum customers never saw a deep discount, that term is fitted from almost nothing and its estimate is unstable.
- If you include the product of two features, why keep both features on their own?This is the hierarchy principle. Dropping a parent term forces the fitted surface through a restricted shape that the data usually does not support, and makes the fit sensitive to arbitrary choices about how the inputs are scaled. Keep the parents in and let the fit decide how much of the effect belongs to the product.
saying these in an interview costs you the question
- Assumes a linear model discovers interactions on its own
- Adds every pairwise product without checking cell counts
- Keeps only the product term and drops its parents
- Confuses an interaction with correlation between two features
- Uses a single product column for a multi-level categorical partner