How does a decision tree route a row whose split feature is missing, without imputing it?
answer
- the node needs a direction, not a value
- backup cut that mimics the primary
- fall back to the majority branch
- try missing rows both ways, keep the better
- the pattern of absence is being learned
basics
~20 sTwo tree-native mechanisms exist. CART learns surrogate splits, backup features whose cut best reproduces the primary split, and uses the best available one. Boosted trees instead learn a default direction per node by testing which side gives more gain.
solid answer
~50 sA node asks `feature <= threshold`, so a missing value leaves the question unanswerable — the tree needs a routing rule, not a filled-in number. CART's answer is **surrogate splits**: after choosing the primary cut, it ranks other features by how well a cut on them reproduces the primary left/right assignment on rows where the primary feature is observed, and stores the best few. If declared income is missing on an incoming loan row, a surrogate on employment length decides the direction; if every surrogate is also missing, the row takes the majority direction. Gradient-boosting algorithms such as XGBoost and LightGBM take the other approach: during split search they try sending the missing rows left and then right and keep whichever yields the higher gain, storing that as the node's **default direction**. The trade-off is that this learns the training missingness pattern, so if the reason values go missing changes in production, rows are routed by a rule fitted to a different population.
go deeper
Know that a node needs only a direction, not a value, and that trees have built-in ways to send a row down a branch when the split feature is absent.
Explain both mechanisms concretely: a surrogate cut chosen to agree with the primary split, and a default direction chosen because it gave the higher split gain.
Demonstrate the operational judgment — informative missingness as a feature you get for free, and the monitoring that catches a shifted missingness mechanism before it silently misroutes traffic.
Own the policy question of whether missingness should be modelled implicitly at all, given that it couples the model to pipeline behaviour that other teams change without notice.
## The problem a tree actually has A node asks `feature <= threshold` and needs a left-or-right answer. When the value is absent, there is no answer. Note what the tree does *not* need: it never needs a plausible number for the missing field, only a direction. That is why tree-native handling can be smarter than substituting a value — and it is a genuinely different question from choosing a fill strategy as a preprocessing step, which is a separate concern with its own trade-offs. ## Mechanism one: surrogate splits This is CART's answer. After the primary split at a node is chosen, the algorithm looks for **backup cuts that mimic it**. For each other feature it finds the threshold whose left/right assignment agrees most closely with the primary split's assignment, measured on the rows where both the primary feature and that candidate are observed. Those candidates are ranked by agreement and the best few are stored on the node alongside the primary rule. At scoring time the logic is a cascade: if the primary feature is present, use it. If not, walk the surrogate list and use the first surrogate whose feature is present. If none is available, fall back to the **majority rule** — send the row down whichever branch took most of the training rows at that node. What makes surrogates work, and what breaks them: - They require **redundancy**. A surrogate only helps if some other feature carries overlapping information. If declared income is missing and employment length correlates with it, the surrogate reconstructs the routing decently. If nothing correlates with income, the best surrogate barely beats the majority rule and you are effectively guessing. - They must be **judged against the majority rule**, not against perfection. A surrogate with 55% agreement, where the majority branch already takes 60% of rows, is worse than useless and should not be stored. - They **cost training time** — for every node you search thresholds on every other feature a second time — and they add per-node state to the model. - They can quietly become a **proxy for something else**. A surrogate is fitted to mimic the split on the rows where the primary was observed; the rows where it is missing may be a different population, and the surrogate has never been validated on them. ## Mechanism two: the learned default direction Gradient-boosting algorithms including XGBoost and LightGBM use sparsity-aware split finding instead. While evaluating a candidate split, the rows whose value for that feature is missing are excluded from the threshold scan, then tried as a block on the left child and as a block on the right child. Whichever assignment produces the higher split gain becomes the node's stored **default direction**, and at scoring time every missing row simply takes that branch. The consequences are worth spelling out: - The route is **learned from the target**, not assumed. If applicants with no declared income default more often, the default direction will send them toward the higher-risk region because that is what maximised gain — no indicator feature required. - **Missingness becomes a signal.** This is a real advantage when the absence is informative, which in operational data it very often is: a field is blank because the customer skipped it, because a device was offline, because an upstream service timed out. Each of those carries information about the row. - It is **per node**, so the same feature can default left high in the tree and right further down, matching whichever population reaches each node. - It is **cheap**: one extra comparison per candidate split, no second threshold search. ## The senior risk: the missingness mechanism shifts Both mechanisms fit the *pattern* of missingness present in training, and that pattern is a property of your pipeline, not of the world. Make the income field mandatory on the application form and the missing rows disappear — but the population that used to be missing is now routed by their declared value, and any calibration you inherited from the default direction is gone. Let an upstream enrichment service start failing and a new, different population suddenly floods the default branch, receiving whatever score was learned for a completely different group. Neither event produces an error. The model keeps scoring. So the production discipline is: record the **per-feature missing rate at training time**, monitor it against live traffic, and alert on drift. Track the share of scored rows that take a default route, and watch the score distribution of those rows separately from the rest. And test explicitly what happens to a feature that was **never missing during training** — the node had no evidence with which to learn a route, so its behaviour is a fallback rather than a fitted decision, and you want to discover that in a test rather than in production. ## How to answer the trade-off question Surrogates buy you a routing decision grounded in another feature's value, which degrades gracefully when missingness is random and uninformative, at the cost of training time and a dependency on redundancy. Default directions buy you a routing decision grounded in the target, which is strictly better when missingness is informative and stable, and worse when it is unstable — because it commits harder to a pattern that may not hold. Knowing that both exist, and that neither substitutes for monitoring the missingness rate, is the whole of the answer.
- When do surrogate splits fail to help?When no other feature carries overlapping information with the primary one. The best surrogate's agreement then barely exceeds the majority rule, so routing is effectively a weighted coin flip toward the larger branch. A surrogate is only worth storing when its agreement clearly beats that baseline.
- Why is a learned default direction risky when the missingness mechanism shifts?The direction was chosen because it maximised gain on the rows that were missing during training — a specific population with a specific outcome profile. If a form field becomes mandatory or an upstream feed starts failing, a different population takes that branch and inherits a score fitted for someone else. Nothing errors.
- What would you monitor in production for a model that relies on this?Per-feature missing rate against the training baseline, the share of rows taking a default or surrogate route at each level, and the score distribution of those rows compared with fully observed rows. A sudden move in any of the three is the signal that routing assumptions have broken.
- What happens if a feature was never missing in training but arrives missing in production?The node had no missing rows to evaluate, so no direction was learned from evidence and the behaviour is whatever the fallback happens to be. Treat it as untested: construct rows with that field absent, score them, and confirm the outcome is sane before the case appears live.
saying these in an interview costs you the question
- Says trees simply drop rows with missing values at scoring time
- Claims a surrogate split is just mean filling under another name
- Thinks the default direction is a fixed convention rather than learned
- Assumes missingness carries no information about the target
- Ignores that the missingness pattern can shift after deployment