Why can removing one training row change a decision tree's root split entirely?
answer
- greedy argmax over candidate cuts
- top candidates are nearly tied
- the root decides what children see
- correlated features make ties routine
- different structure, often similar predictions
basics
~20 sSplit choice is a greedy pick of the single best-scoring cut, and the top candidates are often nearly tied. A one-row change can flip which cut wins, and because the choice is at the root, every subtree below is rebuilt on a different partition.
solid answer
~50 sAt each node the algorithm scores every candidate cut and takes the argmax. That is a hard, discrete decision on top of a continuous score, so two candidates separated by a hair in score produce completely different trees. Delete one row from a 500-row table and the winning root feature can change; everything beneath it is then refit on a different partition, so the whole structure looks new even when held-out accuracy barely moves. Correlated or near-duplicate features make this far worse, because their scores are nearly identical by construction and noise decides. The practical consequence is about interpretation, not accuracy: you cannot present the root split as a discovered fact. Before quoting a rule to stakeholders, refit on several perturbed samples or cross-validation folds and report only the splits that recur, with thresholds stated as approximate ranges.
go deeper
Know that the tree picks the single best-scoring cut at each node, and that when two candidates score almost the same, a tiny data change can swap them.
Explain the greedy plus hierarchical mechanism, why correlated features produce near-ties, and why a fully grown tree is a low-bias, high-variance model.
Show how you check it: refit across folds or perturbed samples, compare recurring splits and held-out predictions, and decide whether the instability is cosmetic or real.
Own the communication policy — what the organisation is allowed to claim from a tree's structure, how rules are versioned across refits, and when interpretability requires a deliberately constrained model.
## Where the instability comes from A tree is grown greedily. At each node the algorithm enumerates candidate cuts — every feature crossed with every threshold between adjacent observed values — scores each one, and commits to the best. Two properties of that procedure combine badly: 1. **The decision is discrete.** The score is a continuous quantity, but the outcome is a winner-take-all choice. A candidate that wins by 0.0001 gets the node; the runner-up leaves no trace. 2. **The decision is hierarchical.** The root split determines which rows the two children ever see. Change the root and both subtrees are refit on different data, so a single flip at the top rewrites everything below it. Put those together and one deleted row out of five hundred is quite enough. That row shifts the score of the leading candidates by a tiny amount, the ranking flips, and the printed tree is unrecognisable. ## Why near-ties are common, not exotic - **Correlated features.** Two features that measure nearly the same thing — declared income and reported salary, height in centimetres and a rounded height band — produce nearly identical partitions and therefore nearly identical scores. Which one wins is decided by noise, and the loser then looks unimportant everywhere downstream, because the winner has already absorbed the signal. - **Threshold granularity.** Candidate thresholds sit between adjacent observed values. Removing a row removes candidate thresholds and shifts others, which perturbs every score in that feature. - **Shrinking samples down the tree.** Each level halves the rows a node sees, so deep nodes choose among candidates scored on a handful of rows. Structural instability is worst exactly where the tree is deepest. - **Small data overall.** With 500 rows, one row is 0.2% of the evidence, and score gaps of that magnitude are routine. ## Structure variance is not the same as prediction variance This distinction is what separates a fluent answer from a memorised one. Two trees can look entirely different and still carve almost the same regions of the input space — a split on centimetres at 178 and a split on the rounded band at "tall" partition the same people. So a wildly unstable structure can coexist with stable predictions. The converse also happens: a flipped root on genuinely different features can move predictions for a whole segment. So measure the thing you care about. If you care about the explanation, measure structural agreement: refit on repeated perturbed samples or on each cross-validation fold, and count how often each feature appears at the root, how often a given split recurs anywhere, and how tightly the chosen thresholds cluster. If you care about the predictions, measure prediction agreement on a held-out set across those same refits — the correlation or the disagreement rate tells you whether the instability is cosmetic. ## Why interviewers care Single trees are sold on interpretability, and interpretability is exactly what instability undermines. The moment a stakeholder reads "customers with tenure above 14 months churn less" off a tree, they treat it as a discovered rule, plan around it, and are entitled to be annoyed when next month's refit says 11 months on a different feature. The mature stance is to treat any single tree's structure as **one sample from a distribution of trees the data could have produced**, and to communicate only what survives resampling: the features that keep appearing, the rough direction of the effect, thresholds as ranges rather than exact numbers. The same caution applies to the seductive claim that the root split is "the most important feature." The root is the best single cut under this particular sample and this particular scoring rule, in a search that never reconsiders it. With two correlated candidates it is a coin flip which one is crowned, and the winner is not more causal, more important, or more actionable than the loser. ## Why this is a bias-variance story A fully grown tree has very low bias — it can carve any partition of the training data given enough depth — and pays for it with high variance: small changes in the sample produce large changes in the fitted model. That is the defining trade of the model family, and it is the reason the natural next question in an interview is what to do about it. Limiting how deep the tree may grow, or requiring a minimum improvement before a split is accepted, trades a little of that flexibility back for stability, and there are model-level answers beyond a single tree as well. ## What a good answer sounds like "The tree picks the argmax of a score over candidate cuts, and the leading candidates are usually nearly tied, especially with correlated features. Because the choice is greedy and hierarchical, a tiny perturbation at the root rewrites the whole tree. I check whether that matters by refitting across folds and comparing both the recurring splits and the held-out predictions, and I only quote rules to stakeholders that survive that check."
- Does an unstable structure necessarily mean unstable predictions?No. Two trees can split on interchangeable features and still carve nearly the same regions, so held-out predictions can agree closely while the printed rules look unrelated. Measure both: structural agreement across refits if you sell the explanation, prediction agreement on held-out rows if you sell the score.
- Why do correlated features make this worse?Two near-duplicates produce nearly identical partitions and therefore nearly identical split scores, so noise decides the winner. Worse, once one takes the node it absorbs the shared signal, and the other looks uninformative downstream — which then corrupts any story you tell about which feature matters.
- How would you report a tree's rules to a stakeholder given all this?Refit across resamples or folds first, then report only what recurs: the features that keep appearing, the direction of the effect, and thresholds as ranges rather than exact numbers. State explicitly that the root split is the best cut in this sample, not a ranking of importance or a causal claim.
It is like a knockout tournament decided by a single point in the first round: change one point and a different bracket plays out all the way down.
saying these in an interview costs you the question
- Treats the root split as proof of the most important feature
- Assumes a different-looking tree must give different predictions
- Says instability is a bug in the implementation
- Quotes an exact threshold to stakeholders as a discovered rule
- Thinks more depth makes the structure more reliable