Under leaf-wise growth, why is a cap of 63 leaves not equivalent to a depth cap of 6?
answer
- the bound only runs one way
- depth caps leaves, leaves barely cap depth
- sixty-three leaves permits a long chain
- interaction order versus total capacity
- check which cap actually bound
basics
~10 sBoth permit roughly 64 leaves, but they constrain different things. A depth cap of 6 limits every prediction path to six splits. A cap of 63 leaves permits a chain 62 splits deep.
solid answer
~50 sA depth cap bounds the length of every root-to-leaf path, and therefore also bounds leaves at `2^depth` — depth 6 allows at most 64 leaves and at most six features interacting on any one path. A leaf-count cap bounds only the number of terminal regions; with 63 leaves the shallowest possible tree is depth 6, but the deepest is a chain of depth 62. Under leaf-wise growth the second case is not hypothetical, because best-first expansion concentrates splits wherever the measured gain is largest. So the two caps constrain orthogonal quantities: leaves control total capacity, depth controls interaction order and indirectly the row support of the deepest leaves. The usual practice is to set both, plus a floor on the rows required in a leaf, and to keep the leaf cap comfortably below `2^depth` so the leaf cap is the binding constraint rather than the depth cap silently truncating the search.
go deeper
Know that under best-first growth the number of leaves and the depth are separate settings, and that limiting one does not pin down the other. Recall that a tree with 63 leaves may be six deep or sixty deep.
Explain the one-way arithmetic: depth d allows at most 2^d leaves, while L leaves allow depth up to L - 1. Say which quantity each cap really governs — total capacity versus interaction order and row support.
Show the operating discipline: leaf count as the tuned capacity dial, depth as a guardrail, a minimum-support floor to protect tail branches, and a habit of inspecting the trees you actually got to see which constraint bound.
Own the standard. Decide whether your organisation caps depth for reviewability and stable retraining even at some accuracy cost, and make sure tuning searches cannot silently trade one cap against another in ways nobody can reproduce.
## The arithmetic that makes them look interchangeable A binary tree of depth `d` has at most `2^d` leaves, so a depth cap of 6 admits at most 64 leaves, and a cap of 63 leaves sounds like the same budget from the other end. It is the same budget only in the balanced case. The two caps are one-directional: depth bounds leaves (`leaves <= 2^depth`), but leaves bound depth only trivially (`depth <= leaves - 1`). Sixty-three leaves permits any shape between a balanced depth-6 tree and a 62-deep chain. ## Why the gap is real under leaf-wise growth Level-wise growth never explores the lopsided end of that range: it splits a whole frontier at a time, so a tree with 63 leaves is necessarily close to depth 6. Best-first growth has no such tendency — it puts the next split wherever the gain is largest, and gains are usually largest in whichever region already contains structure. Deep, one-sided branches are the *typical* output, not the pathological one. On an 800-row employee-attrition pilot extract this is easy to observe: with only a leaf cap in force, best-first growth runs to depth 20 along a single branch while level-wise growth on the same data stops at depth 6. ## What each cap actually buys **A leaf-count cap buys capacity control.** The number of terminal regions is the honest measure of how many distinct predictions one tree can make, so this is the dial that trades bias for variance most directly. Halve the leaf cap and you roughly halve the tree's expressiveness regardless of its shape. **A depth cap buys interaction-order control and support.** A path of `k` splits can involve at most `k` features, so depth is a ceiling on how many features may jointly define a region — which matters both for generalisation and for anyone who has to read the model. Depth also correlates with row support: each split divides the rows reaching it, so leaves at depth 20 on 800 rows sit on a handful of examples each, whatever the leaf cap says. ## The trap of setting only one Set only the leaf cap and you get the deep, low-support branches above. Set only the depth cap under leaf-wise growth and something subtler happens: because expansion is gain-ordered, the tree may stop far short of `2^depth` leaves — gain thresholds and minimum-support rules bite first — so the depth cap is not really controlling capacity, and two runs with the same cap can produce very differently sized trees. Worse, if the depth cap is the binding constraint, the leaf cap you tuned has no effect at all, and a later change to it looks inert. ## A workable configuration discipline 1. Choose the leaf-count cap as the primary capacity dial and tune it. 2. Set a depth cap as a *guardrail* that is loose relative to the leaf cap — loose enough that it rarely binds on large data, tight enough to forbid the 40-deep branch on small data. 3. Add a floor on the number of rows (or the summed hessian) required to form a leaf, which is what actually protects the tail branches; on a few hundred rows this floor does more work than either cap. 4. Verify which constraint bound by inspecting the trees you got — actual leaf counts and actual depths — rather than assuming your settings were the active ones. ## Small and noisy data The smaller the dataset, the more the caps matter and the more the depth guardrail earns its place, because the gain estimates driving best-first expansion are themselves noisy. A defensible default on a few hundred rows is a small leaf budget, a shallow depth guardrail and a meaningful minimum leaf support — or level-wise growth, whose balanced shape is a crude but effective form of the same restraint. ## What an interviewer is checking Whether you understand that these caps are not two spellings of one setting. The candidate who says "63 leaves is depth 6" has assumed balance, which is exactly the assumption leaf-wise growth abandons.
- If you set only a depth cap under leaf-wise growth, what goes wrong?Capacity stops being under your control. Best-first expansion stops when gain thresholds or support floors bite, so the tree often ends far short of the `2^depth` leaves the cap allows, and two runs on similar data can produce very different tree sizes. If the depth cap does bind, it silently becomes the active constraint and any leaf-count tuning you did has no effect.
- Which knob protects the deep, low-support branches most directly?A floor on the support required to form a leaf — a minimum number of rows, or a minimum summed hessian. Caps on leaves and depth constrain the tree's shape globally, but the concrete danger is a split whose two sides hold a handful of rows each; a support floor refuses exactly that split, wherever in the tree it appears, and it scales with the dataset in a way the shape caps do not.
- Does a leaf-count cap bound the number of features a single prediction can depend on?Only very loosely. A prediction depends on the features tested along its root-to-leaf path, so the relevant bound is depth, not leaf count. With 63 leaves a path could test up to 62 splits and therefore up to 62 distinct features. If you need to promise that no region is defined by more than a handful of features — for review or for explanation — you must cap depth.
saying these in an interview costs you the question
- Says 63 leaves and depth 6 are the same constraint
- Assumes leaf-wise trees are roughly balanced
- Thinks a leaf cap bounds interaction order
- Tunes the leaf cap while a depth cap is silently binding
- Relies on shape caps alone with no minimum leaf support