skip to content

How do you set a decision tree's minimum leaf size on a 900-row cohort with only 36 positive cases?

level: seniorimportance: should knowfreq 36%

answer

  1. count events, not rows
  2. 36 positives is the real budget
  3. leaf probabilities need a denominator
  4. expected positives equals prevalence times leaf size
  5. required events divided by prevalence

basics

~20 s

Size leaves by expected positive cases, not total rows: at a 4% rate, 50-row leaves hold about two positives, so their probabilities are noise. Divide the events you need behind a prediction by the prevalence.

solid answer

~50 s

The binding resource here is 36 positives, not 900 rows. With a minimum leaf size of 1 the tree isolates individual patients, reaches 100% training accuracy, and produces leaf probabilities of exactly 0 or 1 that carry no uncertainty and will not survive a fold change. With a floor of 50 you get at most 18 leaves, but each holds only about two expected positives, so even those probability estimates swing wildly. My approach is to set the floor from the minority class backwards: decide the minimum number of events I want behind any predicted probability — say ten — which at a 4% rate means leaves of roughly 250 rows, so three or four leaves at most. Then I validate with repeated stratified cross-validation and check whether the tree structure holds across folds. If it does not, this cohort supports a small tree or a simpler model, not a better-tuned one.

go deeper

for a junior

Remember that a leaf holding very few rows gives an unreliable probability, and that with a rare outcome the number of positive cases matters more than the total number of rows.

for a middle

Be able to do the arithmetic out loud: 4% of 900 is 36 events, a 50-row floor allows at most 18 leaves with about 2 expected positives each, and required events divided by prevalence gives the leaf size you actually need.

for a senior

Demonstrate the whole workflow: sizing from the minority class backwards, stratified repeated validation, checking structural stability across folds, and keeping any pruning step inside the resampling so it does not leak.

for a principal

Own the call that the cohort may not support a useful tree at all, and be ready to argue for collecting more events, choosing a lower-variance model, or shipping a two-split tree with honest uncertainty rather than a score-tuned one.

## Count events, not rows A 900-row cohort with a 4% outcome rate holds 36 positive cases. Every leaf-size decision on this dataset is really a decision about how those 36 events are divided, because the majority rows are plentiful and carry little information about where the positive region lies. This is the reframing an interviewer wants to hear. A candidate who talks about 900 rows will pick a leaf floor that sounds reasonable — 20, or 50 — and never notice that the resulting leaves contain one or two events. ## What each extreme actually does **Floor of 1.** The tree may split until every leaf is pure, isolating individual patients. Training accuracy is 100%, which is a statement about lookup, not prediction. Each leaf reports a probability of exactly 0 or 1, and those numbers have no usable uncertainty attached: a leaf built around one positive patient says "100% risk" for anyone who lands there. Change the random seed of a cross-validation split and the tree's shape changes wholesale, because a single event moving between folds re-routes an entire branch. **Floor of 50.** Now at most 900/50 = 18 leaves exist. Better — but the expected positives per leaf is 0.04 * 50 = 2. A leaf with 2 out of 50 reports 4% and a leaf with 5 out of 50 reports 10%, and the difference between those two leaves is well within what resampling would produce by chance. You have controlled tree size without making any individual leaf's estimate trustworthy. ## Sizing from the minority class backwards The usable rule is to fix the number of events you need behind a prediction and divide by the prevalence: ``` leaf_size ~= required_events / prevalence ``` Want at least 10 positives supporting each probability estimate at a 4% rate? That is 10 / 0.04 = 250 rows per leaf, so about three leaves across the whole cohort — a tree with one or two splits. This feels shockingly small, and it is the correct conclusion. Thirty-six events do not support a fifteen-leaf decision structure, whatever the training score says. Clinical prediction work has long used a similar order-of-magnitude discipline: the number of events, not the number of records, bounds how many effects you can estimate. ## The other knobs, and their limits - **Class weights.** Upweighting the minority class changes the impurity arithmetic so the tree stops ignoring the rare outcome. Combined with a minimum *weighted* fraction of samples per leaf, the floor then applies to weighted mass rather than raw rows. This is worth doing — but note carefully what it does not do: it creates no new events. The information ceiling set by 36 positives is unchanged. - **Maximum depth.** A cap of 2 or 3 is a reasonable belt-and-braces limit here, but it is blunter than the leaf floor, because a depth-3 branch carrying few rows is still allowed to make a claim about them. - **Cost-complexity pruning.** Perfectly usable, and often preferable to guessing a floor: grow a modest tree, prune it, and let cross-validated performance pick the size. On data this small, wrap the whole procedure — growth and pruning together — inside the resampling, or the pruning decision leaks the validation data. ## Validating the choice Use repeated stratified cross-validation so each fold keeps roughly 4% positives; with 36 events, unstratified folds can easily produce a fold with two positives in it. Two things are worth inspecting beyond the headline score. First, the spread of leaf probability estimates across repeats: if the same input region is assigned 5% in one repeat and 40% in another, the leaves are too small regardless of the mean score. Second, the tree structure itself: if the root split changes feature across folds, you are reading noise, and no amount of leaf-size tuning fixes it. ## The senior conclusion Sometimes the answer to "how do I set the leaf size" is "this dataset does not support a tree with interesting structure". A stump or two-split tree that a clinician can read, with honest wide uncertainty on its leaf rates, is a better deliverable than a fifteen-leaf tree tuned to a cross-validation score computed on 36 events. Saying that out loud — and proposing either collecting more events or moving to a lower-variance model — is the answer that distinguishes someone who has shipped a model on rare outcomes from someone who has only tuned hyperparameters.

  • How does weighting the minority class change this leaf-size decision?
    Weights change the impurity arithmetic so the rare outcome is no longer ignored, and pairing them with a minimum weighted fraction per leaf makes the floor apply to weighted mass rather than raw row counts. What it cannot do is manufacture events: with 36 positives, the ceiling on how much structure the data supports is unchanged, so weighting improves where the tree looks, not how much it can learn.
  • How would you detect that the leaves are too small rather than just reading the score?
    Run repeated stratified cross-validation and look past the mean metric. Check the spread of each region's predicted probability across repeats, and check whether the tree's root split even keeps the same feature. Wildly varying leaf rates or a shifting root split mean you are fitting noise, and no leaf-size value will rescue a structure that is not reproducible.
  • Would you use a single tree on this cohort at all?
    Often not. With 36 events the tree that survives honest validation is a stump or a two-split tree, so a simpler, higher-bias model can match it with far less variance, and averaging many trees is another route to stability. If the deliverable requires readable rules, ship the small tree with explicit uncertainty on its leaf rates rather than a larger tree that looks more informative than the data allows.

saying these in an interview costs you the question

  • Sizes leaves from total rows and ignores prevalence
  • Treats 100% training accuracy on tiny leaves as success
  • Reads a leaf probability of 1.0 from one patient as certainty
  • Thinks class weights create more usable minority information
  • Tunes leaf size on unstratified folds with 36 events

context