skip to content

When should a train/validation/test split be stratified rather than purely random?

level: middleimportance: should knowfreq 48%

answer

  1. unbiased on average, still noisy
  2. small partition, uncommon class
  3. binomial spread around the expected count
  4. match the proportions by construction
  5. same eleven rows, not more information

basics

~20 s

Stratify whenever a partition is small enough that random assignment could land a class mix quite different from the full table. Sampling within each class keeps the proportions matched, so partitions are comparable and no class nearly disappears.

solid answer

~50 s

Random assignment is right on average but has real variance, and the variance is what bites when a partition is small or a class is uncommon. Take a 900-row table whose target is a 3-level seniority band with the top band at 6% of rows: a random 20% test set gets about 11 top-band rows in expectation, but the spread is about 3, so a given seed can hand you 5 or 17. Any per-band metric computed on that is noise, and the training partition's class mix wobbles too, so a model comparison partly measures split luck. Stratified assignment fixes the proportions by drawing the requested share from within each class, which removes that variance and guarantees no partition is missing a level. It does not manufacture information: eleven rows are still eleven rows, just reliably eleven.

go deeper

for a junior

Be ready to say what stratified assignment means - each partition keeps the same class proportions as the whole table - and that it matters most when a class is uncommon or a partition is small.

for a middle

Quantify it: give the expected count of the rare class in the partition and the spread around it, and explain why that spread corrupts both the estimate and the comparison between models.

for a senior

Show the judgment about limits - that stratification fixes counts rather than precision, that it says nothing about whether the sample matches production, and what you do when even a stratified partition is too thin to report on.

for a principal

Own the convention: which variables the organisation stratifies on by default, how the choice is recorded so results are reproducible across teams, and when the added complexity of multi-variable stratification stops paying for itself.

## What the two schemes actually do **Random assignment** shuffles the rows and cuts them into the requested proportions. Each row's partition is independent of its label, so in expectation every partition has the same class mix as the full table. **Stratified assignment** groups the rows by a chosen variable - usually the target - and cuts *each group* into the requested proportions, then concatenates. Each partition now has the class mix of the full table by construction, up to rounding. Both are unbiased in the sense of not systematically favouring a class. The difference is variance. ## Where the variance bites Take a 900-row table whose target is a 3-level seniority band, with the top band at 6% of rows - 54 rows in total. Split 60/20/20, so the test partition is 180 rows. Under random assignment the number of top-band rows in the test set behaves like a binomial count: mean `180 * 0.06` = 10.8, standard deviation `sqrt(180 * 0.06 * 0.94)` = about 3.2. Two standard deviations either side spans roughly 4 to 17 rows. So depending on the seed: - The per-band recall for the top band is computed from somewhere between four and seventeen observations, and it will swing wildly between reruns for reasons that have nothing to do with the model. - The *training* partition's top-band count moves in the opposite direction, so the model itself changes with the seed. - Two candidate models compared across two different seeds are partly being compared on split luck. Stratified assignment removes all of that. Each partition gets 6% top band by construction, the reruns become comparable, and the pathological case - a partition with zero rows of a level, which makes some metrics undefined and can make a classifier that never predicts that level look fine - cannot happen. ## What stratification does not do This is where interviews separate people. - **It does not create information.** A stratified test set with 11 top-band rows still estimates that band's recall from 11 observations. The estimate is stable in its *sample size*, not in its precision. - **It does not make the estimate match the deployment population.** If the table itself over-represents a band relative to where the model will run, stratifying reproduces that over-representation faithfully in every partition. How the rows were gathered in the first place is a separate question. - **It does not balance the classes.** Stratifying preserves the existing 6/x/y proportions; it is not resampling and it is not reweighting. - **It does not protect against rows that belong together.** Matching proportions says nothing about whether related rows ended up on both sides of the split. ## When plain random assignment is fine - Every partition is large and every class is common - a 40,000-row test set with no class under a few percent will land within a hair of the table's proportions on its own. - The target is continuous and well spread, with no long tail you specifically report on. - You are going to average over many splits anyway, so the per-split variance averages out. The cost of stratifying when you did not need to is essentially zero, which is why it is a sensible default for classification and why the interesting answer is about *why*, not about *whether*. ## Stratifying on something other than the target You can stratify on any variable whose mix must match across partitions - a region code, a product line, a data-collection wave - when that variable strongly shifts the outcome and is unevenly distributed. The limit is combinatorial: stratifying on the target and two covariates at once means the cross-product of their levels, and once a cell holds a handful of rows the scheme cannot allocate them in the requested proportions and quietly degrades. Pick the one or two variables that matter and let the rest fall randomly. ## Reporting Whatever you choose, record it. "Stratified on the target band, seed fixed, 60/20/20" is a one-line note that makes every downstream number reproducible, and its absence is the reason two colleagues get different scores from the same code.

  • Does stratifying a split change the expected performance estimate, or just its variance?
    Mainly the variance. A random split already gives the right class mix in expectation, so stratifying does not shift the average estimate; it removes the run-to-run swing that makes small partitions unreliable and reruns incomparable. It cannot make a thinly populated class informative - it only makes its row count predictable.
  • Even after stratifying, the rarest band's test score jumps between runs. What is going on?
    The count is now fixed but still tiny, so the metric is estimated from a handful of observations and moves with which particular rows landed there and with any randomness in training. Report an interval, state the row count behind every per-class figure, and get the class-level signal by resampling over the non-test rows instead.
  • Can you stratify on a variable other than the target?
    Yes, on any covariate whose mix must match across partitions - a region or product-line code that strongly shifts the outcome, for instance. The constraint is combinatorial: each extra stratification variable multiplies the number of cells, and once cells contain only a few rows the proportions can no longer be honoured.

saying these in an interview costs you the question

  • Says random splitting systematically under-samples rare classes
  • Believes stratification makes a tiny class informative
  • Thinks stratification balances the classes
  • Claims a stratified split matches the deployment population
  • Never checks the class counts that landed in each partition

context