skip to content

When do you need group k-fold or leave-one-subject-out instead of ordinary k-fold?

level: middleimportance: must knowfreq 62%

answer

  1. rows here are not independent
  2. ask who each row belongs to
  3. whole entity into one fold
  4. one subject per fold at the limit

basics

~20 s

Grouped folds are for data with repeated rows per entity — many windows from one wearer, many rows from one store. Every row of a group goes into one fold, so the score measures generalisation to unseen entities.

solid answer

~50 s

Ordinary k-fold assumes rows are interchangeable. They are not when the data has repeated measurements: a gait dataset of 30 subjects contributing 200 windows each has 6,000 rows but only 30 independent units. Split those windows at random and nearly every validation window has siblings from the same person in training, so the model can score well by recognising that person's stride rather than the activity, and the estimate is optimistic about a new wearer. Grouped folds assign whole subjects to folds — group 5-fold holds out six subjects at a time, leave-one-subject-out holds out one and runs 30 fits. Choose the group key by asking what the model will see fresh in production: a new person, a new store, a new patient. Then decide k by cost and by how many groups you have.

go deeper

for a junior

Be ready to recognise the pattern: many rows per person, store or patient means the rows are not independent. Know that grouped folds keep an entity's rows together and say why that matters.

for a middle

Explain the construction — groups, not rows, are partitioned — and the two variants, group k-fold and leave-one-group-out. Expect a question on how you choose the group key and how many folds to run.

for a senior

Demonstrate judgment: deriving the group key from the deployment question, handling wildly unequal group sizes when aggregating, and reading the per-group score distribution rather than only its mean.

for a principal

Own the standard: which entity keys are grouped by default across the team's datasets, what the extra fits cost in compute, and how grouped estimates are used in go/no-go decisions when a model must serve entities that do not yet exist.

## Rows versus independent units Cross-validation estimates how a model behaves on data it has not trained on. Plain k-fold delivers that estimate under one assumption: the rows are exchangeable, so a randomly chosen held-out row is a fair stand-in for a future row. Repeated measurement breaks the assumption. A wearable gait study with 30 subjects, each contributing 200 sliding windows, has 6,000 rows and 30 independent units. The windows within one subject share that person's height, cadence, sensor placement and walking habits. Under a random split, roughly 199 of any validation window's 200 siblings sit in the training set. A flexible model can exploit that: instead of learning what a stumble looks like in general, it learns what this particular wearer's signal looks like and reads the label off subject identity. The cross-validated number then answers the question 'how well do we predict new windows from people we already recorded?', which is rarely the question the project is asking. ## What grouped folds do Group k-fold takes a group key — subject id, store id, patient id — and partitions the *groups* into k parts rather than the rows. Every row of a group lands in exactly one fold, so a group is either wholly in training or wholly in validation for a given fold. The resulting estimate answers 'how well do we predict for an entity never seen in training?'. Leave-one-group-out is the extreme where k equals the number of groups: with 30 subjects it runs 30 fits, each holding out one person. Leave-one-store-out across 40 retail branches is the same construction, so a branch never appears on both sides of a fold. ## Choosing the group key The key is not a property of the data; it is a property of the deployment question. If the model will be rolled out to newly opened branches, the branch is the group. If it will always run on the same 40 branches whose history is already in training, the branch is arguably not a group at all. When several candidate keys exist — a transaction has both a customer and a store — pick the one you must generalise to; if both must hold, group on the coarser key, or on connected components when customers span stores, and accept that the folds get chunkier. ## Choosing the number of folds With 30 subjects the options run from group 5-fold (five fits, six subjects held out each time) to leave-one-subject-out (30 fits, one subject each). The tradeoff is not the same as choosing k in ordinary k-fold, because here the fold count is capped by the number of groups. Leave-one-group-out uses the largest possible training set each time and produces a per-subject score distribution, which is genuinely useful: it shows whether the model fails badly on a handful of people rather than uniformly. Its downsides are cost — one fit per group — and that a single fold's score, computed on one subject's 200 windows, is very noisy on its own, so the individual numbers should not be over-read. Because the training sets of different folds overlap almost completely, the fold scores are correlated and a naive standard error across folds understates the true uncertainty. Grouped 5-fold is cheaper and each fold score rests on six subjects, so the individual numbers are steadier. With very few groups — say eight stores — leave-one-group-out is often the only way to get a usable number of folds at all. ## Aggregating when groups differ in size Groups are rarely the same size: one branch may contribute 20,000 rows and another 800. Grouped folds therefore contain different row counts, and the plain mean of the per-fold metrics silently weights a small fold as heavily as a large one. Two defensible choices: weight each fold's metric by its row count, or pool all out-of-fold predictions into one vector and compute a single metric over it. Report the spread across groups alongside the headline number either way — with grouped folds that spread is the interesting part, because it describes how the model behaves on the worst entity rather than on the average one. ## What grouping costs Grouping is not free. Holding out whole entities means the training set loses all of their rows, folds become less balanced in size, and stratifying on the label at the same time becomes approximate rather than exact, because you can only move whole groups and each group brings its own label mix. The score also usually drops relative to a random split. That drop is information about what the model was relying on, not a defect in the split.

  • With 30 subjects, how do you choose between leave-one-subject-out and grouped 5-fold?
    Cost first: 30 fits versus five. Leave-one-subject-out trains on the most data and hands you a per-subject score distribution, which exposes the people the model fails on; each single fold score is noisy because it rests on one person, and the fold scores are correlated so their spread understates uncertainty. Grouped 5-fold is cheaper and steadier per fold. With very few groups, leave-one-group-out is often the only way to get enough folds.
  • Group sizes differ by an order of magnitude — how do you aggregate the fold scores?
    Do not take a plain mean of per-fold metrics; unequal fold sizes make that an unequal weighting of rows. Either weight each fold's metric by its row count, or pool the out-of-fold predictions and compute one metric over the whole pooled vector. Then report the per-group spread too, since with grouped folds the worst-entity behaviour is usually the number people actually care about.
  • A row could belong to two different groups — how do you pick the key?
    Pick the entity the model must generalise to at deployment; that is the whole point of grouping. If both keys genuinely matter and they interleave, group on the coarser one, or on connected components so that any two rows linked through a shared entity stay together. The folds become chunkier and fewer, which is the price of the stronger guarantee.

Testing a speech recogniser on the same speakers it trained on tells you it recognises those voices, not that it recognises speech.

saying these in an interview costs you the question

  • Splits 6,000 windows from 30 subjects at random across folds
  • Picks a group key unrelated to what deployment sees fresh
  • Averages fold scores as though the folds were equal size
  • Thinks shuffling more thoroughly solves repeated-subject rows
  • Reads a single leave-one-subject-out fold score as a reliable number

context