Why shuffle the rows before splitting them into k-fold cross-validation folds?
answer
- folds are contiguous slices
- files arrive with a sort order
- the sort key becomes the split
- one class per fold in the worst case
- except when order means time
basics
~20 sExported files usually arrive sorted — by label, by date, by customer. Cutting such a file into contiguous blocks gives folds that are not representative, sometimes single-class, so every fold score misleads. Shuffling row order first breaks that ordering.
solid answer
~50 sFold assignment is usually just slicing the row order into k contiguous blocks, so whatever order the file arrived in becomes the fold structure. A CSV sorted by label is the classic case: with two classes and five folds, the first folds are all class 0 and the last are all class 1, so each model trains on data missing a class and is tested on the class it never saw. The score is garbage, and it is garbage in a way that looks like a modelling problem rather than a plumbing problem. Shuffling the index order before slicing fixes it, and you fix the shuffle's seed so the run reproduces. The one place you must **not** shuffle is when row order carries meaning you are trying to respect — chronological data, where shuffling lets a model train on the future and predict the past.
code
python · 14 linesimport random
labels = [0] * 10 + [1] * 10 # the file arrived sorted by label
order = list(range(20))
def folds(order, k=5):
n = len(order)
return [order[i * n // k:(i + 1) * n // k] for i in range(k)]
print([[labels[i] for i in f] for f in folds(order)])
random.seed(0)
random.shuffle(order)
print([[labels[i] for i in f] for f in folds(order)])go deeper
Be ready to say that folds are contiguous slices of the row order, so a file sorted by label produces folds containing a single class. Shuffling before assigning folds is the fix.
Explain the mechanism and its limits: a shuffle randomises fold membership on average but does not guarantee class proportions, and it needs a fixed seed to stay reproducible.
Show the diagnostic instinct. When a score looks absurd, inspect fold composition before touching features, and know that shuffling is exactly wrong for chronological rows or repeated records of the same entity.
Own the guardrail. Decide whether fold composition checks and seed recording belong in the team's evaluation harness, so a badly ordered export cannot quietly produce a number someone presents to a stakeholder.
## Where the problem comes from Splitting into folds is mechanically trivial: take the rows in whatever order they sit, cut the sequence into k contiguous chunks, and call each chunk a fold. That is fast and deterministic, and it silently makes the file's existing order into the experiment's design. Real files are rarely in random order. They come out of a database with an ORDER BY, they are concatenated from per-class or per-month exports, they are sorted by an identifier that correlates with something, or they were assembled by appending the positive examples to the negatives. None of that is visible when you look at the first ten rows. ## The label-sorted case Take a table of 1,000 rows, 500 of class 0 followed by 500 of class 1, and cut it into five contiguous folds of 200. Folds 1 and 2 are pure class 0, fold 3 is half and half, folds 4 and 5 are pure class 1. Now run the cross-validation. When fold 1 is held out, the model trains on folds 2-5, which contain 300 rows of class 0 and 500 of class 1, and is scored on 200 rows that are all class 0. Every fold is an out-of-distribution test with respect to its own training set. Accuracy collapses, or — worse — comes out at some plausible-looking middling value that you then spend a day trying to improve with better features. The model is fine; the split is broken. The same failure occurs in subtler forms. A table sorted by date means each fold is a different time window, so seasonality, drift or a schema change lands entirely inside one fold. A table sorted by customer ID means folds correspond to ID ranges, which often means regions or signup cohorts. A regression target sorted ascending means each fold covers a different slice of the target range, and the model is asked to extrapolate every time. ## What shuffling does and does not do Shuffling permutes the row order before the contiguous slicing, so fold membership becomes a random draw rather than a function of position. Each fold then looks, in expectation, like a random sample of the whole table: roughly the right class proportions, the whole date range, all customer segments. Three limits are worth knowing: - **Shuffling is random, so it only fixes things on average.** With a rare class — say 40 positives in 5,000 rows — a random shuffle can still hand one fold 3 positives and another 13. Getting exact class proportions in every fold requires a fold-assignment scheme that enforces them, which is a separate mechanism from shuffling. - **Shuffling makes the run non-reproducible unless you fix the seed.** Two people running the same code get two different numbers, and the difference can be larger than the model change they are arguing about. Fix the seed and record it. - **Shuffling destroys order that sometimes matters.** If rows are chronological and the deployed model will predict the future from the past, shuffling lets training rows sit after test rows in time. The resulting score is optimistic and the mechanism is invisible. Likewise, if several rows describe the same underlying entity, shuffling scatters them across folds and the model sees near-duplicates of its test rows — that needs a fold assignment that keeps related rows together, not a shuffle. ## Leave-one-out needs no shuffle Leave-one-out cross-validation holds out exactly one row at a time and uses every other row for training, for all `n` rows. Fold membership is therefore not a choice at all: there is exactly one possible partition, order is irrelevant, and the result is fully deterministic. Any scheme with k < n has a partition to choose, and therefore a shuffle to think about. ## What to say in the interview Name the mechanism, not just the rule. "Folds are contiguous slices of the row order, and files arrive sorted, so without a shuffle the file's sort key becomes the split" shows you understand why the advice exists. Then add the two caveats — fix the seed, and do not shuffle when the row order encodes time — because that is what separates someone repeating a tip from someone who has debugged this.
- When would shuffling before folding be the wrong thing to do?When row order encodes time and the deployed model predicts forward. Shuffling puts future rows in the training set and past rows in the test fold, so the score is optimistic in a way nothing in the output reveals. Chronological data needs a split that respects the ordering.
- Why fix the shuffle's random seed?Reproducibility and comparability. Without a fixed seed, two runs of identical code draw different partitions and report different numbers, and on a small table that difference can exceed the effect you are trying to measure. Fix and record the seed so a score change is attributable to the change you made.
- Does leave-one-out cross-validation need shuffling?No. With one row held out at a time there is only one possible partition of the data, so row order cannot influence the result and the whole run is deterministic. Shuffling is only a question when k is smaller than the number of rows.
saying these in an interview costs you the question
- Assumes exported data already arrives in random order
- Shuffles chronological rows and reports the optimistic score
- Treats a broken split as a modelling problem to fix with features
- Never fixes the seed, then debates non-reproducible numbers
- Thinks a shuffle guarantees equal class counts in every fold