Training loss oscillates with an identical pattern in every epoch — what data-loading bug does that suggest?
answer
- the same wiggle in every epoch
- periodic, not random
- what order does the file sit in?
- print one batch's label histogram
- reshuffle indices before each epoch
basics
~20 sThe batches are being read in the stored file order without shuffling. If that order is sorted by class, every batch holds one class, the loss swings as the class changes, and the same swings repeat each epoch.
solid answer
~50 sA loss curve whose wiggles line up epoch after epoch says the sample order is fixed and correlated with the label. The classic cause is reading a catalogue in the shard order it was written — say a 40-class product-photo set stored class by class — with shuffling off, so each batch contains a single class. The model drives the batch loss down by predicting whatever class it is currently fed, then pays for it the moment the class changes; those spikes recur at the same iteration index every epoch because the order never changes. Each minibatch gradient is then a strongly biased estimate rather than a noisy estimate of the full-data gradient. Fix it by reshuffling the sample indices before every epoch, then confirm that the label histogram of one batch resembles the overall class distribution.
go deeper
Know that training batches must be shuffled and that shuffling is repeated each epoch. Be ready to say what a single-class batch does to a gradient step and to name the one-minute check: print the labels in one batch.
Explain why unbiasedness of the minibatch gradient depends on random sampling, and why momentum amplifies the damage when consecutive batches share a label. Be able to read epoch-length periodicity in a loss curve as evidence about ordering.
Show the diagnosis path on a real pipeline: overlay two epochs' loss curves, count distinct labels per batch, and check whether a streaming shuffle buffer is actually larger than the class run length in the files.
Own the guardrail rather than the incident. Argue for a startup assertion on batch label diversity, for how datasets get written so that block structure cannot silently reappear, and for what such a check costs when the run is legitimately single-class.
## The symptom You plot training loss against iteration and see structure, not noise: a repeating pattern of dips and spikes whose shape is the *same* in epoch 2 as in epoch 1, aligned to the same iteration offsets. Validation is poor and the model's predictions look skewed toward a handful of classes. Nothing has errored; the run is green. That exact periodicity is the tell. Ordinary minibatch noise is aperiodic — it does not reproduce itself. A pattern with period equal to one epoch can only come from something that repeats with period one epoch, and the only thing that does is **the order in which samples are visited**. ## Why order matters at all Stochastic gradient descent rests on one assumption: the gradient computed on a minibatch is an *unbiased* estimate of the gradient on the full training set, differing from it only by noise. That holds when the batch is a random sample of the data. It does not hold when the batch is a contiguous slice of an ordered file. Concretely, take a 40-class product-photo catalogue written class by class — all of class 0, then all of class 1, and so on — and read it in stored order with shuffling disabled. With a batch size well below the size of one class block, nearly every batch contains a single label. The cheapest way to reduce the loss on such a batch is to collapse the output distribution onto that one class. The next few batches reward the same collapse. When the block boundary arrives, the model is confidently wrong on every sample at once and the loss spikes; it then chases the new class. Momentum makes this worse, because accumulated velocity keeps pushing toward the previous class for several steps after the class has changed. Any statistic computed per batch is likewise estimated on a single class rather than on a representative mix. Because the visiting order is identical on the next pass, the entire pattern replays: same dips, same spikes, same iterations. Late in training there is also a recency effect — whichever class was seen last has the strongest claim on the final weights. ## Confirming it in a minute 1. Pull one batch straight out of the loader and print the histogram of its labels. If the batch holds one or two distinct labels while the dataset has forty, you are done. 2. Count distinct labels per batch across the first twenty batches. A properly shuffled loader gives roughly the number you would expect from sampling the class distribution. 3. Overlay the loss-versus-iteration curves of two consecutive epochs. If the spikes superimpose, the order is fixed. ## The fix Reshuffle the sample indices **before every epoch**, not once at load time. A single shuffle removes the class correlation but freezes the batch partition, so the same groups of samples ride together forever and the gradient noise stops being independent across passes; reshuffling is cheap and strictly better. When the data is streamed from shards too large to hold in memory, a permutation of the whole index set is not available. The working recipe is two-level: randomize the order of shards, interleave reads from several shards at once, and draw training samples through a shuffle buffer. The buffer only helps if it is **large relative to the run length of a single class** — a buffer of 1,000 samples cannot break up blocks of 10,000 identical labels. Sizing the buffer without checking the block structure of the files is the most common way this bug survives a supposed fix. ## Boundaries and exceptions Shuffling applies to *training* order. Evaluation order is irrelevant when the metric is an average over the whole set, so there is nothing to fix there. Sequence data is the case people cite as an exception, and it is usually a misunderstanding: you still shuffle the *sequences*, you simply do not shuffle timesteps within a sequence. Order is meaningful along the time axis, not along the sample axis. Finally, distinguish this from a genuinely noisy loss. A learning rate that is too large produces large, erratic swings that do **not** align across epochs, and it usually degrades from the start rather than showing clean structure. If someone reaches for the learning rate on seeing an epoch-periodic pattern, they are treating a data-ordering bug with a model-side knob, and the pattern will still be there afterwards.
- Is shuffling once when the dataset is loaded enough, or must it be redone every epoch?Once is not enough. A single shuffle breaks the class correlation but freezes the batch partition, so the same samples travel together on every pass and the gradient noise repeats instead of being independent across epochs. Reshuffling the index order before each epoch costs almost nothing and gives fresh batch compositions, which is why it is the default expectation.
- You stream from shards too large to fit in memory. How do you shuffle then?Two levels: randomize the order of shards, interleave reads from several shards at once, and draw samples through a shuffle buffer. The buffer must be large compared with the run length of a single class — a buffer of a thousand samples cannot break up blocks of ten thousand identical labels, so check the block structure of the files before trusting the buffer size.
- Which is more suspicious — a loss that oscillates with epoch-length periodicity, or one that is noisy with no periodicity?The periodic one. Aperiodic noise is expected from minibatch sampling, and if it is severe it usually points to a learning rate that is too high or a batch that is too small. Periodicity locked to the epoch length can only come from a repeating visit order, which makes it a data-pipeline problem rather than an optimization one.
It is like revising from a stack of flashcards sorted by subject that nobody ever cuts: you look brilliant inside each block and go blank at every boundary, and the same blank spots come back on the next pass.
saying these in an interview costs you the question
- Blames the learning rate for a strictly epoch-periodic loss pattern
- Thinks one shuffle at load time covers every epoch
- Says shuffling matters for the validation set too
- Never inspects the label histogram of an actual batch
- Assumes any oscillating loss means the batch size is too small