A training run pins its initialization seed but still varies run to run - what else is random?
answer
- count the random draws, not the layers
- training randomness has more than one source
- shuffle order is randomness too
- augmentation and dropout draw every step
- child processes carry their own generators
basics
~20 sInitialization is only one of three random streams. The epoch shuffle that sets data order and the per-step stochastic operations - augmentation parameters, dropout masks, random masking - draw too, and parallel loading workers hold their own generators.
solid answer
~40 sA training run consumes randomness from at least three places, and pinning one of them fixes only that one. First, parameter initialization. Second, data order: the epoch permutation, plus any random negative sampling, bucketing or class-balanced sampler. Third, everything stochastic per step - augmentation parameters, dropout masks, stochastic depth, random token masking. On top of that, parallel data-loading worker processes each carry their own generator that the parent process's seed never touched, and resuming from a checkpoint restarts the streams unless the generator states were saved with it. The robust pattern is one master seed from which you *derive* an explicit seed per purpose - init, sampler, augmentation, per-worker - rather than one global seed everything shares, because a shared stream couples components: add one extra draw anywhere and every value downstream shifts.
go deeper
Be ready to name the obvious sources of randomness in a training run - the initial weights, the order examples arrive in, and anything that randomly transforms or drops values during a step.
Explain the mechanics: a generator is deterministic given its seed and its consumption order, which is why one shared stream couples components and why child worker processes need their own derived seeds.
Show the operational habit - one master seed, derived per-purpose seeds, generator state saved in checkpoints, resolved seeds logged at startup, and a two-run diff at fixed steps as the standard verification.
Own the distinction between reproducibility and replicability. Argue for what the team standardizes on: seeds are infrastructure for varying the run deliberately, not a way to keep publishing one lucky result.
## What a seed actually fixes A pseudorandom generator is a deterministic machine. It holds an internal state, and each draw returns a value and advances that state. Seeding sets the starting state, so the *same seed produces the same sequence of values* - but only if the draws are requested in the same order, in the same quantity, from the same generator. Every reproducibility failure below is a violation of one of those three conditions. ## The three streams in a training run **1. Initialization.** Weight matrices, biases, embedding tables and any randomly initialized buffers are sampled once at construction. This is the stream people remember to seed, because it is the visible one. **2. Data order.** Each epoch draws a permutation of the dataset. Beyond the plain shuffle, this stream also covers random negative sampling, random cropping of long documents into training windows, length bucketing with randomized bucket order, class-balanced or importance samplers, and random subsampling of an oversized corpus. Data order alone can move a final metric measurably: the model sees the same examples but in a different sequence, so it takes a different path through the loss surface and lands somewhere else. **3. Per-step stochastic operations.** Augmentation parameters (crop box, flip, colour jitter, the mixing coefficient of a mixup-style blend), dropout masks, stochastic depth decisions, random token masking for a masked-prediction objective, and any sampling inside the model itself. These draw thousands of times per epoch, so they are the stream most likely to be under-seeded and the one that makes two runs diverge from step one. ## The three things people forget **Worker processes.** Data loading is usually parallelized across worker processes. A child process gets its own generator; seeding the parent does not set it. The classic symptom is augmentations that repeat within a worker or differ between runs even though the script "seeded everything". **Replicas.** In a multi-replica run every replica needs a *different* augmentation seed - otherwise all replicas apply identical augmentations - while the sampler must partition the data disjointly. Same-seed-everywhere and different-seed-everywhere are both wrong; derived seeds are right. **Resume.** Restarting from a checkpoint that stored weights and optimizer state but not generator states gives a resumed run a fresh random sequence. The run is then not reproducible across its own restart boundary, which is exactly when you least want a mystery. ## Why one global seed is fragile even when it works If initialization, the sampler and augmentation all draw from a single shared stream, they are coupled by consumption order. Add a layer, remove a dropout call, reorder two lines, or take a conditional branch that draws once more - and every value after that point shifts. The seed is unchanged, the code is "equivalent", and the run is different. Deriving a seed per purpose (a stable hash of the master seed and a component name, for instance) decouples them, so changing one component leaves the others' sequences intact. ## How to verify Run the script twice with everything pinned and compare. If the loss at step 1 already differs, the initialization or the first batch differs. If step 1 matches but step 100 does not, a per-step stream (augmentation, dropout) is unpinned - or the divergence is not seed-related at all but comes from nondeterministic parallel reductions, which no seed can fix. Log the resolved seed of every stream at startup, and log a hash of the first batch's example indices: that single line turns "why is this different?" into a two-second check. ## Reproducibility is not the real goal Bit-identical reruns are a debugging convenience. What matters for decisions is *replicability*: does the conclusion survive when the seed changes? A pipeline that is perfectly reproducible under one seed and swings widely across seeds has not been made trustworthy, only made repeatable. Seed control exists so that you can deliberately vary the seed and measure the spread - not so that you can report one lucky run forever.
- How would you check that every stream is actually pinned?Run the script twice and compare the loss at steps 1, 10 and 100, plus a hash of the first batch's example indices. A mismatch at step 1 implicates initialization or data order; a match at step 1 followed by divergence implicates a per-step stream such as augmentation or dropout. Log the resolved seed of every stream at startup so the check takes seconds.
- Why can adding one extra random draw break reproducibility even with an unchanged seed?Because components sharing one global stream are coupled by consumption order. The generator returns a fixed sequence; inserting or removing a draw shifts every later component onto different values, so the same seed yields a different shuffle and different masks. Deriving an independent seed per purpose from one master seed removes that coupling.
- In a multi-replica run, should every replica use the same seed?No. Replicas need different augmentation and dropout seeds, or they all apply identical transforms and you lose the diversity augmentation was meant to add. The data sampler, by contrast, must partition examples disjointly rather than draw independently. Derive both from one master seed plus the replica index so the whole run stays describable by a single logged number.
Fixing the initialization seed is like agreeing on the order of a deck of cards, then letting three other people cut it, deal it and swap a few cards at random.
saying these in an interview costs you the question
- Claims one global seed guarantees a reproducible run
- Forgets that shuffle order is a randomness source
- Assumes child loading processes inherit the parent's generator state
- Thinks turning off augmentation removes all per-step randomness
- Resumes from a checkpoint without restoring generator state