skip to content

Why do five runs of one deep RL configuration with different seeds end at wildly different returns?

level: middleimportance: must knowfreq 58%

answer

  1. the agent collects its own data
  2. early exploration decides what is learnable
  3. bootstrapped targets compound early error
  4. outcomes cluster into took-off and never-took-off

basics

~20 s

A deep RL agent generates its own training data, so random exploration and initialisation change which transitions it ever sees. Bootstrapped value estimates then amplify that early divergence, and the run ends in a completely different policy.

solid answer

~50 s

In supervised learning the dataset is fixed, so a seed only perturbs initialisation, shuffling and augmentation, and runs converge to similar losses. In RL the policy chooses the data: a different seed means different initial weights, different sampled actions, different environment start states and different replay ordering, so two runs see disjoint experience within the first few thousand steps. That difference compounds, because value learning bootstraps off its own estimates and the policy is trained on the states it already prefers. A seed that stumbles onto rewarding behaviour early gets gradient signal and keeps improving; one that collapses its action entropy first can sit at zero forever. The result is usually bimodal rather than a tight spread — for example five seeds of an unchanged locomotion configuration returning 0, 900, 2,600, 2,800 and 3,100 — which is why any single run tells you almost nothing.

go deeper

for a junior

Be ready to say that in RL the agent's own actions decide what data it trains on, so a different random seed leads to different experience and a different final policy. Name the seeded pieces: weights, action sampling, environment start states.

for a middle

An interviewer expects the mechanism. Explain the closed loop between behaviour and data, why bootstrapped value targets propagate early error instead of averaging it away, and why an entropy collapse before the first reward is effectively terminal.

for a senior

Show that you plan for it. Log seeds, treat the seed as an experimental factor, and describe the bimodal outcome shape you actually see in production runs and what you report instead of a mean.

for a principal

Own the tradeoff. Decide whether a high-variance method is acceptable at all given your compute budget and how often downstream teams can afford to rerun, and be able to argue that reliability across seeds is a first-class selection criterion, not a footnote.

## The claim Run the *same* deep RL code, the *same* hyperparameters and the *same* environment five times, changing only the random seed, and you can get final returns of 0, 900, 2,600, 2,800 and 3,100. Nothing is broken. This spread is a property of the method, and understanding where it comes from is the difference between a candidate who trusts a single run and one who does not. ## Why supervised intuition does not transfer In supervised training the data is a fixed set of pairs. A seed changes the initial weights, the order of minibatches and any stochastic augmentation, and the optimiser walks a slightly different path over the *same* loss surface. Runs usually land at similar losses, and the seed-to-seed gap is small next to the gap between two genuinely different models. Reinforcement learning breaks the fixed-data assumption. The agent's own policy decides which states it visits and therefore which transitions it ever trains on. Data collection and learning form a closed loop, and a seed perturbs the loop at the very first step. ## The four sources of randomness - **Initial parameters.** The policy and value networks start from different weights, so the first action distribution differs. - **Action sampling.** Almost every deep RL method acts stochastically — sampling from a policy distribution, or acting greedily with an epsilon-random override. Which actions get tried first is pure chance. - **The environment.** Start states, and any stochastic dynamics inside the simulator, are drawn from their own stream. - **Data ordering.** Replay-buffer sampling order, and the shuffling of a batch of collected trajectories, differ per seed. ## Why small differences do not average out Three mechanisms turn a tiny initial perturbation into a different outcome. **The data distribution is non-stationary and self-selected.** After a few thousand steps the two runs hold different experience. Anything a run never tries, it never learns about — a state that is unvisited contributes no gradient. Two runs are therefore not noisy estimates of one optimisation problem; they are solving two different, diverging problems. **Bootstrapping compounds error.** Value-based methods regress the value of a state onto a target built from the network's own prediction of the next state. An early over-estimate becomes the target for the states that lead into it, so errors propagate backwards instead of washing out. Combine that with function approximation and off-policy data and the estimate can chase itself for a long time before it settles — or not settle at all. **Entropy is a one-way door in practice.** A policy that narrows its action distribution before finding reward has stopped exploring; it collects only the experience its current behaviour produces, which confirms the current behaviour. This is the usual explanation for the run that returns 0: it did not fail slowly, it stopped searching early. ## The distribution is bimodal, not Gaussian The most useful thing to say in an interview is about the *shape* of the outcome. Returns across seeds frequently cluster into 'took off' and 'never took off'. That has two consequences. A mean is dragged between the two modes and describes no run that actually happened, and a standard deviation implies a symmetric spread that is not there. Report how many seeds solved the task, alongside the distribution of the ones that did. ## What this does *not* mean It does not mean the failed seed is a bug. It also does not mean more training steps fixes it — a collapsed policy can sit flat for the whole budget. And it does not mean the variance is uniform across methods: some algorithms and some environments are far more reliable than others, and reliability across seeds is a legitimate axis on which to compare methods, not a nuisance to be hidden. ## What to do about it Fix what you can and measure the rest. Seed everything explicitly and log the seed with the run so a result can be reproduced. Then treat the seed as an experimental factor rather than a setting: run many of them, and judge a configuration by the distribution of its outcomes rather than by any single trace. Reducing the variance itself is a design question — more exploration pressure, gentler value targets, better input conditioning — but the first professional move is to stop drawing conclusions from one run.

  • If the environment were fully deterministic, would the seed spread go away?
    No. Deterministic dynamics remove one source, but initial weights, sampled actions and data ordering still differ, and those alone change which trajectories the agent ever collects. Determinism narrows the spread in easy environments; in anything with an exploration problem the feedback loop still dominates, and you will still see runs that never leave the floor.
  • Would training the failed seed for ten times longer close the gap?
    Usually not. A run that has collapsed its action entropy is no longer generating novel experience, so extra steps mostly re-confirm the current behaviour. Longer budgets help when a run is improving slowly; they rarely rescue a run that is flat at zero. Treat a flat seed as evidence about the method's reliability, not as a run to extend until it agrees with the others.
  • Is high seed variance ever a legitimate reason to reject an algorithm outright?
    Yes. If a method reaches a high score on three seeds out of ten and zero on the rest, its expected value under a fixed compute budget can be worse than a duller method that works every time. Reliability is part of the result, especially when downstream users cannot afford to rerun until they get lucky.

Supervised learning is five students taking the same exam paper; RL is five students each choosing their own reading list on day one, then being examined on what they chose to read.

saying these in an interview costs you the question

  • Claims a seed only changes weight initialisation
  • Says a deterministic environment guarantees identical runs
  • Calls a zero-return seed a bug and discards it
  • Assumes averaging two or three runs is enough
  • Says more training steps always closes the gap

context