skip to content

How should you report deep RL returns across seeds instead of showing the best run's curve?

level: seniorimportance: should knowfreq 46%

answer

  1. three seeds prove nothing
  2. a mean hides bimodal outcomes
  3. middle 50% of runs, not the best one
  4. stratified bootstrap intervals over ten seeds
  5. re-run the sweep winner on fresh seeds

basics

~20 s

Run about ten seeds per configuration and report an aggregate that is robust to failed runs, such as the median or interquartile mean, with stratified bootstrap confidence intervals and every per-seed curve visible. Never report the best run.

solid answer

~50 s

Decide the number of seeds before you look at any results, and make it around ten rather than three — with returns as spread as 0 to 3,100, three runs give an interval so wide it cannot separate anything. Aggregate with a statistic that survives failed seeds: the median, or the interquartile mean, which averages the middle 50% of runs and so is less noisy than the median while still ignoring the extremes. Attach uncertainty with stratified bootstrap confidence intervals rather than a standard-deviation band, because the outcome distribution is usually bimodal and not symmetric. Plot all seeds, or a shaded interval, never one curve. Report the fraction of seeds that reached the goal separately, since that is the number a reader most wants. And state the protocol — seed count, evaluation episodes, which checkpoint — because a learning curve with no shaded band is unreadable without it.

go deeper

for a junior

Be ready to say why a single learning curve is not a result, and that a deep RL number should come with the number of seeds behind it. Knowing that the best run is never the number to quote is enough at this level.

for a middle

Explain the mechanics of each aggregate: why a mean is dragged between two modes, what the interquartile mean discards, and what a bootstrap interval resamples. Be able to describe how you would build the interval from a set of per-seed final scores.

for a senior

Demonstrate protocol discipline. Fix the seed count before looking at results, hold evaluation seeds out from training, state which checkpoint you report, and re-run any configuration that a sweep selected for you before quoting its score.

for a principal

Own the standard for the team. Decide the minimum seed budget a result must carry before it can be cited internally, weigh that compute against the cost of acting on a false positive, and set the norm that reliability across seeds is reported alongside peak performance.

## What is being asked A deep RL result is a *distribution* over runs, not a number. The reporting question is how to compress that distribution honestly, and how many samples of it you need before compressing at all. ## How many seeds The common practice of three seeds comes from compute cost, not statistics. With outcomes that can range from 0 to 3,100 for one unchanged configuration, three samples give a confidence interval wide enough to contain almost any claim, and the probability that all three land in the same mode is high enough to produce a confidently wrong conclusion in either direction. Ten is the usual working minimum for a claim of the form 'A beats B'; more if the two methods are close or if the outcome is bimodal. Fix the number in advance. Choosing when to stop after seeing the numbers is a way of selecting on noise. ## Which aggregate - **Mean.** Sensitive to both modes. With four failures and six successes it reports a value no run achieved, and one extra failure moves it a lot. - **Median.** Robust, but throws away most of the sample and is itself noisy at small seed counts. - **Interquartile mean (IQM).** Discards the top and bottom quarters of runs and averages the rest. It keeps half the data, so it is less noisy than the median, while remaining insensitive to a couple of collapsed or lucky seeds. This is the aggregate recommended by the recent literature on statistical practice in deep RL, together with performance profiles that show the whole distribution. - **Maximum.** Not an aggregate. See below. ## Which uncertainty band A mean plus or minus one standard deviation encodes an assumption of a symmetric, roughly normal spread. Deep RL outcomes routinely violate it. Stratified bootstrap confidence intervals — resample runs with replacement, within each task, recompute the aggregate, and take the empirical interval — make no distributional assumption and behave sensibly with a handful of runs per task. On a learning curve, shade that interval over training steps rather than drawing a single line. ## Why 'best of five' is a claim about luck A paper that plots the best of five seeds is reporting the maximum of five draws from an unknown distribution. The maximum is a biased estimator of what you would get on a fresh run: it is optimistic by construction, and the bias grows with the number of runs you selected from. Worse, it is not comparable across methods unless both were selected from the same number of runs, which is usually not stated. The reader cannot recover the quantity they care about — what happens when *they* run it once. ## The same problem inside a hyperparameter sweep This is where the mistake usually hides. Suppose you sweep a grid, one run per cell, and take the best cell. If the seed-to-seed spread *inside* a single cell is larger than the gap between the best and second-best cells, you have mostly ranked noise, and the winning cell's score is an over-estimate for the same reason a best-of-five curve is. The winner will regress when you re-run it. The fix is to re-evaluate the top few cells with a fresh set of seeds, and to report the winner's score from those fresh runs rather than from the sweep that selected it. It also reframes the sweep result: if the best cell's advantage does not exceed the within-cell spread, the honest conclusion is that the hyperparameter did not matter over the range you searched. ## Evaluation protocol details that change the number State them, because readers cannot compare otherwise. How many evaluation episodes per seed, and were they run with the stochastic policy or a deterministic one. Which checkpoint — the final one, or the best checkpoint by evaluation score, which reintroduces exactly the selection bias described above. Whether evaluation episodes used environment seeds held out from training. And, for on-policy methods, whether the reported curve is training-time episode return, which is measured under exploration noise and is not the same quantity as a clean evaluation. ## Comparing two methods When the question is 'is A better than B', the useful statements are the interval around the difference of aggregates, and the probability that a random run of A beats a random run of B — a rank-based quantity that needs no normality assumption and is meaningful even when both distributions are bimodal. A p-value from a t-test on six runs assumes a shape the data does not have; prefer the bootstrap and the rank statistic, and say plainly when the seeds cannot separate the methods.

  • Why prefer the interquartile mean over the plain median?
    Both ignore a collapsed or lucky seed, but the median uses effectively one run's worth of information and is jumpy at ten seeds. The interquartile mean averages the middle half, so it keeps five times as much data and moves less between reruns while still being immune to the extremes. It is a robustness-versus-noise compromise, not a different claim.
  • Your sweep's best cell beats the baseline cell by less than the spread inside either cell. What do you conclude?
    That the sweep ranked noise. Take the top few cells, re-run each with a fresh set of seeds you did not select on, and report those numbers. If the gap still fits inside the within-cell spread, the honest conclusion is that the hyperparameter did not matter over the searched range — publish that rather than the selected maximum.
  • Is it acceptable to report the best checkpoint of each run by evaluation score?
    Only if you say so, and only with evaluation episodes held out from the ones used to pick the checkpoint. Selecting the peak of a noisy evaluation curve is a maximum over draws and is optimistic by the same argument as best-of-five seeds. Reporting the final checkpoint, or a fixed-step checkpoint, avoids the issue entirely.

saying these in an interview costs you the question

  • Reports the best of five runs as the result
  • Uses mean plus or minus standard deviation on bimodal returns
  • Runs three seeds and calls the gap significant
  • Picks the sweep's best cell without fresh seeds
  • Reuses training environment seeds for evaluation episodes
  • Plots one curve with no interval and no seed count

context