skip to content

Your RL variant beats a baseline that lacked observation normalisation — is the gain real?

level: principalimportance: nice to knowfreq 28%

answer

  1. compare like with like
  2. normalisation is not an algorithmic change
  3. equal tuning budget for both arms
  4. ablate each detail with the idea fixed
  5. effect must exceed within-arm seed spread

basics

~20 s

Not as stated. Observation normalisation, reward scaling and advantage standardisation often move returns more than an algorithmic change does, so the comparison must give both arms the same details and the same tuning budget before any gain can be attributed to the algorithm.

solid answer

~50 s

The honest answer is that the experiment cannot separate the two causes yet. Normalising observations by a running mean and standard deviation is routinely the difference between a run that learns and one that never leaves the floor, with no change to the algorithm at all — and the same is true of reward scaling and per-batch advantage standardisation. So the first move is a matched comparison: turn the detail on for both arms, give both the same tuning budget over the same search space, and run both over the same set of seeds. Then run an ablation that toggles each detail with the algorithmic change held fixed, so you can see how much of the gap each one owns. If the algorithmic delta is smaller than the swing from the details, say that in the write-up. As a lead I would make matched-details, matched-budget comparison the default, because otherwise the team's reported gains are unreproducible by anyone who happens to have the details already on.

go deeper

for a junior

Be ready to say that scaling inputs by a running mean and standard deviation can decide whether an RL run learns at all, and that a fair comparison must apply the same preprocessing to both sides.

for a middle

Explain the mechanics: why unscaled observation dimensions make one learning rate unworkable, what dividing advantages by a batch standard deviation does to effective step size, and how reward scale reaches the critic's targets.

for a senior

Show how you would run the attribution — a factorial ablation over the same seeds, matched tuning budgets, and effect sizes compared against within-arm spread — and how you handle normalisation statistics at checkpoint and evaluation time.

for a principal

Own the standard: require matched details and matched tuning budgets before any internal claim of an algorithmic win, and be able to argue the cost of the alternative, which is teams building on gains that were really configuration.

## The situation Someone reports that a new algorithmic idea beats a baseline. The baseline ran without observation normalisation. The question is what fraction of the reported gap belongs to the idea. ## Why the details are not details A family of small, usually undocumented choices sits between an algorithm's equations and a working implementation, and empirically they carry a large share of measured performance. **Observation normalisation.** Inputs are scaled by a running mean and standard deviation maintained over the observations seen so far. Without it, an observation vector whose components differ by orders of magnitude — joint angles near 1, velocities in the tens — makes the first layer's response dominated by the large components, and a single learning rate cannot suit all of them. On continuous-control benchmarks this alone flips a run between learning and flat. **Reward and return scaling.** Value targets inherit the reward scale. A critic trained on targets in the thousands needs a different effective step size from one trained on targets near one, so scaling rewards by a running estimate of return magnitude changes how the critic and policy losses trade off. **Per-batch advantage standardisation.** Subtracting the batch mean from advantages is, in expectation, a baseline shift that leaves the policy gradient's expectation essentially unchanged; dividing by the batch standard deviation is not neutral, because it rescales the gradient magnitude and therefore acts as an adaptive step size. On clipped policy-gradient methods it also changes how often the clipping threshold binds. Others in the same family: gradient-norm clipping, learning-rate annealing, the initialisation scale of the final policy layer, and how the value loss is weighted against the policy loss. ## Two failure modes in the comparison **Unmatched details.** If the new arm has normalisation and the baseline does not, the experiment measures the union of the two changes. The literature on reproducing policy-gradient results has repeatedly found that re-adding these details to a baseline erases much of a published algorithmic gap. **Unmatched tuning budget.** If the variant was tuned over hundreds of configurations and the baseline runs the original paper's defaults on a new environment, part of the gap is tuning, not method. This one is easy to miss because it looks like diligence. Both are made worse by the seed spread. If a single run per arm decides the comparison, and within-arm seed spread is large, the reported difference may be neither algorithm nor detail but chance. ## How to actually attribute the gain 1. **Fix the protocol first.** Same environment, same interaction budget, same evaluation procedure, same set of seeds, decided before results are seen. 2. **Give both arms every detail.** Turn normalisation, reward scaling and advantage standardisation on for both, or off for both. The comparison is only meaningful with the detail axis held constant. 3. **Match the tuning budget.** Same search space, same number of trials per arm, and re-tune the baseline on this environment rather than importing defaults. 4. **Ablate one axis at a time.** With the algorithmic change fixed, toggle each detail and measure the swing. This produces the sentence that makes the result credible: how much of the gap each cause owns. 5. **Size the effect against the noise.** If the algorithmic delta is smaller than the seed spread inside one arm, you have not demonstrated it, regardless of how clean the ablation looks. ## Operational consequences beyond the paper Normalisation statistics are state. The running mean and standard deviation are part of the model: they must be checkpointed with the weights, frozen at evaluation and deployment, and applied identically to serving inputs. Evaluating with statistics that keep updating, or with statistics reset, silently changes the policy's inputs and is a common source of an offline score that will not reproduce online. Reward scaling has the same property when a running return estimate feeds it. ## The lead's call As the person who sets the standard, the useful policy is not 'ban implementation details' — the details are often genuinely necessary. It is to require that comparisons hold them constant, that tuning budgets are matched and stated, and that the write-up separates 'this idea helps' from 'this configuration is strong'. The failure this prevents is the expensive one: a team building on a claimed algorithmic contribution that was really a normalisation switch, and finding a year later that their stack already had it on.

  • How would you design the ablation that separates the algorithm from the details?
    A small factorial: the algorithmic change on or off, crossed with each detail on or off, every cell run over the same ten seeds. Report the aggregate per cell with intervals. The interesting number is the swing along the detail axis compared with the swing along the algorithm axis; if the former dominates, the write-up must say so.
  • What operational risk does observation normalisation introduce once the policy is deployed?
    The running mean and standard deviation are model state. If they keep updating during evaluation, are reset, or are not shipped with the weights, the deployed policy sees differently scaled inputs than it was trained on and can degrade silently. Checkpoint them with the parameters and freeze them outside training.

saying these in an interview costs you the question

  • Credits the algorithm without ablating the normalisation
  • Tunes the new arm hard and the baseline not at all
  • Calls observation normalisation a trivial detail
  • Runs one seed per arm in the ablation
  • Ships weights without the frozen normalisation statistics
  • Claims a delta smaller than the within-arm seed spread

context