skip to content

Why does PPO run several gradient epochs over the same batch of trajectories?

level: middleimportance: should knowfreq 54%

answer

  1. environment steps cost more than gradients
  2. the ratio corrects for stale data
  3. importance weighting degrades as policies separate
  4. later epochs are increasingly clipped
  5. single-digit passes, commonly three to ten

basics

~20 s

The importance ratio corrects the objective for data gathered by the previous policy, so one batch stays usable for a few passes. Reuse amortises expensive environment interaction, and the clip keeps that correction trustworthy as the policy drifts.

solid answer

~50 s

Collecting the batch is usually the expensive part - simulator or hardware steps, not gradient steps - so one update per batch wastes most of that cost. Because the surrogate weights each sample by `r = pi_new(a|s) / pi_old(a|s)`, it estimates the *current* policy's expected advantage from data the *old* policy collected, which stays reasonable while the two policies remain close. The clip is what enforces closeness, so the mechanisms are a pair: the ratio licenses reuse, the clip decides how long the licence holds. In practice that means a single-digit number of passes, commonly three to ten, over the batch split into minibatches. Past that, more samples sit outside the clip band and contribute nothing, the advantages computed at collection time describe a policy that no longer exists, and the update is driven by whichever samples happen to remain inside the band.

go deeper

for a junior

Know that the batch is collected once and then passed over several times in minibatches, and that this exists because talking to the environment is the slow, expensive part of training.

for a middle

Explain why the reuse is legitimate rather than just conventional: the ratio reweights old data toward the current policy, and that estimate is only trustworthy while the policies are close. Name what degrades - clipped samples, stale advantages, accumulating drift.

for a senior

Bring the diagnostics. Say which quantity you would log to decide the epoch count for a new environment, and how epoch count, batch size and clip range trade against each other in a run you have actually tuned.

for a principal

Frame it as a budget decision: how much policy movement per unit of collected experience the project can afford, given what interaction costs and what a bad update costs. Be ready to argue when the answer is a different algorithm family rather than a different epoch count.

## The economics In most reinforcement learning settings the scarce resource is environment interaction, not compute. A robot arm, a hardware-in-the-loop rig, or a slow simulator produces transitions far more slowly than a network consumes them. If each collected batch yielded exactly one gradient step, nearly all of that expensive data would be thrown away after a single, very noisy update. So the practical question is: how many gradient passes can you legitimately squeeze out of one batch before the updates start lying to you? ## Why more than one pass is legitimate at all After the first gradient step the parameters have moved, so the batch was no longer collected by the policy you are now optimising. The importance ratio is the correction for exactly that mismatch. The surrogate `L = E_[s,a ~ pi_old] [ (pi_new(a|s) / pi_old(a|s)) * A(s,a) ]` reweights each sample by how much more or less likely the current policy is to take that action, which makes it an estimate of the current policy's expected advantage computed from old data. The estimate is good when the reweighting factors are near 1 and degrades as they spread out - the classic behaviour of importance sampling, where the variance of the estimate grows as the two distributions separate. That is the connection to clipping. The clip range keeps the per-sample ratios in a narrow band, which is precisely the regime where the reweighted estimate is still meaningful. Reuse and clipping are not two independent tricks: the clip is what makes the reuse defensible. ## What degrades across epochs Three things get worse as the epoch count rises. **Samples go silent.** Every epoch pushes more ratios past the band in the direction their advantage favours, and once a sample is outside it contributes zero gradient. Take a concrete run on a procedurally generated platformer: a clip range of 0.2, a 2048-step batch, ten epochs over it in minibatches. In the first epoch essentially nothing is clipped, since every ratio starts at 1. By the last epoch a large fraction of each minibatch is clipped, so the gradient is being produced by a shrinking, non-random subset - the samples the policy happens to have moved least on. You are paying full compute for a progressively narrower update. **Advantages go stale.** The advantage estimates were computed once, from the batch and the value estimates available at collection time. They describe how good each action was relative to the *old* policy's behaviour. Ten epochs later the policy has changed, so the target the update is chasing describes a policy that no longer exists. **Drift compounds.** Each epoch moves the policy a little further from the collecting policy. Clipping bounds each sample's contribution individually but does not bound the aggregate movement, so a high epoch count is one of the most reliable ways to end an update far from where it started. ## Picking the number There is no universal answer, but the shape of the tradeoff is stable. Low epoch counts waste interaction; high epoch counts waste compute and risk overshooting. Environments where interaction is cheap and abundant tolerate fewer epochs, because you can simply collect more data; environments where each step is expensive push you toward more reuse, and toward a smaller clip range to keep that reuse honest. The epoch count, the clip range and the batch size are not independent knobs - they jointly determine how far one batch is allowed to move the policy. The useful instrumentation is the fraction of samples clipped per epoch. If that fraction is still near zero in the final epoch, the batch is being under-used and you can afford more passes or a tighter clip. If it is high by the third epoch, the extra passes are burning compute to move a shrinking subset of the data, and the epoch count should come down. ## The contrast worth knowing This reuse is bounded and short-lived. It is not the same as the unlimited replay a value-based learner enjoys, where old transitions stay usable indefinitely because what is being learned - the value of a state-action pair - does not depend on the policy that produced the data in the same way. Here the licence to reuse expires as soon as the ratios spread, which is exactly what the clip range measures.

  • How would you tell from a training run that the epoch count is too high?
    Log the fraction of samples whose ratio falls outside the clip band, per epoch. If that fraction is already large by the middle epochs, the later passes are producing gradient from a shrinking, biased subset of the batch and mostly burning compute. Rising update magnitude with flat or falling returns points the same way. The fix is fewer epochs, a tighter clip, or a larger batch.
  • Why do the advantage estimates get worse as the epochs go on?
    They are computed once, before the update, from the returns in the batch and the value estimates of the moment. They express how good each action was relative to the policy that was acting then. As the epochs move the policy, that reference point drifts, so the update optimises against numbers that describe a policy it has already left behind.
  • If interaction is very expensive, why not simply raise the epoch count much further?
    Because the reuse licence comes from the ratios staying near 1, not from the data being present. Past a point the extra epochs move only the samples still inside the band, chase stale advantages, and push the policy further from the collecting distribution. The better levers are a larger batch per update and a smaller clip range, which buy reuse without letting the policy run away.

saying these in an interview costs you the question

  • Says the batch can be reused indefinitely like replay memory
  • Cannot connect the reuse to the importance ratio at all
  • Thinks more epochs always means better sample efficiency
  • Claims the advantages are recomputed each epoch
  • Treats epoch count as unrelated to the clip range

context