skip to content

Why must a REINFORCE run discard its collected trajectories after a single gradient update?

level: seniorimportance: should knowfreq 43%

answer

  1. look at the subscript on the expectation
  2. the samples belong to a specific policy
  3. one step and the sampling distribution moved
  4. state visitation drifts, not just actions
  5. importance weights blow up with divergence

basics

~20 s

The policy gradient is an expectation over trajectories drawn from the current policy. One update moves the policy, so the batch now comes from a different distribution and estimates the wrong expectation. Every step needs fresh interaction.

solid answer

~50 s

The estimator is `E over trajectories from pi_theta [ ... ]`, with the sampling distribution indexed by the very parameters you are updating. After a step, `theta` has moved, so the stored batch is a sample from the old policy and averaging over it estimates the old policy's gradient, not the current one. That is what on-policy means, and it is a correctness constraint, not a style choice. The cost is brutal: for a Gaussian policy pacing hourly ad budgets, where each episode is a full campaign flight and today's spend changes the budget state left for the rest of it, every gradient step costs a fresh batch of real flights. Importance sampling can buy limited reuse by reweighting each action by the ratio of new to old probability, but the weights' variance explodes as the policies diverge, so the reuse window is a handful of steps at best.

go deeper

for a junior

Recall the label and its consequence: REINFORCE is on-policy, so the episodes you collect are spent after one update and the next step needs fresh ones. Contrast that with methods that can replay stored experience.

for a middle

Explain the mechanism precisely. The gradient is an expectation over trajectories from the current policy, so once the parameters move the stored batch estimates the wrong expectation, and the mismatch covers visited states as well as chosen actions.

for a senior

Show you have paid this bill. Reason about batch size against update count, when importance-weighted reuse is worth its variance, and what it means operationally when an episode is a whole campaign flight or a physical run rather than a simulator tick.

for a principal

Own the budget conversation. Decide up front how many environment interactions the problem can justify, whether a simulator faithful enough to absorb them exists, and at what point a cheaper decision-making approach is the better engineering answer.

## Where the constraint comes from The policy gradient is `grad J(theta) = E over tau drawn from pi_theta [ sum_t G_t * grad log pi_theta(a_t | s_t) ]` Read the subscript on the expectation carefully. The distribution being averaged over is *the current policy's*. A Monte-Carlo estimate of an expectation is only valid if the samples come from the distribution named under the E. Your batch of trajectories came from whatever policy was in force when you collected it. The instant you apply an update, that is no longer the policy you hold, and reusing the batch silently estimates a different quantity: the gradient at the old parameters. For one tiny step the discrepancy is small, and in practice people do sometimes take two or three steps on a batch and accept the error. But it is an error, growing with how far the policy has moved, and unlike ordinary optimisation noise it does not average out: it is a systematic bias pointing back at where the policy used to be. ## Why it is worse than "the data is a bit stale" The drift is not only in which actions were chosen; it is in which **states were visited**. Consider a Gaussian policy pacing an hourly advertising budget, where the state includes how much budget remains for the rest of the flight. A policy that spends aggressively in the morning produces afternoon states with little budget left. Update the policy toward more conservative morning spending, and the afternoon states in your stored batch are states the new policy would rarely reach at all. You are computing a gradient over a state distribution that has become counterfactual. This compounding through the state visitation distribution is why off-policy correction in sequential decision problems is much harder than in a bandit or a supervised setting. ## The bill The accounting is unforgiving. Each gradient step consumes one batch of complete episodes, and each of those episodes is real interaction with the environment. In the ad-pacing case an episode is an entire campaign flight, so a training run that needs a few thousand updates needs a few thousand batches of flights. On a physical system, each episode also carries wall-clock time, wear, and the risk of a bad action being genuinely bad rather than merely uninformative. This is the structural difference from value-based learning. A Q-learning-style update is valid for transitions collected under any behaviour that covered the relevant actions, so a replay buffer can be revisited many times. A policy gradient batch has one shot. ## Partial escapes, and what each costs **Importance sampling.** You can reuse a batch collected under `pi_old` while holding `pi_new` by reweighting: multiply each term by the ratio `pi_new(a|s) / pi_old(a|s)`. In principle this is unbiased. In practice the ratios multiply along a trajectory and their variance grows quickly as the policies diverge; after a few updates a handful of trajectories with huge weights dominate the average and the estimate is worse than no reuse at all. The honest description is that importance weights extend the useful life of a batch from one update to a small number, provided you also keep the policy from moving far, and that constraining the step is itself a separate piece of machinery. **Bigger batches.** Collecting more episodes per update lowers the noise in each gradient step but does not reduce total interaction; if anything it raises it, because you spend more data to take the same number of steps. It buys stability, not economy. Choosing the batch size is a genuine trade between gradient quality and number of updates you can afford. **Parallel collection.** Running many copies of a simulator concurrently does not reduce the number of interactions required; it reduces the wall-clock time to gather them. That is the right fix when interaction is cheap but slow, and no help at all when interaction is intrinsically scarce or expensive. **Variance reduction.** Anything that lowers the variance of the estimator, a baseline being the cheapest, means fewer episodes are needed per usable gradient step. This is the highest-leverage lever available inside plain REINFORCE, because it attacks the number of samples needed rather than trying to recycle samples you already spent. ## What an interviewer is checking That you can name the exact reason: the expectation is indexed by the current policy. Weak answers say "the data is stale" or "on-policy means you learn online", both of which miss the mechanism. Strong answers connect it to the operating consequence, that gradient steps are metered in units of environment interaction, and to the honest limits of the workarounds: importance weights degrade with policy divergence, larger batches trade stability against update count, and parallelism buys clock time rather than data.

  • How does importance sampling extend a batch's life, and where does it break down?
    Each sampled action is reweighted by the ratio of its probability under the new policy to its probability under the collecting policy, which restores the correct expectation in principle. The ratios compound along a trajectory, so their variance grows as the policies diverge; a few enormous weights come to dominate the average, and the reuse window closes after a handful of steps.
  • Would collecting a much larger batch per update solve the sample cost?
    No. It lowers the noise in each gradient step, which makes training more stable, but the total interaction to reach a given policy quality does not fall and can rise, because you spend more episodes per step. It is a stability lever, not an economy one. The lever that actually reduces episodes needed is variance reduction in the estimator itself.
  • Why is this constraint harsher here than for a value-based method learning from a replay buffer?
    A value-based update targets a quantity defined per transition, so transitions gathered under other behaviour are still usable as long as they cover the relevant actions. A policy gradient is an expectation over trajectories from the current policy, and the mismatch compounds through the state visitation distribution as well as the action choice.

An exit poll taken before you redesigned the ballot describes the old ballot's voters. You can reweight it toward the new electorate, but the further the design has moved, the more the estimate rests on a few unusual respondents.

saying these in an interview costs you the question

  • Says old trajectories can be reused freely with no correction
  • Thinks on-policy simply means learning online or incrementally
  • Claims importance weights stay well behaved at any divergence
  • Treats a larger batch as a cure for total interaction cost
  • Ignores that the visited-state distribution shifts too

context