skip to content

In REINFORCE, why does subtracting a baseline from the return cut gradient variance?

level: middleimportance: should knowfreq 57%

answer

  1. all-positive returns push everything upward
  2. you want relative, not absolute, quality
  3. centre the weights on something
  4. any shift independent of the action
  5. the expected score is exactly zero

basics

~10 s

Weighting by return minus a baseline centres the weights, so better-than-average actions are pushed up and worse-than-average ones down, instead of everything rising. Any baseline independent of the chosen action leaves the gradient unbiased.

solid answer

~50 s

Replace the weight `G_t` with `G_t - b`. Unbiasedness survives because the expected score is zero: `E over a from pi[ grad log pi(a|s) ] = 0`, so subtracting anything that does not depend on the action contributes nothing in expectation. Variance falls because of what the raw weights look like. On a portfolio-rebalancing agent whose episode returns are all positive and swing wildly, every sampled action gets its probability pushed up; only the magnitudes differ, and those magnitudes are dominated by the overall scale of returns rather than by which action was good. Centring on a running average of recent returns turns the weight into an advantage-like signal: positive for above-average episodes, negative for below-average ones. The cheapest practical baseline is that running mean; a state-dependent baseline is also valid, as long as it never depends on the action actually sampled.

code

python · 20 lines
python
import math, random, statistics

# One-parameter policy over two actions: P(a=1) = sigmoid(theta).
theta = 0.0
p = 1 / (1 + math.exp(-theta))
reward = {1: 1.0, 0: 0.2}            # mean reward under this policy = 0.6

def sample_gradient(baseline):
    a = 1 if random.random() < p else 0
    score = (1 - p) if a == 1 else -p  # d/dtheta of log P(a)
    return (reward[a] - baseline) * score

random.seed(0)
plain = [sample_gradient(0.0) for _ in range(20000)]
based = [sample_gradient(0.6) for _ in range(20000)]

print("no baseline :", round(statistics.mean(plain), 3), "+/-", round(statistics.stdev(plain), 3))
print("baseline 0.6:", round(statistics.mean(based), 3), "+/-", round(statistics.stdev(based), 3))
# no baseline : 0.198 +/- 0.3
# baseline 0.6: 0.2 +/- 0.0

go deeper

for a junior

Recall the rule of thumb: weight by return minus a running average rather than by raw return, so actions from below-average episodes get discouraged instead of merely encouraged less.

for a middle

Prove the unbiasedness. Show that the probabilities sum to one, so the expected gradient of the log-policy is zero, and explain why that means any action-independent shift is free. Then explain concretely why all-positive returns make raw weights noisy.

for a senior

Show judgment about which baseline. Talk about a running scalar versus a per-state or per-phase average, the failure mode of accidentally letting the baseline see the sampled outcome, and how you would detect that the variance reduction is actually helping.

for a principal

Frame variance reduction as a budget decision: every unit of variance removed is environment interaction you do not have to buy. Be able to argue where that engineering effort pays off versus simply collecting larger batches.

## The problem with raw returns as weights The REINFORCE estimator weights each step's score `grad log pi_theta(a_t | s_t)` by the return `G_t`. Consider a portfolio-rebalancing agent evaluated on episode profit, where returns are large, always positive, and swing wildly between episodes: one week returns 40, the next 900, the next 120. Every one of those weights is positive. So every sampled action, good or bad, gets its probability pushed *up*. The only thing distinguishing a good action from a bad one is that its nudge is somewhat larger, and that difference is buried under the enormous variation in the overall scale of returns. The estimator still points uphill in expectation, but a single batch's estimate is dominated by which episodes happened to land in a rich period, not by which actions were better than the alternatives. That is variance in the pure statistical sense: the estimator is right on average and useless per sample. ## The identity that makes baselines free Subtract a constant `b` from every weight: `grad_estimate = sum over t of (G_t - b) * grad log pi_theta(a_t | s_t)` The bias introduced is `-b * E[ grad log pi_theta(a | s) ]`. That expectation is exactly zero, and the proof is two lines: `E over a from pi[ grad log pi(a|s) ] = sum over a of pi(a|s) * grad log pi(a|s) = sum over a of grad pi(a|s) = grad ( sum over a of pi(a|s) ) = grad 1 = 0` The probabilities have to sum to one no matter what `theta` is, so the gradient of that sum is zero. This is the **expected score is zero** identity, and it is the single most useful fact about score-function estimators. It means you may subtract any quantity that is constant with respect to the action being averaged over, and the estimator's expectation is untouched. Only the variance changes. ## What the baseline changes about the signal With `b` set to roughly the mean return, the weight `G_t - b` becomes a comparison: how much better than typical was this episode? Above-average episodes get positive weights and their actions become more likely; below-average episodes get negative weights and their actions become less likely. The estimator now carries information about *relative* quality, which is what actually determines the optimal policy, instead of absolute scale, which does not. Shifting every reward in a task by +1000 changes nothing about the optimal policy, and with a baseline it changes nothing about the estimator either. Without one it inflates the variance enormously. The cheapest useful baseline is a **running average of recent episode returns**, updated as training proceeds. It costs one scalar, requires no extra model, and captures most of the benefit. A more refined choice makes the baseline depend on the state, since some states are simply richer than others; a per-timestep average handles tasks where early and late parts of an episode pay very differently. ## What a baseline may and may not depend on The validity condition is precise: the baseline must not depend on the **action** whose log-probability is being differentiated. It may depend on the state, on the timestep, on the entire training history, on a running statistic of past episodes. All of those are constants with respect to the action expectation, so they factor out of the sum and hit the zero identity. Make `b` depend on the sampled action and it no longer factors out. The `-b(s,a) * grad log pi(a|s)` term keeps a non-zero expectation, and the estimator is biased: it will converge to the optimum of some other objective. This is the single most common way people break a baseline in practice, usually by accident, for example by computing the baseline from the same sample's outcome rather than from a statistic that was fixed before the action was drawn. ## Is the mean return the best baseline? Not exactly. The variance-minimising baseline is a weighted average of returns, weighted by the squared magnitude of the score, rather than the plain mean. In practice nobody bothers: the plain running mean captures the bulk of the reduction, and the optimal baseline requires estimating second moments that are themselves noisy. The distinction is worth knowing because it shows the baseline is a variance-reduction device with an optimum, not a normalisation ritual. ## Common confusions worth clearing - **A baseline is not a reward change.** The rewards, the objective, and the optimal policy are all identical with and without it. Only the noise of the estimator differs. - **A baseline is not the same as scaling.** Dividing the weights by their standard deviation is a separate trick; it changes the effective step size and is not covered by the zero-expectation argument in the same way. - **A baseline does not eliminate the need for samples.** It reduces variance per sample; it does not make a single episode informative. The practical summary: it costs one running scalar, keeps the estimator unbiased by an identity that holds exactly, and often shrinks variance by an order of magnitude. There is essentially no reason to run REINFORCE without one.

  • Why must the baseline never depend on the action that was sampled?
    Unbiasedness rests on the expected score being zero, which only lets you pull a factor out of the action expectation if that factor is constant across actions. An action-dependent baseline stays inside the sum, contributes a non-zero term, and biases the estimator toward a different objective than expected return.
  • What baseline would you start with before you have any value estimate?
    A running average of recent episode returns. It is one scalar, needs no extra model, and removes most of the scale-driven variance. If returns differ sharply by phase of the episode or by state, refine it to a per-timestep or per-state average, keeping it independent of the sampled action.
  • Does subtracting the batch mean change which policy REINFORCE converges to?
    No. The estimator stays unbiased, so it targets the same expected-return objective with the same optima. What changes is the noise: fewer wildly misdirected steps, so you can use a larger learning rate and reach the same optimum with far fewer episodes.

Praising every student who shows up, more loudly for the ones with higher marks, is a weak teaching signal when everyone passes. Grading against the class average tells each student whether to do more of what they did or less.

saying these in an interview costs you the question

  • Says the baseline biases the gradient toward average behaviour
  • Uses a baseline computed from the action that was taken
  • Thinks the baseline changes the objective being optimised
  • Believes only more samples can reduce estimator variance
  • Cannot state that the expected score function is zero

context