skip to content

In generalized advantage estimation, what changes as lambda moves from 1.0 down toward 0?

level: seniorimportance: should knowfreq 44%

answer

  1. one dial between two estimators
  2. exponential weights over TD residuals
  3. lambda zero trusts the critic fully
  4. lambda one telescopes to return minus value
  5. lambda tunes the estimator, gamma the objective

basics

~20 s

Lambda sets how heavily the advantage estimate leans on the critic. Near 1 it sums many observed rewards: low bias, high variance. Near 0 it is one bootstrapped residual: low variance, but the critic's error passes straight through.

solid answer

~50 s

Generalized advantage estimation forms `A_t = sum over l of (gamma * lambda)^l * delta_(t+l)`, where `delta_t = r_t + gamma * V(s_(t+1)) - V(s_t)`. At `lambda = 0` this collapses to the single residual `delta_t`, which is cheap and low-variance but inherits the critic's error directly. At `lambda = 1` the residuals telescope into the observed discounted return minus `V(s_t)`, which barely depends on critic accuracy but carries the full variance of the sampled trajectory. Typical values sit around 0.9 to 0.97. The knob to remember is that lambda changes only the *estimator* of the advantage - the thing you are optimising is unchanged - whereas gamma changes the objective itself by redefining which future rewards count. Pick lambda by critic quality and horizon: an untrained or badly fitting critic argues for higher lambda, an accurate critic on a long noisy horizon argues for lower.

code

python · 12 lines
python
rewards = [0.0, 0.0, 0.0, 1.0]
values = [0.5, 0.6, 0.7, 0.8, 0.0]   # V(s0)..V(s4); s4 is terminal
gamma, lam = 0.99, 0.95

adv = [0.0] * len(rewards)
running = 0.0
for t in reversed(range(len(rewards))):
    delta = rewards[t] + gamma * values[t + 1] - values[t]
    running = delta + gamma * lam * running
    adv[t] = running

print([round(a, 3) for a in adv])

go deeper

for a junior

Know that lambda is a number between 0 and 1 that controls how much of the observed future reward, versus the critic's prediction, goes into the advantage estimate.

for a middle

Be able to state both endpoints: lambda zero is the single one-step residual, lambda one is the observed discounted return minus the state value, and explain which one is noisier and which one leans on the critic.

for a senior

Show that you set lambda from evidence - critic explained variance, reward delay, rollout length - and that you can name the symptom each bad setting produces in a real run rather than quoting a default.

for a principal

Own the framing that lambda is an estimator knob and gamma is a problem definition. Be ready to argue when a team should invest in a better critic instead of tuning lambda at all, and what that investment costs.

## The estimator Write the one-step residual at time `t` as `delta_t = r_t + gamma * V(s_(t+1)) - V(s_t)` Generalized advantage estimation weights these residuals with an exponentially decaying window: `A_t = delta_t + (gamma*lambda) * delta_(t+1) + (gamma*lambda)^2 * delta_(t+2) + ...` which is computed in one backward pass over a rollout using the recursion `A_t = delta_t + gamma * lambda * A_(t+1)`, with `A` initialised to zero at the end of the rollout. ## The two endpoints At **lambda = 0** only the first term survives and `A_t = delta_t`. This is the cheapest possible estimate. It uses exactly one observed reward and otherwise trusts the critic completely. Its variance is small - one reward sample - and its error is whatever the critic's error is. If the critic is biased in a region of state space, every advantage in that region is biased the same way, and the actor is pushed confidently in the wrong direction. At **lambda = 1** the sum telescopes. Expanding the residuals, all the intermediate value terms cancel and what is left is `A_t = (r_t + gamma*r_(t+1) + gamma^2*r_(t+2) + ...) - V(s_t)` that is, the observed discounted return with the critic's value used only as a subtracted reference. The critic's accuracy now affects only the size of the variance reduction, not the expected direction of the update. The cost is that the estimate contains the full randomness of the whole future trajectory: every stochastic transition and every stochastic reward from `t` to the end of the rollout is baked in. Everything between the endpoints is a continuous dial between those two regimes. `1 / (1 - gamma*lambda)` is a useful reading of the effective number of steps the estimate looks ahead before the critic takes over. ## Choosing lambda in practice The question to ask is: *how much do I trust the critic, over how long a horizon?* On a long-horizon dispatch task - a ride-hailing agent whose decision to send a car across the city pays off twenty minutes later - sweeping lambda from 0.9 to 1.0 shows the trade directly. At 0.9 the advantage estimates are tight and gradients are quiet, but early in training, when the value function has not yet learned that a repositioning move is worth anything, the advantage for that move is close to zero and the policy simply never learns to reposition. Near 1.0 the delayed payoff is visible in the estimate immediately - it is in the observed rewards - but the estimate also absorbs every unrelated fluctuation in demand over those twenty minutes, and the gradient becomes noisy enough that the policy wanders. Practical heuristics that follow from this: - Start higher when the critic is weak or the reward is delayed and sparse; the critic cannot yet carry the credit assignment. - Lower it once the critic explains most of the variance in observed returns; then bootstrapping is nearly free and you are simply discarding noise. - Very high lambda with a good critic is wasteful: you pay full trajectory variance for bias you had already removed. - Very low lambda with a poor critic is dangerous: it is confident and wrong, which is worse than noisy and right on average. ## Lambda is not a second discount factor This is the distinction interviewers probe. Gamma is part of the **problem**: it defines the return you are trying to maximise, so lowering it makes the agent genuinely more short-sighted and changes which policy is optimal. Lambda is part of the **estimator**: whatever value it takes, the quantity being estimated is the advantage under the same gamma, and the optimal policy is untouched. Two runs with different gammas are solving different problems; two runs with different lambdas are solving the same problem with different amounts of bias and noise in the gradient. ## Interaction with rollout length Rollouts are finite. If you collect `T` steps before each update, the sum is truncated at the boundary and the value at the boundary state is bootstrapped in. That means the lambda = 1 endpoint is not truly unbiased in practice - it still leans on the critic at the truncation point - and it also means that a lambda whose effective horizon `1/(1 - gamma*lambda)` far exceeds `T` buys you nothing: the terms you were paying variance for have been cut off anyway. Choose lambda and rollout length together. ## Diagnosing a bad setting Track the critic's explained variance against observed returns alongside the gradient norm. Low explained variance plus a low lambda is the classic broken combination: the actor is confidently following a critic that knows nothing. High lambda plus an erratic gradient norm and a policy that improves and then regresses points the other way - reduce lambda, or collect larger batches, before touching the learning rate.

  • How does the length of each rollout interact with your choice of lambda?
    The weighted sum is truncated at the rollout boundary and the boundary state's value is bootstrapped in. So a lambda whose effective horizon `1/(1 - gamma*lambda)` far exceeds the rollout length buys nothing - those terms were cut off - while still costing you variance from the terms that remain. It also means lambda = 1 is not genuinely unbiased with short rollouts, because the critic still enters at the truncation point.
  • Would you change lambda over the course of a run?
    It is defensible: the critic improves, so bootstrapping gets cheaper and a lower lambda discards noise you no longer need. In practice most teams keep it fixed, because changing the estimator mid-run makes runs incomparable and makes a regression hard to attribute. If you do schedule it, schedule it alone and keep everything else frozen.
  • How would you tell empirically that lambda is set too low?
    Look at the critic's explained variance against observed returns. If it is low while lambda is low, the advantages are mostly the critic's error, and the symptom is a policy that stops improving on exactly the delayed-payoff decisions - it never learns actions whose reward arrives outside the bootstrap. Raise lambda, or fix the critic, before touching the actor's learning rate.

saying these in an interview costs you the question

  • Treats lambda and gamma as interchangeable discount factors
  • Says lambda equals one is unbiased even with truncated rollouts
  • Claims lowering lambda reduces bias
  • Picks lambda without regard to critic accuracy
  • Thinks lambda changes which policy is optimal

context