skip to content

Advantage Actor-Critic

A learned value network is the baseline, so the actor follows an advantage rather than the raw return, and the GAE lambda tunes the variance it absorbs. Interviewers probe the two-network split.

on this pageshow

questions

4

In an advantage actor-critic, what does the critic's learned value function contribute to the policy update?

level: middleimportance: must knowfreq 70%

answer

  1. the critic supplies a reference point
  2. signed signal, not a raw return
  3. r + gamma * V(s') - V(s)
  4. zero-centred: better than average goes up
  5. critic error becomes gradient bias

basics

~20 s

The critic learns a state-value estimate V(s). The actor is updated with the advantage - reward plus gamma times the next state's value, minus the current state's value - a signal centred on zero that reinforces only actions better than average.

solid answer

~50 s

The actor is the policy `pi(a|s)`; the critic is a second head that estimates `V(s)`, the expected discounted return from that state under the current policy. The critic never picks actions. Its value is used to form the advantage `A(s,a) = Q(s,a) - V(s)`, estimated from a single transition as the TD residual `delta = r + gamma * V(s') - V(s)`. The policy gradient becomes `grad log pi(a|s) * A`, so a positive advantage pushes that action up and a negative one pushes it down, no matter how large the raw rewards are. The critic itself is trained by regressing on the bootstrapped target `r + gamma * V(s')` with that target held fixed. The payoff is a signed, roughly zero-mean, bounded learning signal available at every step; the price is that a systematically wrong critic biases the advantage, not just the noise around it.

go deeper

for a junior

Be ready to say which network does what: the actor outputs a distribution over actions, the critic outputs one number per state. Know that the actor is trained on the advantage rather than on the raw reward.

for a middle

Expect to write the one-step advantage as reward plus gamma times the next state value minus the current state value, explain why that is a sample of Q minus V, and say how the critic itself is trained and why its target is frozen.

for a senior

Show you can tell a broken critic from a broken policy in a real run: explained variance, entropy, advantage scale. Be able to say why an inaccurate critic biases the update rather than merely adding noise.

for a principal

Own the argument for when a learned critic is worth its bias at all. On very short horizons or extremely noisy value targets, a simpler score can beat a critic that never becomes accurate, and you should be able to state what evidence would change your mind.

## The two learned functions An advantage actor-critic agent trains two things at once. The **actor** is a policy network `pi(a | s)`: it maps a state to a distribution over actions and is the thing that actually acts. The **critic** estimates `V(s)`, the expected discounted return obtained from state `s` when the current policy is followed to the end. The critic never selects an action and never takes an argmax over actions - its only job is to say how good the situation already was, before the action was chosen. ## Why the policy needs a reference point A policy network is improved with a gradient of the shape `E[ grad log pi(a|s) * W ]`, where `W` is a scalar score attached to the action that was actually taken. The vector `grad log pi(a|s)` points in the parameter direction that makes exactly that action more probable; the scalar decides how hard you push and in which direction. If `W` is a raw return in an environment where all rewards are positive, then every sampled action gets pushed up and the difference between a good and a bad action shows up only as a difference in magnitude. Learning then depends on rare large numbers rather than on a clear sign, which is a slow and noisy way to move a policy. ## The advantage Define `Q(s,a)` as the expected discounted return after taking action `a` in state `s` and following the policy afterwards, and `V(s)` as the same expectation when the action is drawn from the policy. The **advantage** is `A(s,a) = Q(s,a) - V(s)`: it answers "was this action better or worse than what this state was worth anyway?". Averaged over actions drawn from the policy, the advantage is exactly zero, so it is a naturally centred score. Actions with positive advantage are made more likely, actions with negative advantage less likely, and the size of the reward scale mostly cancels out. ## Estimating it from one transition `Q` is not available, but a single observed transition gives a sample of it: after taking `a` in `s` you see reward `r` and next state `s'`, and `r + gamma * V(s')` is a one-sample estimate of `Q(s,a)`. Subtracting the critic's `V(s)` gives the TD residual `delta = r + gamma * V(s') - V(s)` which is the cheapest advantage estimate there is. It is available the moment the transition is observed, so the actor can be updated from partial rollouts rather than waiting for an episode to end. ## Training the critic The critic minimises the squared error between `V(s)` and the bootstrapped target `r + gamma * V(s')`. The target is treated as a **constant**: no gradient is allowed to flow through the `V(s')` inside it. If you let it, the optimiser can reduce the loss by dragging the target toward the prediction instead of the other way round, and the value function drifts or collapses. Practically the two objectives are combined into one scalar - a policy term, a value term scaled by a coefficient, and usually an entropy term - and optimised together. ## What the critic buys, and what it costs What it buys: a **signed** signal (better-than-average versus worse-than-average), a signal whose magnitude stays bounded even as returns grow, and updates that do not need episode boundaries. What it costs is **bias**. A baseline that depends only on the state leaves the direction of the gradient correct in expectation, but the estimate above also *bootstraps* - the critic's own prediction `V(s')` sits inside the target. If the critic is systematically wrong in some region of the state space, the mean of the advantage estimate is wrong there, not merely noisy, and the actor is pushed in a confidently wrong direction. Early in training the critic is near-random and its advantages carry little information, which is one reason on-policy agents often sit flat for a while before improving. The standard diagnostic is the **explained variance** of the critic's predictions against observed returns: near zero means the critic explains nothing, and any conclusion the actor draws from its advantages is suspect. ## Practical detail: standardising advantages Most implementations standardise the advantages within a batch - subtract the batch mean, divide by the batch standard deviation - so the effective step size does not depend on the reward scale of the task. Two cautions. First, the subtracted mean is estimated from the same batch, so this is not a clean state-independent baseline; it is accepted in practice because the effect is small relative to the stabilisation it buys. Second, if every advantage in the batch has the same sign - common when a task has all-positive rewards and the critic is still poor - standardising forces roughly half of them negative, so merely below-average-good actions are actively suppressed. That is usually the intended relative ranking, but with a small or nearly homogeneous batch the standard deviation is tiny and dividing by it amplifies noise into large, arbitrary updates. ## The confusion to avoid The critic is not a controller. It does not rank actions, it does not choose one, and it is not a Q-network in disguise. It supplies a per-state reference level, and everything the agent does about actions comes from the actor.

  • If the critic is badly wrong, does the actor's update still point somewhere useful?
    Not reliably. A state-only baseline would leave the direction correct in expectation, but this estimate bootstraps: the critic's own prediction sits inside the target, so systematic value error shifts the mean of the advantage, not just its spread. In a region where the critic overestimates, genuinely good actions get negative advantages and are suppressed. Watch explained variance of the critic against observed returns before trusting any policy movement.
  • Why do implementations standardise advantages inside each batch, and when does that backfire?
    Standardising makes the update scale-free, so the same learning rate works across tasks with wildly different reward magnitudes. It backfires when the batch is small or the advantages are nearly identical: the standard deviation is tiny, dividing by it inflates noise into huge updates. It also converts an all-positive batch into roughly half-negative signals, which is the intended ranking but surprises people who expect positive reward to mean positive push.
  • What exactly does the critic minimise, and why is its regression target held fixed?
    It minimises the squared difference between `V(s)` and `r + gamma * V(s')`. The target is treated as a constant with no gradient flowing through it. Otherwise the optimiser can shrink the loss by moving the target toward the prediction rather than correcting the prediction, and the value function drifts toward a degenerate solution instead of tracking real returns.

A raw return tells you the salesperson closed 40,000 in deals. The advantage tells you they closed 8,000 more than anyone working that territory would have. Only the second number tells you anything about the salesperson.

saying these in an interview costs you the question

  • Says the critic selects the action or takes an argmax
  • Calls the advantage the raw episode return
  • Assumes a positive advantage means the reward was positive
  • Claims a learned critic keeps the gradient unbiased
  • Lets gradients flow through the bootstrapped value target

context

open as a page

Why add an entropy bonus to the policy loss when training an actor-critic agent?

level: middleimportance: should knowfreq 55%

basics

~20 s

An entropy bonus rewards a spread-out action distribution, so the policy does not collapse onto one action before the alternatives have been tried. Its coefficient trades exploration against exploitation and is usually small and decayed toward zero.

open as a page

In generalized advantage estimation, what changes as lambda moves from 1.0 down toward 0?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Lambda sets how heavily the advantage estimate leans on the critic. Near 1 it sums many observed rewards: low bias, high variance. Near 0 it is one bootstrapped residual: low variance, but the critic's error passes straight through.

open as a page

When would you share one encoder trunk between the actor and critic instead of training two networks?

level: principalimportance: should knowfreq 38%

basics

~20 s

Share a trunk when observations are high-dimensional and the encoder dominates the cost: the value head's dense signal also shapes features early. Keep them separate when the state is small or the value loss swamps the policy gradient.

open as a page