skip to content

Policy Optimization

Learning the policy itself with a network: a critic that cuts the variance of the score-function gradient, the advantage it estimates, and clipped updates that stop one step wrecking the policy.

on this pageshow

explore

questions

13

In an advantage actor-critic, what does the critic's learned value function contribute to the policy update?

level: middleimportance: must knowfreq 70%

answer

  1. the critic supplies a reference point
  2. signed signal, not a raw return
  3. r + gamma * V(s') - V(s)
  4. zero-centred: better than average goes up
  5. critic error becomes gradient bias

basics

~20 s

The critic learns a state-value estimate V(s). The actor is updated with the advantage - reward plus gamma times the next state's value, minus the current state's value - a signal centred on zero that reinforces only actions better than average.

solid answer

~50 s

The actor is the policy `pi(a|s)`; the critic is a second head that estimates `V(s)`, the expected discounted return from that state under the current policy. The critic never picks actions. Its value is used to form the advantage `A(s,a) = Q(s,a) - V(s)`, estimated from a single transition as the TD residual `delta = r + gamma * V(s') - V(s)`. The policy gradient becomes `grad log pi(a|s) * A`, so a positive advantage pushes that action up and a negative one pushes it down, no matter how large the raw rewards are. The critic itself is trained by regressing on the bootstrapped target `r + gamma * V(s')` with that target held fixed. The payoff is a signed, roughly zero-mean, bounded learning signal available at every step; the price is that a systematically wrong critic biases the advantage, not just the noise around it.

go deeper

for a junior

Be ready to say which network does what: the actor outputs a distribution over actions, the critic outputs one number per state. Know that the actor is trained on the advantage rather than on the raw reward.

for a middle

Expect to write the one-step advantage as reward plus gamma times the next state value minus the current state value, explain why that is a sample of Q minus V, and say how the critic itself is trained and why its target is frozen.

for a senior

Show you can tell a broken critic from a broken policy in a real run: explained variance, entropy, advantage scale. Be able to say why an inaccurate critic biases the update rather than merely adding noise.

for a principal

Own the argument for when a learned critic is worth its bias at all. On very short horizons or extremely noisy value targets, a simpler score can beat a critic that never becomes accurate, and you should be able to state what evidence would change your mind.

## The two learned functions An advantage actor-critic agent trains two things at once. The **actor** is a policy network `pi(a | s)`: it maps a state to a distribution over actions and is the thing that actually acts. The **critic** estimates `V(s)`, the expected discounted return obtained from state `s` when the current policy is followed to the end. The critic never selects an action and never takes an argmax over actions - its only job is to say how good the situation already was, before the action was chosen. ## Why the policy needs a reference point A policy network is improved with a gradient of the shape `E[ grad log pi(a|s) * W ]`, where `W` is a scalar score attached to the action that was actually taken. The vector `grad log pi(a|s)` points in the parameter direction that makes exactly that action more probable; the scalar decides how hard you push and in which direction. If `W` is a raw return in an environment where all rewards are positive, then every sampled action gets pushed up and the difference between a good and a bad action shows up only as a difference in magnitude. Learning then depends on rare large numbers rather than on a clear sign, which is a slow and noisy way to move a policy. ## The advantage Define `Q(s,a)` as the expected discounted return after taking action `a` in state `s` and following the policy afterwards, and `V(s)` as the same expectation when the action is drawn from the policy. The **advantage** is `A(s,a) = Q(s,a) - V(s)`: it answers "was this action better or worse than what this state was worth anyway?". Averaged over actions drawn from the policy, the advantage is exactly zero, so it is a naturally centred score. Actions with positive advantage are made more likely, actions with negative advantage less likely, and the size of the reward scale mostly cancels out. ## Estimating it from one transition `Q` is not available, but a single observed transition gives a sample of it: after taking `a` in `s` you see reward `r` and next state `s'`, and `r + gamma * V(s')` is a one-sample estimate of `Q(s,a)`. Subtracting the critic's `V(s)` gives the TD residual `delta = r + gamma * V(s') - V(s)` which is the cheapest advantage estimate there is. It is available the moment the transition is observed, so the actor can be updated from partial rollouts rather than waiting for an episode to end. ## Training the critic The critic minimises the squared error between `V(s)` and the bootstrapped target `r + gamma * V(s')`. The target is treated as a **constant**: no gradient is allowed to flow through the `V(s')` inside it. If you let it, the optimiser can reduce the loss by dragging the target toward the prediction instead of the other way round, and the value function drifts or collapses. Practically the two objectives are combined into one scalar - a policy term, a value term scaled by a coefficient, and usually an entropy term - and optimised together. ## What the critic buys, and what it costs What it buys: a **signed** signal (better-than-average versus worse-than-average), a signal whose magnitude stays bounded even as returns grow, and updates that do not need episode boundaries. What it costs is **bias**. A baseline that depends only on the state leaves the direction of the gradient correct in expectation, but the estimate above also *bootstraps* - the critic's own prediction `V(s')` sits inside the target. If the critic is systematically wrong in some region of the state space, the mean of the advantage estimate is wrong there, not merely noisy, and the actor is pushed in a confidently wrong direction. Early in training the critic is near-random and its advantages carry little information, which is one reason on-policy agents often sit flat for a while before improving. The standard diagnostic is the **explained variance** of the critic's predictions against observed returns: near zero means the critic explains nothing, and any conclusion the actor draws from its advantages is suspect. ## Practical detail: standardising advantages Most implementations standardise the advantages within a batch - subtract the batch mean, divide by the batch standard deviation - so the effective step size does not depend on the reward scale of the task. Two cautions. First, the subtracted mean is estimated from the same batch, so this is not a clean state-independent baseline; it is accepted in practice because the effect is small relative to the stabilisation it buys. Second, if every advantage in the batch has the same sign - common when a task has all-positive rewards and the critic is still poor - standardising forces roughly half of them negative, so merely below-average-good actions are actively suppressed. That is usually the intended relative ranking, but with a small or nearly homogeneous batch the standard deviation is tiny and dividing by it amplifies noise into large, arbitrary updates. ## The confusion to avoid The critic is not a controller. It does not rank actions, it does not choose one, and it is not a Q-network in disguise. It supplies a per-state reference level, and everything the agent does about actions comes from the actor.

  • If the critic is badly wrong, does the actor's update still point somewhere useful?
    Not reliably. A state-only baseline would leave the direction correct in expectation, but this estimate bootstraps: the critic's own prediction sits inside the target, so systematic value error shifts the mean of the advantage, not just its spread. In a region where the critic overestimates, genuinely good actions get negative advantages and are suppressed. Watch explained variance of the critic against observed returns before trusting any policy movement.
  • Why do implementations standardise advantages inside each batch, and when does that backfire?
    Standardising makes the update scale-free, so the same learning rate works across tasks with wildly different reward magnitudes. It backfires when the batch is small or the advantages are nearly identical: the standard deviation is tiny, dividing by it inflates noise into huge updates. It also converts an all-positive batch into roughly half-negative signals, which is the intended ranking but surprises people who expect positive reward to mean positive push.
  • What exactly does the critic minimise, and why is its regression target held fixed?
    It minimises the squared difference between `V(s)` and `r + gamma * V(s')`. The target is treated as a constant with no gradient flowing through it. Otherwise the optimiser can shrink the loss by moving the target toward the prediction rather than correcting the prediction, and the value function drifts toward a degenerate solution instead of tracking real returns.

A raw return tells you the salesperson closed 40,000 in deals. The advantage tells you they closed 8,000 more than anyone working that territory would have. Only the second number tells you anything about the salesperson.

saying these in an interview costs you the question

  • Says the critic selects the action or takes an argmax
  • Calls the advantage the raw episode return
  • Assumes a positive advantage means the reward was positive
  • Claims a learned critic keeps the gradient unbiased
  • Lets gradients flow through the bootstrapped value target

context

open as a page

How does PPO's clipped surrogate objective bound how far one update moves the policy?

level: middleimportance: must knowfreq 76%

basics

~20 s

PPO maximises min(r*A, clip(r, 1-eps, 1+eps)*A), where r is the new-over-old action probability ratio and A the advantage. Once r passes the bound in the direction the advantage favours, the objective flattens and that sample stops pushing.

open as a page

How does DDPG's deterministic actor get its gradient from the critic?

level: middleimportance: must knowfreq 62%

basics

~20 s

DDPG's critic scores a state-action pair and is differentiable in the action, so the actor is trained by pushing its output uphill on that score: the gradient is dQ/da evaluated at the actor's own action, chained with the actor's parameter gradient.

open as a page

Why does TD3 train two critics and take the smaller of their target values?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Because a deterministic actor is a dedicated maximiser of its critic, any upward error in the value estimate is actively sought out and then bootstrapped into future targets. Two independently initialised critics rarely err upward in the same place, so taking the smaller target cancels most of that optimism.

open as a page

Why add an entropy bonus to the policy loss when training an actor-critic agent?

level: middleimportance: should knowfreq 55%

basics

~20 s

An entropy bonus rewards a spread-out action distribution, so the policy does not collapse onto one action before the alternatives have been tried. Its coefficient trades exploration against exploitation and is usually small and decayed toward zero.

open as a page

Why does PPO run several gradient epochs over the same batch of trajectories?

level: middleimportance: should knowfreq 54%

basics

~20 s

The importance ratio corrects the objective for data gathered by the previous policy, so one batch stays usable for a few passes. Reuse amortises expensive environment interaction, and the clip keeps that correction trustworthy as the policy drifts.

open as a page

How does SAC's maximum-entropy objective differ from maximising return alone?

level: middleimportance: should knowfreq 50%

basics

~20 s

SAC maximises expected return plus a temperature times the policy's entropy at every visited state, so the optimal policy is stochastic rather than a single best action. The entropy term also enters the critic's bootstrap target, so the values learned are entropy-augmented, not ordinary returns.

open as a page

In generalized advantage estimation, what changes as lambda moves from 1.0 down toward 0?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Lambda sets how heavily the advantage estimate leans on the critic. Near 1 it sums many observed rewards: low bias, high variance. Near 0 it is one bootstrapped residual: low variance, but the critic's error passes straight through.

open as a page

When would you share one encoder trunk between the actor and critic instead of training two networks?

level: principalimportance: should knowfreq 38%

basics

~20 s

Share a trunk when observations are high-dimensional and the encoder dominates the cost: the value head's dense signal also shapes features early. Keep them separate when the state is small or the value loss swamps the policy gradient.

open as a page

When does PPO's clip alone suffice, and when do you add explicit KL control?

level: principalimportance: should knowfreq 36%

basics

~20 s

The clip bounds per-sample ratios, not the overall policy shift, and gives no restoring force once a ratio is already outside its band. Add measured-divergence control when a wrecked policy is expensive or the environment cannot be re-run cheaply.

open as a page

For a bounded physical plant, would you deploy a deterministic actor with injected exploration noise or a maximum-entropy stochastic policy?

level: principalimportance: should knowfreq 40%

basics

~20 s

Prefer the entropy-regularised stochastic learner for training, because its exploration is state-dependent and self-tuning, then ship its median action behind an external rate limiter. Choose the deterministic actor only when a fixed, auditable control law and a reproducible exploration process matter more.

open as a page

How does TRPO's hard KL constraint differ from PPO's clipped objective?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

TRPO maximises the surrogate subject to an explicit constraint that average KL between old and new policy stays under a threshold, solved by a natural-gradient step plus a line search. PPO approximates that idea first-order, with clipping.

open as a page

Why must a tanh-squashed continuous action correct its log-probability?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Squashing is a nonlinear change of variables, so it compresses probability mass near the bounds. The density of the squashed action equals the pre-squash density divided by the derivative 1 - tanh(u)^2, so the log-probability must have log(1 - tanh(u)^2) subtracted per action coordinate.

open as a page