skip to content

How does DDPG's deterministic actor get its gradient from the critic?

level: middleimportance: must knowfreq 62%

answer

  1. the gradient flows through the critic
  2. chain rule, not a log-probability
  3. dQ/da at the actor's own action
  4. actor maximises the critic's own score
  5. critic weights frozen for that step

basics

~20 s

DDPG's critic scores a state-action pair and is differentiable in the action, so the actor is trained by pushing its output uphill on that score: the gradient is dQ/da evaluated at the actor's own action, chained with the actor's parameter gradient.

solid answer

~50 s

DDPG keeps two networks: a deterministic actor `mu(s)` that emits an action vector, and a critic `Q(s, a)` that scores it. Because the critic takes the action as a real-valued input, it is differentiable with respect to that input, and that is the whole trick. The actor's objective is `J = mean over states of Q(s, mu(s))`, and its gradient is `dQ/da` at `a = mu(s)` times `d mu/d theta` — a plain chain rule through the critic and back into the actor. In implementation terms the actor's loss is the negative mean critic score of its own actions, and the critic's parameters are held fixed while that step is taken. No sampled action log-probability appears anywhere, which is why this is called the deterministic policy gradient rather than a score-function estimator. The critic itself is trained separately by regressing onto bootstrapped TD targets.

go deeper

for a junior

Recall the shape: one network chooses the action, another scores it, and the chooser is trained to make the scorer happy. Know that a deterministic actor needs noise added during data collection or it never explores.

for a middle

Be ready to derive the update: the actor's objective is the critic's score of the actor's own action, and the gradient is dQ/da chained with the actor's parameter gradient. Say explicitly that the critic's weights are frozen during that step.

for a senior

Show you have debugged it. Talk about action scaling and saturated outputs killing the actor's gradient, and about the actor reliably finding whatever region the critic is wrong about.

for a principal

Own the trade being made: an analytic, low-variance gradient bought with complete dependence on one learned critic's derivative. Be able to argue when that dependence is acceptable and what pessimism or ensembling you would add before trusting it on a real plant.

## The problem this solves In continuous control the action is a vector of real numbers — joint torques, valve openings, steering angles — so an agent cannot pick an action by scanning a finite list of Q-values. Something has to perform the maximisation over an uncountable action set. DDPG's answer is to *learn the maximiser*: a network whose output is, by training, an action the critic rates highly. ## The two networks - **Actor** `mu_theta(s)`: state in, a single action vector out. It is deterministic — one state, one action, no distribution. - **Critic** `Q_phi(s, a)`: state and action in, one scalar out, an estimate of the expected discounted return from taking `a` in `s` and following the actor afterwards. The critic is a differentiable function of *both* inputs. Everyone thinks of the state input; the action input is the one that matters here. ## The deterministic policy gradient The actor's objective is the critic's own opinion of the actor: J(theta) = E over states s of Q_phi(s, mu_theta(s)) Differentiate by the chain rule: dJ/d theta = E_s [ (dQ_phi/da at a = mu_theta(s)) * (d mu_theta(s)/d theta) ] The first factor is a vector the size of the action: it says, for this state, in which direction each action coordinate should move to raise the predicted return. The second factor maps that action-space direction back into parameter space. In practice you write the actor's loss as the negative mean of `Q(s, mu(s))` over a batch of states, run the backward pass through the critic's layers into the action, and continue into the actor's weights. The critic's own parameters are excluded from this update — the critic acts purely as a differentiable scoring function for one step, and is trained by its own regression step. ## Why this is not the score-function gradient A stochastic policy gradient estimates `E[ grad log pi(a|s) * (something like a return) ]`: it perturbs actions by sampling and reweights by how good the sample turned out. The deterministic gradient never samples in action space at all — the policy is a point mass, so the expectation over actions collapses and only states are averaged over. The consequence is a low-variance estimator: you get the *analytic* local slope of the value surface instead of a noisy sample-based estimate of it. The price is bias and total dependence on the critic. The score-function estimator only needs an unbiased-ish scalar signal; the deterministic gradient needs the critic's **derivative** to point the right way. A critic can predict values with small average error and still have a badly wrong action-gradient in some region — and the actor will follow it there. ## The critic's training loop The critic regresses onto bootstrapped targets built from transitions `(s, a, r, s')` collected earlier and stored off-policy: y = r + gamma * Q(s', mu(s')) (with a slowly-moving copy used for the target side) critic loss = (Q(s, a) - y)^2 Because the actor supplies `mu(s')` directly, no maximisation over actions is ever performed explicitly. ## Exploration A deterministic actor returns the same action for the same state forever, so it explores nothing by itself. During data collection you add noise to the emitted action — independent Gaussian noise, or temporally correlated noise when the plant needs sustained pushes rather than jitter. Exploration is therefore an external, hand-scheduled component of the algorithm, not something the objective produces. ## Practical failure modes to be able to name - **Actions on the wrong scale.** If different action coordinates span wildly different physical ranges, the critic's action-gradient is badly conditioned and the actor moves mostly along one coordinate. Normalising the action space to a common range is standard. - **A saturated output nonlinearity.** If the actor squashes its output and drives deep into saturation, `d a / d theta` shrinks toward zero and the actor stops learning even though the critic still wants a bigger action. - **Following critic errors.** The actor is a dedicated optimiser aimed at the critic, so wherever the critic is spuriously optimistic, the actor goes there. This is the single most important weakness of the vanilla algorithm and the reason later methods add pessimism to the critic. ## What an interviewer is listening for That you can say "the critic is differentiable in the action, so we backpropagate through it" without hedging, that you know the critic's weights are frozen for that step, and that you can name what the actor's optimisation does to critic errors.

  • Why does the deterministic policy gradient usually have lower variance than a score-function estimator?
    Because it never samples in action space. The policy is a point mass, so the expectation over actions disappears and only states are averaged; the update uses the critic's analytic derivative instead of a noisy sampled signal reweighted by a log-probability. The trade is bias for variance: the estimator is only as good as the critic's action-gradient.
  • Are the critic's parameters updated by the actor's loss?
    No. The backward pass runs through the critic's activations to reach the action, but only the actor's parameters take the step. The critic is trained separately by regressing onto bootstrapped TD targets. Letting the actor loss touch the critic would let the agent raise its own score by editing the scorer.
  • What happens when the critic's action-gradient is wrong in some region of action space?
    The actor walks straight into it. The actor is a dedicated maximiser of the critic, so a spuriously high patch of the value surface is exactly what it seeks out, and the resulting action then feeds the critic's own bootstrap targets. Predicted values climb while real return stalls or collapses.

The critic is a differentiable altimeter for action space: the actor does not guess-and-check heights, it reads the local slope off the altimeter and steps uphill.

saying these in an interview costs you the question

  • Says the actor is updated by a sampled action's log-probability
  • Thinks a deterministic actor explores on its own
  • Claims the critic's weights are trained by the actor's loss
  • Says the algorithm needs an argmax over a discretised action grid
  • Ignores that the critic must be differentiable in the action input

context