skip to content

Why does TD3 train two critics and take the smaller of their target values?

level: seniorimportance: must knowfreq 55%

answer

  1. the actor maximises the critic's errors
  2. max of a noisy estimate is biased up
  3. the error bootstraps into earlier targets
  4. two independent critics, take the minimum
  5. pessimism does not self-amplify

basics

~20 s

Because a deterministic actor is a dedicated maximiser of its critic, any upward error in the value estimate is actively sought out and then bootstrapped into future targets. Two independently initialised critics rarely err upward in the same place, so taking the smaller target cancels most of that optimism.

solid answer

~50 s

In a single-critic deterministic actor-critic, the actor is trained to maximise `Q(s, mu(s))`, so it finds whatever action the critic is spuriously optimistic about. That inflated value then enters the bootstrap target for earlier states, and the error compounds — a locomotion agent will show predicted Q climbing steadily while its actual return flatlines or collapses. TD3 keeps two critics with independent initialisations, trains both on the *same* target, and builds that target from the minimum of the two: `y = r + gamma * min(Q1_target, Q2_target)(s', a')`. Two networks can each be wrong, but they are unlikely to be optimistically wrong about the *same* action, so the minimum strips off most of the overestimation. The cost is a deliberate underestimation bias, which is accepted because pessimism does not self-amplify: an action the agent undervalues is one the actor stops visiting, so the error stops being reinforced.

go deeper

for a junior

Recall that two critics are trained and the smaller estimate is used when forming the learning target, and that this exists to stop value estimates inflating. Know that the two critics differ only by their random initialisation.

for a middle

Explain the mechanism: maximising a noisy value function returns whichever action had the largest positive error, and bootstrapping feeds that error backwards. Write down the target with the minimum in it and say that both critics regress onto that same target.

for a senior

Show the diagnosis. Describe watching predicted value against observed return, and be ready to explain why deliberate pessimism is a sound trade and what the delayed actor updates and target smoothing noise each fix.

for a principal

Own the estimator-design argument: an asymmetric bias is justified when the consequences of the two errors are asymmetric under a maximising policy. Be able to say where more pessimism stops paying and what you would monitor before trusting these values in a real deployment.

## Where the overestimation comes from A critic is a regression model, so its output is the true value plus error, and that error is not tiny in the regions the agent has barely visited. Now add an actor whose entire job is `argmax_a Q(s, a)`. Maximising a noisy function does not return the maximum of the true function — it returns whichever point had the largest positive error. Formally, the expectation of a maximum is at least the maximum of the expectations, so the maximisation step is *systematically* biased upward even when the critic's errors are zero-mean. With a deterministic actor the maximisation is not a one-off: it is a trained network that spends every update searching for the critic's optimistic patches, and it does so faster than the critic can be corrected by new data. ## Why it compounds The critic bootstraps. The target for a transition `(s, a, r, s')` is `r + gamma * Q(s', a')`, where `a'` is the actor's action at the next state — precisely the action the actor has been optimising to maximise `Q`. So the inflated value at `s'` becomes the regression target at `s`, which inflates `s`, which inflates its predecessors. On a locomotion benchmark this shows up as an unmistakable diagnostic: the mean predicted Q on batches from the buffer rises smoothly and without bound, while the episode return plateaus and then falls. The critic has drifted into fantasy values, and the actor is faithfully chasing them. ## The fix: clipped double Q-learning TD3 instantiates **two critics**, `Q1` and `Q2`, with independent random initialisations. They see the same data and are trained on the same shared target, and that target uses the smaller of the two slowly-updated target critics: a' = actor_target(s') (plus smoothing noise, see below) y = r + gamma * min( Q1_target(s', a'), Q2_target(s', a') ) loss = (Q1(s, a) - y)^2 + (Q2(s, a) - y)^2 The two critics stay different only because of their initialisation and their independent optimisation paths — that is enough. The intuition is a vote for pessimism: for the minimum to be optimistic, *both* networks must be optimistic about the same action at the same state, which is far rarer than either one being optimistic alone. Note that the actor is trained against one designated critic (conventionally `Q1`), not the minimum; the pessimism is what matters in the target that drives learning. ## Why deliberate underestimation is acceptable The minimum is not unbiased — it pushes the estimate low. That is accepted on purpose because the two biases behave very differently under a maximising actor: - **Overestimation is self-amplifying.** The actor seeks out the overvalued action, so the wrong value gets rehearsed in every target and the error grows. - **Underestimation is self-correcting-ish.** An undervalued action is one the actor avoids, so the error simply stops being reinforced; the worst case is that a good action is explored less than it deserves. Asymmetric consequences justify an asymmetric estimator. The failure mode of too much pessimism is real, though — pushing the idea to many critics and taking the minimum over all of them makes the agent timid and slow to commit, which is why two is the standard choice rather than five. ## The two companions in the same algorithm The twin critics are one of three changes, and an interviewer usually wants all three: - **Delayed policy updates.** The actor and the slow target copies are updated once per two critic updates. The actor is then optimising against a critic that has had time to settle, rather than chasing a value surface that is itself moving fast — which reduces the chance of locking onto transient errors. - **Target policy smoothing.** Small clipped noise is added to the action used in the target, so the target value is effectively an average of `Q` over a neighbourhood of that action rather than its value at one exact point. Without it, a narrow spurious spike in the critic is a perfectly good thing for the actor to climb; a quadruped will discover a razor-thin peak in the critic and produce a gait that scores wonderfully and moves nowhere. Smoothing enforces the reasonable prior that similar actions should have similar value. ## Diagnosing it in practice Log the mean predicted value on replayed batches alongside the mean observed discounted return. Healthy training keeps them in the same neighbourhood. A widening gap with the prediction on top is the overestimation signature, and it is the fastest single diagnostic in continuous-control debugging.

  • What do delayed policy updates add on top of the twin critics?
    They update the actor and the slow target copies once per two critic updates, so the actor optimises against a critic that has had time to settle instead of chasing a fast-moving value surface. It reduces the chance of the actor locking onto transient critic errors, at the cost of half as many policy improvement steps per unit of data.
  • What failure does target policy smoothing prevent?
    Adding small clipped noise to the action used in the target makes the target an average of the critic over a neighbourhood rather than its value at one exact point. Without it a narrow spurious spike in the critic is a valid thing to climb, and the actor produces a gait or motion that scores highly and does nothing useful.
  • Doesn't taking the minimum introduce an underestimation bias?
    Yes, deliberately. It is accepted because the two errors are not symmetric in consequence: an overvalued action is exactly what the actor seeks out and rehearses into future targets, while an undervalued action is simply avoided and stops being reinforced. Pushing the idea further, by minimising over many critics, does make the agent excessively timid.
  • How would you detect this problem from training logs alone?
    Plot the critic's mean predicted value on replayed batches against the mean observed discounted return. They should track each other. A prediction curve that climbs steadily while return plateaus or collapses is the overestimation signature, and it appears well before the return curve turns down.

saying these in an interview costs you the question

  • Says the two critics are averaged to reduce variance
  • Thinks the twin critics share weights or one target network only
  • Blames overestimation on off-policy data rather than the maximisation
  • Claims underestimation is equally harmful, so the minimum is unprincipled
  • Cannot name the delayed update or the target smoothing noise

context