Why does the max in a DQN target overestimate action values, and how does Double DQN fix it?
answer
- max over estimates, not over truths
- expected max exceeds max of expectations
- the biggest error wins the argmax
- split selection from evaluation
- online picks the action, target scores it
basics
~20 sTaking a max over noisy Q-estimates picks whichever action's error is largest, so targets are biased upward even when true values are equal. Double DQN selects the action with the online network but scores it with the target network.
solid answer
~50 sStandard Q-learning bootstraps with `y = r + gamma * max_a' Q_target(s', a')`. That max is applied to *estimates*, not to true values, and a max over noisy numbers returns something systematically too large: whichever action happens to carry the biggest positive error wins the argmax, and that error is copied straight into the target. Even when every action in a state has the same true value and the estimation noise is zero-mean, the expected max is strictly positive. Bootstrapping then propagates the inflation backwards, and because the bias is larger where estimates are noisier, it distorts which action looks greedy rather than shifting all values equally. Double DQN splits the two jobs the max was doing: the online network *selects* the argmax action, the target network *evaluates* it, so `y = r + gamma * Q_target(s', argmax_a' Q_online(s', a'))`. Two partly decorrelated estimators rarely err upward on the same action, so the bias shrinks.
code
python · 20 linesimport random, statistics
N_ACTIONS, N_TRIALS = 4, 20000
def single_estimator():
# every action's TRUE value is 0; estimates carry zero-mean noise
est = [random.gauss(0.0, 1.0) for _ in range(N_ACTIONS)]
return max(est) # select AND evaluate with one set
def double_estimator():
a_est = [random.gauss(0.0, 1.0) for _ in range(N_ACTIONS)]
b_est = [random.gauss(0.0, 1.0) for _ in range(N_ACTIONS)]
best = max(range(N_ACTIONS), key=lambda i: a_est[i]) # select with A
return b_est[best] # evaluate with B
random.seed(0)
print(statistics.mean(single_estimator() for _ in range(N_TRIALS)))
print(statistics.mean(double_estimator() for _ in range(N_TRIALS)))
# single: about +1.03, though the true value of every action is 0
# double: about 0.0go deeper
Recall the shape of the Q-learning target and that the next-state value comes from a max over estimated action values. Be able to say that estimates carry error and that a max tends to pick the lucky one.
Explain why the expected max of noisy estimates exceeds the max of the true values, and write the Double DQN target with the online network choosing the action and the target network scoring it. Say what the standard target does differently.
Show how you would spot the bias in a real run by comparing predicted start-state values with realised discounted returns, and be honest that Double DQN only shrinks the bias because the two networks are correlated.
Argue when overestimation is worth attacking at all. Weigh it against exploration, reward scaling and data throughput, and be able to say what evidence would convince you the bias, not something else, is limiting the agent.
## The target a value network is actually fitted to One-step Q-learning updates an action-value estimate towards a bootstrapped target: `Q(s,a) <- Q(s,a) + lr * (r + gamma * max_a' Q(s',a') - Q(s,a))`, where `r` is the immediate reward, `gamma` the discount factor and `s'` the next state. In a deep Q-network the table becomes a network, and the bracketed quantity becomes a regression label `y = r + gamma * max_a' Q_target(s',a')` that the network is fitted to by minimising a squared or Huber error. Everything below is about that one `max`. ## Why a max over estimates is biased upward Write the estimate as truth plus error: `Q_hat(s',a') = q(s',a') + e(a')`, with the errors roughly zero-mean. The maximum is a convex function of its arguments, so by Jensen's inequality `E[max_a' Q_hat(s',a')] >= max_a' E[Q_hat(s',a')] = max_a' q(s',a')`. The inequality is strict as soon as the errors have spread and no single action dominates outright. The gap grows with the number of actions and with the noise scale. A concrete case makes the size visible. Take a state where every one of four actions has true value exactly zero, and give each estimate an independent standard-normal error. The expected value of the largest of four independent standard normal draws is about 1.03. So the target reports a value near `+1.03 * gamma` for a state whose true value is zero. Nothing about the noise is skewed; the damage comes from *selecting on the error and then reusing the selected number as the estimate* -- the winner's curse. ## Why it does not average out Two answers you should have ready. First, **bootstrapping compounds it**. The inflated target becomes the label for `Q(s,a)` in the predecessor state, and that state's value is itself maxed when its own predecessor is updated. Bias travels backwards along trajectories instead of cancelling. Second, **the bias is not uniform**. It scales with the estimation variance, which differs sharply across states and actions: rarely tried actions, states with high reward variance, and everything early in training are inflated most. A uniform upward shift across all actions would leave the greedy policy untouched and be harmless. A non-uniform one changes which action is greedy, so the agent commits to actions that are merely optimistic rather than good, gathers data under that policy, and reinforces the error. An inventory-replenishment agent shows the shape of the failure. It chooses among a dozen order-up-to levels under stochastic demand, so returns are noisy and the levels tried least have the widest error bars. Vanilla targets keep crowning whichever seldom-tried level happened to see a lucky week; the predicted values climb steadily while realised cost and service level stay flat. ## Double Q-learning: separate the two roles The root cause is that one estimator both *picks* the argmax and *supplies* the value for it. Double Q-learning keeps two estimators, `Q_A` and `Q_B`, trained on different samples. When updating `A`, it selects `a* = argmax_a Q_A(s',a)` and evaluates with the other one, `Q_B(s',a*)`. Because `B`'s error at `a*` is independent of the fact that `A` ranked it first, the evaluation is not inflated by the selection -- it is as likely to be low as high. ## What Double DQN changes in practice Double DQN reuses machinery a deep Q-learning setup already has rather than training a second network. The online network selects; the target network -- a periodically refreshed older copy of the online weights -- evaluates: `y = r + gamma * Q_target(s', argmax_a' Q_online(s', a'))` compared with the standard `y = r + gamma * max_a' Q_target(s', a')`. Note what the standard form really does: the target network both selects and evaluates, so the coupling is intact. The fix costs one extra forward pass of the online network on next states; no extra parameters, no extra buffer, no change to the loss. The honest caveat: the target network is a lagged copy of the online one, so the two estimators are correlated rather than independent. Double DQN therefore *reduces* overestimation substantially but does not remove it, and in some states it under-corrects. It is a cheap approximation to Double Q-learning, not the real thing. ## Detecting it in a training run You rarely have ground-truth values, but you have realised returns. Log the predicted value of episode start states and compare it with the actual discounted return those episodes achieved. A prediction curve sitting above -- and drifting further above -- the realised-return curve is the signature. Value estimates that grow steadily while episode return is flat say the same thing. ## What it does not fix It is a fix for one specific statistical bias in target construction. It does not make an unstable optimisation stable, does not compensate for badly scaled rewards, and says nothing about exploration. If the agent is not learning at all, overestimation is usually not the first thing to chase.
- If overestimation shifted every action's value up by the same amount, would it matter?A perfectly uniform shift would leave the greedy ranking unchanged and be mostly harmless. The problem is that the bias tracks estimation variance, so seldom-tried actions and high-variance states are inflated more than others. That reorders the argmax, the agent acts on the reordering, and bootstrapping carries the distortion into earlier states.
- Why does Double DQN reuse the target network instead of training two independent Q-networks?Cost and simplicity: the target network already exists, so the change is one extra forward pass and no new parameters or buffers. The price is that a lagged copy of the online weights is correlated with them, so the decorrelation is partial and some overestimation survives. True Double Q-learning with two independently trained estimators corrects more but doubles the learning machinery.
- Without ground-truth values, how would you show a running agent is overestimating?Compare predictions with outcomes on the same episodes: record the predicted value of each episode's start state, then the discounted return that episode actually collected, and plot the two together. Systematic, widening separation with predictions on top is the evidence. Predicted values climbing while episode return stays flat is the same signal in cruder form.
Run a hiring loop where every candidate is genuinely equal but each score carries random error, then always promote the top scorer: their score overstates their real ability. Double DQN re-scores that winner with a second, independent panel.
saying these in an interview costs you the question
- Says the max is unbiased because the noise has mean zero
- Claims Double DQN needs two replay buffers or two full agents
- Reverses the roles: target network selects, online network evaluates
- Believes Double DQN removes overestimation completely
- Treats overestimation as harmless because it affects all actions equally
- Proposes a smaller learning rate as the fix for a bias in the target