skip to content

Why is a Monte-Carlo return unbiased while a one-step TD target is biased?

level: middleimportance: should knowfreq 55%

answer

  1. what randomness does each target contain
  2. one step of noise versus a whole episode
  3. the target that reads an estimate
  4. unbiased but jittery, biased but steady
  5. the bias shrinks as estimates improve

basics

~20 s

A first-visit Monte-Carlo return is an actual sampled return, so on average it equals the true value — unbiased, but noisy because it accumulates every random reward to the end of the episode. A TD target uses the current estimate of the next state, which is wrong early on, so it is biased but far lower variance.

solid answer

~50 s

The two methods differ in what their target is made of. The Monte-Carlo target is the realised return `G`, a genuine sample of the quantity being estimated, so a first-visit estimate is unbiased. But `G` accumulates every random action, transition and reward from that state to the end of the episode, so its variance grows with episode length. The TD target `r + gamma * V(s')` contains exactly one random reward and one random transition, so its variance is small — but `V(s')` is the agent's current estimate, not the true value, so the target is systematically off until the estimates improve. That is the bias. In practice TD usually learns faster on long episodes because variance, not bias, is what dominates early; Monte-Carlo is more robust when the value estimates are badly initialised or the state representation is poor, since it never trusts them.

go deeper

for a junior

Be ready to state which target is noisy and which is systematically off, and why: one sums a whole episode of randomness, the other leans on a current estimate of the next state.

for a middle

Explain the mechanism on both sides — what random quantities each target contains, and why bootstrapping on an imperfect estimate produces a bias that shrinks as learning proceeds.

for a senior

Show judgment about when the usual TD default is wrong: poor state representations, non-Markov environments and function approximators whose errors are structural, where bootstrapped error propagates to every predecessor state.

for a principal

Own the argument about which error a team should be willing to carry — a transient, self-correcting bias that buys fast learning, versus an unbiased signal so noisy that the project runs out of interaction budget before it converges.

## The two targets Both methods update the same way — nudge the estimate toward a target — and differ only in the target. **Monte-Carlo:** wait until the episode ends, compute the realised discounted return from state `s`, ``` G = r1 + gamma*r2 + gamma^2*r3 + ... + gamma^(T-1)*rT V(s) <- V(s) + alpha * [G - V(s)] ``` **Temporal-difference:** after one step, ``` V(s) <- V(s) + alpha * [r + gamma * V(s') - V(s)] ``` ## Why Monte-Carlo is unbiased The true value of `s` under the current behaviour is *defined* as the expected return from `s`. The realised return `G` is one draw from that distribution. Averaging draws of a random variable converges on its expectation, so with first-visit Monte-Carlo — where each episode contributes at most one return per state — the estimate is unbiased for every sample size. A precise footnote worth knowing: **every-visit** Monte-Carlo, which averages returns from all visits to a state within an episode, is *biased* for finite data because returns from the same episode are correlated. It is still consistent, converging to the true value as data grows. Interviewers occasionally check whether you say "Monte-Carlo is unbiased" without qualification. ## Why Monte-Carlo has high variance `G` is the sum of everything that happened after `s`. Every stochastic action choice, every stochastic transition, every noisy reward between `s` and the terminal state contributes randomness. On a long episode, two visits to the same state can produce wildly different returns purely by luck downstream. That noise goes straight into the update, so estimates jitter and learning is slow — you need many episodes to average it out. ## Why TD has low variance The TD target depends on exactly two random things: the immediate reward `r` and the sampled next state `s'`. Everything beyond that is summarised by the deterministic lookup `V(s')`. One step of randomness instead of hundreds means dramatically less noise per update. ## Why TD is biased `V(s')` is an estimate, and early in training it is simply wrong — often initialised to zero. The target `r + gamma * V(s')` therefore does not have the true value as its expectation; its expectation is the true one-step-lookahead of a *wrong* value function. That systematic offset is the bias. The bias is **self-correcting rather than permanent**. As the estimates improve, the targets improve, and in the tabular case with a fixed policy and a suitable step-size schedule TD(0) converges to the true value function. The bias shrinks to zero; it is a transient cost, not a ceiling. It becomes a real hazard when the value function is approximated rather than tabulated, because then errors in `V(s')` can be structural — the representation simply cannot express the right value — and bootstrapping propagates that structural error to every predecessor state. ## Which one wins in practice Usually TD, and the reason is that variance dominates early learning. A high-variance unbiased target and a low-variance slightly-wrong target both point roughly in the right direction, but the noisy one requires far more samples to be trusted. Add TD's other advantages — updates after one step, no need to wait for episodes to end, works in continuing tasks — and TD is the default for prediction from streaming experience. Monte-Carlo earns its place when: - **the value estimates cannot be trusted** — poor initialisation, a bad state representation, or a function approximator whose errors are structural. Monte-Carlo never reads `V(s')`, so it cannot inherit its errors; - **the environment is not Markov** — the value of a state genuinely depends on history the state does not encode. TD's whole argument rests on the successor state summarising the future, so a non-Markov state breaks it; Monte-Carlo, which only sums actual reward, degrades far more gracefully; - **episodes are short**, in which case the variance penalty is small and the guarantee of unbiasedness is cheap. There is one more distinction worth carrying, because it explains their different failure modes. Given a fixed finite batch of experience, batch Monte-Carlo converges to the estimate with the smallest squared error on the observed returns, while batch TD(0) converges to the value function of the maximum-likelihood Markov model implied by that data — the *certainty-equivalence* estimate. TD is exploiting the Markov structure of the problem; Monte-Carlo is fitting the data it saw. That is exactly why TD wins when the Markov assumption holds and loses when it does not. ## Stating it correctly The crisp form is: Monte-Carlo targets are **unbiased, high variance**; TD targets are **biased, low variance**. Getting the direction backwards is a first-class error, and the reason is worth saying out loud — the bias comes from bootstrapping on an estimate, the variance from summing a whole trajectory of randomness.

  • If TD is biased, why is it usually the faster learner?
    Because early learning is limited by variance, not bias. A Monte-Carlo return carries the noise of an entire trajectory, so many episodes are needed before the average is trustworthy. TD's target carries one step of noise, so each update is informative, and its bias shrinks on its own as the estimates it bootstraps from improve. On long episodes that tradeoff is heavily in TD's favour.
  • When would you deliberately prefer Monte-Carlo over TD for value prediction?
    When the state representation is poor or the environment is not really Markov, so the successor state does not summarise the future. TD's target trusts `V(s')`, and if that estimate is structurally wrong the error propagates to every predecessor. Monte-Carlo never reads another estimate, so it degrades gracefully. Short episodes also make the variance penalty cheap enough to accept.
  • Is every-visit Monte-Carlo also unbiased?
    No. Averaging returns from multiple visits to the same state within one episode uses correlated samples, which makes the estimator biased for finite data. It remains consistent, converging to the true value as the number of episodes grows, and in practice performs comparably. The unqualified claim that Monte-Carlo is unbiased is strictly true only for the first-visit variant.

saying these in an interview costs you the question

  • Says TD is unbiased and Monte-Carlo biased
  • Claims TD variance is high because it updates often
  • Thinks TD bias never goes away in the tabular case
  • Cannot say where the Monte-Carlo variance comes from
  • Asserts Monte-Carlo is always unbiased without the first-visit caveat

context