How do you choose n in n-step TD returns when the reward arrives far in the future?
answer
- a dial between two extremes
- how far back does one update reach
- more real reward, less estimate
- match it to the delay in the problem
- or average over all horizons instead
basics
~20 sPick n on the delay between action and reward. Small n gives low-variance but myopic targets that propagate credit one state per pass; large n propagates credit fast but carries the noise of a long trajectory. Tune n empirically; intermediate values usually beat both extremes.
solid answer
~50 sAn n-step return uses `n` real rewards and then bootstraps: `G(n) = r1 + gamma*r2 + ... + gamma^(n-1)*rn + gamma^n * V(s_n)`. With `n = 1` it is TD(0); with `n` running to the end of the episode it is Monte-Carlo. Two things move with `n`. Bias falls and variance rises, because more of the target is real reward and less is estimate. Less obviously, **credit propagates faster**: one-step TD moves reward information back only one state per pass, so a signal seven steps away takes many passes to reach the start. Consider a product-onboarding funnel where the activation reward only lands on day seven — one-step bootstrapping is too myopic to connect day-one behaviour to it in reasonable time, while a full return drowns the signal in a week of unrelated noise. Choose `n` near the delay, sweep it, and remember it interacts with the step size.
go deeper
Be ready to say that n-step returns sit between one-step TD and Monte-Carlo, using n real rewards before bootstrapping on the value of the state reached.
Explain both directions of the dial: more real reward means less bias and more variance, and a larger n carries credit further back in a single update.
An interviewer expects a tuning story. Anchor n in the actual delay between action and payoff, sweep it against learning curves, and say how it interacts with the step size and update lag.
Own the alternative framings: whether to blend horizons instead of picking one, whether to redefine the time step so the credit gap shrinks, and when a weak state representation is the real problem.
## The family between TD and Monte-Carlo One-step TD and Monte-Carlo are the two ends of one continuum, not two rival algorithms. The **n-step return** is the general member: ``` G(n) = r_1 + gamma*r_2 + ... + gamma^(n-1)*r_n + gamma^n * V(s_n) V(s) <- V(s) + alpha * [G(n) - V(s)] ``` Take `n = 1` and the bootstrap happens immediately: that is TD(0). Let `n` run to the terminal state and there is no bootstrap term at all: that is Monte-Carlo. Everything in between mixes `n` sampled rewards with one bootstrapped tail. Note the practical cost: an n-step update cannot be applied until `n` steps have actually occurred, so the method lags the agent by `n` transitions and needs a short buffer of recent experience. That is a real implementation consequence, though a mild one. ## The two things n controls **Bias and variance.** As `n` grows, more of the target is genuine sampled reward and less of it is the possibly-wrong estimate `V(s_n)`, so bias falls. At the same time the target accumulates `n` steps of randomness, so variance rises. This is the same tradeoff as the Monte-Carlo versus TD contrast, now available as a dial rather than a binary choice. **Credit-assignment speed.** This is the part people miss, and it is not the same thing as variance. With one-step updates, a reward observed at the end of a chain only improves the value of the state immediately before it. Its predecessor learns nothing until the next pass, when it bootstraps off the newly-improved estimate. Information walks backwards one state per pass. With n-step returns it jumps back `n` states at once. When the payoff is far from the behaviour that caused it, this determines whether the agent learns in tens or thousands of episodes. ## Choosing n on a delayed-reward problem Take a product-onboarding funnel where the reward that matters — activation — only arrives on day seven, while the intermediate days emit little or nothing. - **n = 1** is too myopic. Day-one decisions get a target of roughly `gamma * V(day two)`, which early in training is near zero and carries no information about activation at all. The signal does eventually crawl backwards, but it takes as many passes as there are steps. - **Full returns** connect day one to activation immediately, but the target now includes seven days of unrelated variation — traffic mix, weekday effects, whatever else emitted reward — so the estimate is unbiased and nearly useless per sample. - **n around the delay** (here, roughly the number of steps to the payoff) reaches the informative reward in one update while summing far less noise than the full trajectory. The honest procedure is: start from the structural delay between action and consequence, sweep `n` over a small grid around it, and compare learning curves. Empirically the curve is U-shaped in `n` — intermediate values beat both extremes — and the optimum interacts with the step size, because a larger `n` means noisier targets and typically wants a smaller `alpha`. Tune the pair, not `n` alone. ## When a single n is the wrong shape Committing to one horizon is itself a modelling choice, and it can be avoided. A **λ-return** averages all n-step returns with geometrically decaying weights controlled by a parameter `λ` in `[0, 1]`, giving a smooth blend rather than a hard cutoff: `λ = 0` recovers one-step TD, `λ = 1` recovers Monte-Carlo. Implemented with eligibility traces, this achieves the blend online, updating every recently visited state in proportion to how recently it was visited — which is a direct answer to slow credit assignment without picking a single `n`. Other levers exist and should be named as alternatives rather than substitutes: - **Restructure the problem** so the horizon shrinks — coarser time steps, or defining a step as a meaningful decision rather than a fixed interval. - **Improve the state representation** so intermediate states genuinely predict the eventual payoff, which is what makes short bootstrapping work in the first place. ## What a good answer sounds like Name both effects of `n` — the bias-variance dial and the credit-propagation speed — anchor the starting value in the actual delay of the problem, admit that it is tuned empirically alongside the step size, and mention that averaging over horizons is available if you would rather not commit to one. Claiming a universally best `n` is the giveaway that the tradeoff has not been understood.
- Besides bias and variance, what practical cost does a larger n impose?The update lags. An n-step return cannot be computed until `n` transitions have actually happened, so the learner trails the agent by `n` steps and must buffer that much recent experience. On a long horizon that delays every correction and slightly complicates the implementation, which is one reason very large `n` is rarely worth it once credit already propagates fast enough.
- How does a λ-return differ from committing to a single n?A λ-return averages every n-step return with geometrically decaying weights set by `λ` in `[0, 1]`, rather than cutting off at one horizon. `λ = 0` recovers one-step TD and `λ = 1` recovers Monte-Carlo. Blending horizons is usually more robust than picking one, and eligibility traces implement it online by updating recently visited states in proportion to recency.
- Would increasing n fix a state representation that fails to predict the eventual outcome?Only partially, and it treats the symptom. A larger `n` relies less on the bootstrapped estimate, so a bad value function hurts less — but the target still bootstraps at step `n`, and the extra variance is a real price. If intermediate states carry no signal about the payoff, the durable fix is a representation that does, or a coarser step definition that shortens the horizon.
saying these in an interview costs you the question
- Claims a single best n exists for all problems
- Thinks larger n always increases bias
- Ignores that one-step credit moves back one state per pass
- Treats n-step returns as unrelated to Monte-Carlo
- Tunes n without touching the step size