In deep RL, why does a reward that is zero for hundreds of steps stall learning?
answer
- look at the return, not the network
- every trajectory scores the same
- the estimate is exactly zero, not small
- noise random-walks, it does not travel
- no reward seen, not a late one
basics
~20 sWith no non-zero reward, every return and every value target is zero, so the policy gradient and the TD error are exactly zero. Nothing is learned slowly; there is no signal at all until the agent stumbles onto a reward.
solid answer
~50 sA sparse reward is not a weak signal, it is an absent one. If a trajectory collects no reward, its return is zero, so a policy-gradient estimate `sum_t grad log pi(a_t|s_t) * G_t` is exactly zero for that trajectory; a baseline or a discount factor multiplies zero and still gives zero. In a value-based method every bootstrap target is zero, the value network fits the constant zero perfectly, and the greedy action becomes arbitrary. The agent learns nothing until random behaviour happens to produce the first reward, and undirected noise almost never produces a long specific sequence. Montezuma's Revenge is the standard case: descend a ladder, dodge a skull, reach a key, hundreds of correct actions before the first point. The fix is directed exploration or an easier start state, not a bigger learning rate.
go deeper
Be ready to say plainly that with zero rewards the return is zero, so there is no gradient and no useful value target. Know one concrete example, such as needing hundreds of correct moves before the first point in Montezuma's Revenge.
Show the mechanics: write the policy-gradient estimator and explain why every term vanishes, and explain why a value network trained on all-zero targets converges to the constant zero with an arbitrary greedy action.
Demonstrate the diagnostic instinct. Log the fraction of episodes with a non-zero return before touching the algorithm, rule out a reward that is emitted but lost in the pipeline, and pick between directed exploration, an easier reset distribution and demonstrations based on what the environment allows.
Own the call about whether a sparse-reward task should be attacked with RL at all. Weigh the engineering cost of building an exploration stack against restructuring the task, buying demonstrations, or using a controller for the parts that do not need learning.
## What "sparse" actually means A reward is sparse when almost every state transition returns zero and a non-zero reward appears only after a long, specific sequence of actions. This is different from a reward that is *noisy* (present but high variance) and different from a reward that is *delayed* (it does arrive, but late, and has to be attributed to the right earlier action). Sparse means most training data contains no information about what is good at all. ## Why the learning signal is exactly zero, not merely small Take a policy-gradient method. The estimator on one sampled trajectory is ``` grad J = sum_t grad log pi(a_t | s_t) * G_t ``` where `G_t` is the return from step t. If the trajectory earned no reward, every `G_t` is zero, so every term is zero and the whole estimate is zero. Discounting does not help: `gamma^k * 0 = 0`. A learned baseline does not help either, because with all returns zero the baseline fits zero and the advantage `G_t - V(s_t)` is also zero. The estimator is unbiased with zero mean *and* zero variance, which is the mathematical way of saying there is nothing in the batch to learn from. Value-based methods degenerate the same way. The TD target `r + gamma * max_a Q(s', a)` is zero when `r` is zero and `Q` is initialised near zero, so the network is trained to output zero everywhere, converges to exactly that, and stays there. Every action then has the same estimated value, so the greedy policy is decided by initialisation noise rather than by experience. A value loss that has fallen to near zero looks like healthy convergence on a dashboard and actually means the network has learned the constant function. The consequence is that the usual knobs are irrelevant. A larger learning rate multiplies a zero gradient. A deeper network fits the constant zero just as well. A better optimizer converges faster to the same nothing. The bottleneck is the data, not the optimisation. ## Why undirected exploration does not rescue it Undirected exploration adds noise to actions: a small probability of a random action, or noise added to a continuous action. This produces a random walk in state space. A random walk with per-step noise returns to where it came from far more often than it travels a long way in one direction, so the probability of executing a hundred-step specific sequence by chance is astronomically small. Doubling the exploration noise usually makes it worse, because the walk becomes even less likely to make consistent progress. Montezuma's Revenge is the canonical demonstration. To earn the first point the agent must climb down a ladder, jump a gap, avoid a moving skull and reach a key. Every one of those sub-steps has to be right, in order, before the score changes at all. An agent that only follows the extrinsic reward sees a flat zero forever. ## What to do instead First, diagnose. The most useful number on a sparse-reward run is the fraction of training episodes with a non-zero return. If it is zero, no algorithmic change matters and you should check the environment itself: is the reward actually emitted, is the terminal condition firing, is the reward being masked by a wrapper or lost at an episode boundary? Plenty of "sparse reward" bugs are simply rewards that never reach the learner. If the reward is genuinely reachable but rare, the families of fixes are: - **Directed exploration.** Give the agent an internally generated signal that rewards visiting states it has not seen, so that its behaviour is pushed outward rather than jittering locally. Count-based bonuses (roughly proportional to `1/sqrt(N(s))`) and prediction-error novelty bonuses both do this, and both decay as the state becomes familiar, so they do not permanently distort the task objective. - **Change where episodes begin.** If you can reset the agent near states from which the reward is reachable, and gradually move the reset distribution back toward the true start, the agent sees rewards immediately and the problem becomes ordinary learning. - **Bootstrap from demonstrations.** A handful of human or scripted trajectories that reach the reward gives the learner positive examples it would never have found on its own; the policy can be pre-trained on them or they can be seeded into the data the learner replays. - **Decompose the task.** Sub-goals, or a hierarchy where a high-level controller picks targets that a low-level policy can reach, shortens the sequence that must be produced by chance. ## The line a good answer draws Be explicit that this is an exploration problem, not a credit-assignment problem. If the agent does eventually reach the reward, deciding which of the earlier actions deserves the credit is a separate and solvable question. Here the agent never obtains the observation that would start that process. Confusing the two is the most common way this question is answered badly.
- Does adding a learned baseline or lowering the discount factor change anything when all returns are zero?No. The baseline is fit to the returns, so with all returns zero it fits zero and the advantage `G_t - V(s_t)` is zero too. Discounting multiplies a zero return by a number less than one and still gives zero. Both are variance-reduction devices, and there is no variance to reduce when every sample carries the same value.
- How is this different from an agent that does get a reward, but only at the very end of a long episode?That agent has a signal and has to spread it over the actions that earned it, which value bootstrapping and eligibility-style multi-step returns are designed to do. A sparse-reward agent has no signal at all: it has never observed the outcome that would need attributing. One is an attribution problem, the other is an exploration problem, and they call for different fixes.
- What is the first thing you check before reaching for an exploration bonus?Whether any episode has ever returned a non-zero reward. Log that fraction. Very often the reward is emitted but dropped, for instance clipped away, zeroed on the terminal transition, or lost when an episode is truncated. Only after confirming the reward reaches the learner is it worth changing the algorithm.
saying these in an interview costs you the question
- Says the gradient is small rather than exactly zero
- Blames vanishing gradients inside the network
- Suggests a larger learning rate or a deeper network
- Confuses never seeing a reward with a delayed reward
- Expects action noise to find a hundred-step sequence
- Reads a value loss near zero as successful convergence