skip to content

Sample Cost and Reliability

Why a deep RL result is hard to get and harder to trust: the price of every environment step, rewards that get gamed, and runs whose outcome flips with the random seed.

on this pageshow

explore

questions

10

In deep RL, why does a reward that is zero for hundreds of steps stall learning?

level: juniorimportance: must knowfreq 68%

answer

  1. look at the return, not the network
  2. every trajectory scores the same
  3. the estimate is exactly zero, not small
  4. noise random-walks, it does not travel
  5. no reward seen, not a late one

basics

~20 s

With no non-zero reward, every return and every value target is zero, so the policy gradient and the TD error are exactly zero. Nothing is learned slowly; there is no signal at all until the agent stumbles onto a reward.

solid answer

~50 s

A sparse reward is not a weak signal, it is an absent one. If a trajectory collects no reward, its return is zero, so a policy-gradient estimate `sum_t grad log pi(a_t|s_t) * G_t` is exactly zero for that trajectory; a baseline or a discount factor multiplies zero and still gives zero. In a value-based method every bootstrap target is zero, the value network fits the constant zero perfectly, and the greedy action becomes arbitrary. The agent learns nothing until random behaviour happens to produce the first reward, and undirected noise almost never produces a long specific sequence. Montezuma's Revenge is the standard case: descend a ladder, dodge a skull, reach a key, hundreds of correct actions before the first point. The fix is directed exploration or an easier start state, not a bigger learning rate.

go deeper

for a junior

Be ready to say plainly that with zero rewards the return is zero, so there is no gradient and no useful value target. Know one concrete example, such as needing hundreds of correct moves before the first point in Montezuma's Revenge.

for a middle

Show the mechanics: write the policy-gradient estimator and explain why every term vanishes, and explain why a value network trained on all-zero targets converges to the constant zero with an arbitrary greedy action.

for a senior

Demonstrate the diagnostic instinct. Log the fraction of episodes with a non-zero return before touching the algorithm, rule out a reward that is emitted but lost in the pipeline, and pick between directed exploration, an easier reset distribution and demonstrations based on what the environment allows.

for a principal

Own the call about whether a sparse-reward task should be attacked with RL at all. Weigh the engineering cost of building an exploration stack against restructuring the task, buying demonstrations, or using a controller for the parts that do not need learning.

## What "sparse" actually means A reward is sparse when almost every state transition returns zero and a non-zero reward appears only after a long, specific sequence of actions. This is different from a reward that is *noisy* (present but high variance) and different from a reward that is *delayed* (it does arrive, but late, and has to be attributed to the right earlier action). Sparse means most training data contains no information about what is good at all. ## Why the learning signal is exactly zero, not merely small Take a policy-gradient method. The estimator on one sampled trajectory is ``` grad J = sum_t grad log pi(a_t | s_t) * G_t ``` where `G_t` is the return from step t. If the trajectory earned no reward, every `G_t` is zero, so every term is zero and the whole estimate is zero. Discounting does not help: `gamma^k * 0 = 0`. A learned baseline does not help either, because with all returns zero the baseline fits zero and the advantage `G_t - V(s_t)` is also zero. The estimator is unbiased with zero mean *and* zero variance, which is the mathematical way of saying there is nothing in the batch to learn from. Value-based methods degenerate the same way. The TD target `r + gamma * max_a Q(s', a)` is zero when `r` is zero and `Q` is initialised near zero, so the network is trained to output zero everywhere, converges to exactly that, and stays there. Every action then has the same estimated value, so the greedy policy is decided by initialisation noise rather than by experience. A value loss that has fallen to near zero looks like healthy convergence on a dashboard and actually means the network has learned the constant function. The consequence is that the usual knobs are irrelevant. A larger learning rate multiplies a zero gradient. A deeper network fits the constant zero just as well. A better optimizer converges faster to the same nothing. The bottleneck is the data, not the optimisation. ## Why undirected exploration does not rescue it Undirected exploration adds noise to actions: a small probability of a random action, or noise added to a continuous action. This produces a random walk in state space. A random walk with per-step noise returns to where it came from far more often than it travels a long way in one direction, so the probability of executing a hundred-step specific sequence by chance is astronomically small. Doubling the exploration noise usually makes it worse, because the walk becomes even less likely to make consistent progress. Montezuma's Revenge is the canonical demonstration. To earn the first point the agent must climb down a ladder, jump a gap, avoid a moving skull and reach a key. Every one of those sub-steps has to be right, in order, before the score changes at all. An agent that only follows the extrinsic reward sees a flat zero forever. ## What to do instead First, diagnose. The most useful number on a sparse-reward run is the fraction of training episodes with a non-zero return. If it is zero, no algorithmic change matters and you should check the environment itself: is the reward actually emitted, is the terminal condition firing, is the reward being masked by a wrapper or lost at an episode boundary? Plenty of "sparse reward" bugs are simply rewards that never reach the learner. If the reward is genuinely reachable but rare, the families of fixes are: - **Directed exploration.** Give the agent an internally generated signal that rewards visiting states it has not seen, so that its behaviour is pushed outward rather than jittering locally. Count-based bonuses (roughly proportional to `1/sqrt(N(s))`) and prediction-error novelty bonuses both do this, and both decay as the state becomes familiar, so they do not permanently distort the task objective. - **Change where episodes begin.** If you can reset the agent near states from which the reward is reachable, and gradually move the reset distribution back toward the true start, the agent sees rewards immediately and the problem becomes ordinary learning. - **Bootstrap from demonstrations.** A handful of human or scripted trajectories that reach the reward gives the learner positive examples it would never have found on its own; the policy can be pre-trained on them or they can be seeded into the data the learner replays. - **Decompose the task.** Sub-goals, or a hierarchy where a high-level controller picks targets that a low-level policy can reach, shortens the sequence that must be produced by chance. ## The line a good answer draws Be explicit that this is an exploration problem, not a credit-assignment problem. If the agent does eventually reach the reward, deciding which of the earlier actions deserves the credit is a separate and solvable question. Here the agent never obtains the observation that would start that process. Confusing the two is the most common way this question is answered badly.

  • Does adding a learned baseline or lowering the discount factor change anything when all returns are zero?
    No. The baseline is fit to the returns, so with all returns zero it fits zero and the advantage `G_t - V(s_t)` is zero too. Discounting multiplies a zero return by a number less than one and still gives zero. Both are variance-reduction devices, and there is no variance to reduce when every sample carries the same value.
  • How is this different from an agent that does get a reward, but only at the very end of a long episode?
    That agent has a signal and has to spread it over the actions that earned it, which value bootstrapping and eligibility-style multi-step returns are designed to do. A sparse-reward agent has no signal at all: it has never observed the outcome that would need attributing. One is an attribution problem, the other is an exploration problem, and they call for different fixes.
  • What is the first thing you check before reaching for an exploration bonus?
    Whether any episode has ever returned a non-zero reward. Log that fraction. Very often the reward is emitted but dropped, for instance clipped away, zeroed on the terminal transition, or lost when an episode is truncated. Only after confirming the reward reaches the learner is it worth changing the algorithm.

saying these in an interview costs you the question

  • Says the gradient is small rather than exactly zero
  • Blames vanishing gradients inside the network
  • Suggests a larger learning rate or a deeper network
  • Confuses never seeing a reward with a delayed reward
  • Expects action noise to find a hundred-step sequence
  • Reads a value loss near zero as successful convergence

context

open as a page

Why does a replay-based deep RL learner need far fewer environment steps than a PPO-family run?

level: middleimportance: must knowfreq 62%

basics

~20 s

A replay-based learner stores every transition in a large buffer and resamples each one hundreds of times, so a single collected step funds many gradient updates. A PPO-family batch is passed over about ten times and then thrown away.

open as a page

Why do five runs of one deep RL configuration with different seeds end at wildly different returns?

level: middleimportance: must knowfreq 58%

basics

~20 s

A deep RL agent generates its own training data, so random exploration and initialisation change which transitions it ever sees. Bootstrapped value estimates then amplify that early divergence, and the run ends in a completely different policy.

open as a page

How does a random-network-distillation bonus give a deep RL agent a novelty signal?

level: middleimportance: should knowfreq 42%

basics

~20 s

A frozen, randomly initialised target network maps each observation to an embedding, and a predictor is trained to reproduce it on visited states. The prediction error is the intrinsic reward: large on unfamiliar observations, shrinking as they recur.

open as a page

What does clipping rewards to their sign cost a deep RL agent's learned policy?

level: seniorimportance: should knowfreq 36%

basics

~10 s

Clipping bounds value targets and stabilises training, but it changes the objective. Once every positive event is worth one, the agent maximises how many good events happen rather than how much they are worth.

open as a page

In off-policy deep RL, what breaks when you raise the update-to-data ratio to 20 gradient steps per environment step?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Efficiency improves at first, then the critic overfits the small pool of early transitions and progress stalls — the primacy bias. Usual fixes are ensembled critics with in-target minimisation, stronger regularisation, or periodically reinitialising the network while keeping the buffer.

open as a page

How should you report deep RL returns across seeds instead of showing the best run's curve?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Run about ten seeds per configuration and report an aggregate that is robust to failed runs, such as the median or interquartile mean, with stratified bootstrap confidence intervals and every per-seed curve visible. Never report the best run.

open as a page

With a fixed per-step simulator cost, how do you plan a 10M-step deep RL training budget?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat environment steps and wall-clock time as two separate budgets. Price one step, estimate steps-to-threshold from short pilot runs, reserve most of the budget for search, and pick the algorithm family by whichever resource is scarce.

open as a page

Your simulated robot policy scores by exploiting a physics-engine bug. How do you catch it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Watch the trajectories and log physical invariants such as penetration depth, contact impulse and energy, and track a success metric defined independently of the reward. A rising return inside the same buggy simulator is not evidence.

open as a page

Your RL variant beats a baseline that lacked observation normalisation — is the gain real?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Not as stated. Observation normalisation, reward scaling and advantage standardisation often move returns more than an algorithmic change does, so the comparison must give both arms the same details and the same tuning budget before any gain can be attributed to the algorithm.

open as a page