skip to content

Deep Reinforcement Learning

Replacing the Q-table with a network buys generalization and costs stability: replay buffers, target networks, actor-critic and PPO. Interviewers probe where the tabular guarantees break.

on this pageshow

explore

questions

page 1 of 2

In deep RL, why does a reward that is zero for hundreds of steps stall learning?

level: juniorimportance: must knowfreq 68%

answer

  1. look at the return, not the network
  2. every trajectory scores the same
  3. the estimate is exactly zero, not small
  4. noise random-walks, it does not travel
  5. no reward seen, not a late one

basics

~20 s

With no non-zero reward, every return and every value target is zero, so the policy gradient and the TD error are exactly zero. Nothing is learned slowly; there is no signal at all until the agent stumbles onto a reward.

solid answer

~50 s

A sparse reward is not a weak signal, it is an absent one. If a trajectory collects no reward, its return is zero, so a policy-gradient estimate `sum_t grad log pi(a_t|s_t) * G_t` is exactly zero for that trajectory; a baseline or a discount factor multiplies zero and still gives zero. In a value-based method every bootstrap target is zero, the value network fits the constant zero perfectly, and the greedy action becomes arbitrary. The agent learns nothing until random behaviour happens to produce the first reward, and undirected noise almost never produces a long specific sequence. Montezuma's Revenge is the standard case: descend a ladder, dodge a skull, reach a key, hundreds of correct actions before the first point. The fix is directed exploration or an easier start state, not a bigger learning rate.

go deeper

for a junior

Be ready to say plainly that with zero rewards the return is zero, so there is no gradient and no useful value target. Know one concrete example, such as needing hundreds of correct moves before the first point in Montezuma's Revenge.

for a middle

Show the mechanics: write the policy-gradient estimator and explain why every term vanishes, and explain why a value network trained on all-zero targets converges to the constant zero with an arbitrary greedy action.

for a senior

Demonstrate the diagnostic instinct. Log the fraction of episodes with a non-zero return before touching the algorithm, rule out a reward that is emitted but lost in the pipeline, and pick between directed exploration, an easier reset distribution and demonstrations based on what the environment allows.

for a principal

Own the call about whether a sparse-reward task should be attacked with RL at all. Weigh the engineering cost of building an exploration stack against restructuring the task, buying demonstrations, or using a controller for the parts that do not need learning.

## What "sparse" actually means A reward is sparse when almost every state transition returns zero and a non-zero reward appears only after a long, specific sequence of actions. This is different from a reward that is *noisy* (present but high variance) and different from a reward that is *delayed* (it does arrive, but late, and has to be attributed to the right earlier action). Sparse means most training data contains no information about what is good at all. ## Why the learning signal is exactly zero, not merely small Take a policy-gradient method. The estimator on one sampled trajectory is ``` grad J = sum_t grad log pi(a_t | s_t) * G_t ``` where `G_t` is the return from step t. If the trajectory earned no reward, every `G_t` is zero, so every term is zero and the whole estimate is zero. Discounting does not help: `gamma^k * 0 = 0`. A learned baseline does not help either, because with all returns zero the baseline fits zero and the advantage `G_t - V(s_t)` is also zero. The estimator is unbiased with zero mean *and* zero variance, which is the mathematical way of saying there is nothing in the batch to learn from. Value-based methods degenerate the same way. The TD target `r + gamma * max_a Q(s', a)` is zero when `r` is zero and `Q` is initialised near zero, so the network is trained to output zero everywhere, converges to exactly that, and stays there. Every action then has the same estimated value, so the greedy policy is decided by initialisation noise rather than by experience. A value loss that has fallen to near zero looks like healthy convergence on a dashboard and actually means the network has learned the constant function. The consequence is that the usual knobs are irrelevant. A larger learning rate multiplies a zero gradient. A deeper network fits the constant zero just as well. A better optimizer converges faster to the same nothing. The bottleneck is the data, not the optimisation. ## Why undirected exploration does not rescue it Undirected exploration adds noise to actions: a small probability of a random action, or noise added to a continuous action. This produces a random walk in state space. A random walk with per-step noise returns to where it came from far more often than it travels a long way in one direction, so the probability of executing a hundred-step specific sequence by chance is astronomically small. Doubling the exploration noise usually makes it worse, because the walk becomes even less likely to make consistent progress. Montezuma's Revenge is the canonical demonstration. To earn the first point the agent must climb down a ladder, jump a gap, avoid a moving skull and reach a key. Every one of those sub-steps has to be right, in order, before the score changes at all. An agent that only follows the extrinsic reward sees a flat zero forever. ## What to do instead First, diagnose. The most useful number on a sparse-reward run is the fraction of training episodes with a non-zero return. If it is zero, no algorithmic change matters and you should check the environment itself: is the reward actually emitted, is the terminal condition firing, is the reward being masked by a wrapper or lost at an episode boundary? Plenty of "sparse reward" bugs are simply rewards that never reach the learner. If the reward is genuinely reachable but rare, the families of fixes are: - **Directed exploration.** Give the agent an internally generated signal that rewards visiting states it has not seen, so that its behaviour is pushed outward rather than jittering locally. Count-based bonuses (roughly proportional to `1/sqrt(N(s))`) and prediction-error novelty bonuses both do this, and both decay as the state becomes familiar, so they do not permanently distort the task objective. - **Change where episodes begin.** If you can reset the agent near states from which the reward is reachable, and gradually move the reset distribution back toward the true start, the agent sees rewards immediately and the problem becomes ordinary learning. - **Bootstrap from demonstrations.** A handful of human or scripted trajectories that reach the reward gives the learner positive examples it would never have found on its own; the policy can be pre-trained on them or they can be seeded into the data the learner replays. - **Decompose the task.** Sub-goals, or a hierarchy where a high-level controller picks targets that a low-level policy can reach, shortens the sequence that must be produced by chance. ## The line a good answer draws Be explicit that this is an exploration problem, not a credit-assignment problem. If the agent does eventually reach the reward, deciding which of the earlier actions deserves the credit is a separate and solvable question. Here the agent never obtains the observation that would start that process. Confusing the two is the most common way this question is answered badly.

  • Does adding a learned baseline or lowering the discount factor change anything when all returns are zero?
    No. The baseline is fit to the returns, so with all returns zero it fits zero and the advantage `G_t - V(s_t)` is zero too. Discounting multiplies a zero return by a number less than one and still gives zero. Both are variance-reduction devices, and there is no variance to reduce when every sample carries the same value.
  • How is this different from an agent that does get a reward, but only at the very end of a long episode?
    That agent has a signal and has to spread it over the actions that earned it, which value bootstrapping and eligibility-style multi-step returns are designed to do. A sparse-reward agent has no signal at all: it has never observed the outcome that would need attributing. One is an attribution problem, the other is an exploration problem, and they call for different fixes.
  • What is the first thing you check before reaching for an exploration bonus?
    Whether any episode has ever returned a non-zero reward. Log that fraction. Very often the reward is emitted but dropped, for instance clipped away, zeroed on the terminal transition, or lost when an episode is truncated. Only after confirming the reward reaches the learner is it worth changing the algorithm.

saying these in an interview costs you the question

  • Says the gradient is small rather than exactly zero
  • Blames vanishing gradients inside the network
  • Suggests a larger learning rate or a deeper network
  • Confuses never seeing a reward with a delayed reward
  • Expects action noise to find a hundred-step sequence
  • Reads a value loss near zero as successful convergence

context

open as a page

Why does a deep Q-network train from a replay buffer instead of the transitions as they arrive?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Consecutive transitions are almost identical, so learning from them in order gives correlated, unstable updates. A replay buffer stores past transitions and samples shuffled batches from them, breaking that correlation and letting each transition be reused many times.

open as a page

In an advantage actor-critic, what does the critic's learned value function contribute to the policy update?

level: middleimportance: must knowfreq 70%

basics

~20 s

The critic learns a state-value estimate V(s). The actor is updated with the advantage - reward plus gamma times the next state's value, minus the current state's value - a signal centred on zero that reinforces only actions better than average.

open as a page

How does PPO's clipped surrogate objective bound how far one update moves the policy?

level: middleimportance: must knowfreq 76%

basics

~20 s

PPO maximises min(r*A, clip(r, 1-eps, 1+eps)*A), where r is the new-over-old action probability ratio and A the advantage. Once r passes the bound in the direction the advantage favours, the objective flattens and that sample stops pushing.

open as a page

How does DDPG's deterministic actor get its gradient from the critic?

level: middleimportance: must knowfreq 62%

basics

~20 s

DDPG's critic scores a state-action pair and is differentiable in the action, so the actor is trained by pushing its output uphill on that score: the gradient is dQ/da evaluated at the actor's own action, chained with the actor's parameter gradient.

open as a page

Why does a replay-based deep RL learner need far fewer environment steps than a PPO-family run?

level: middleimportance: must knowfreq 62%

basics

~20 s

A replay-based learner stores every transition in a large buffer and resamples each one hundreds of times, so a single collected step funds many gradient updates. A PPO-family batch is passed over about ten times and then thrown away.

open as a page

Why do five runs of one deep RL configuration with different seeds end at wildly different returns?

level: middleimportance: must knowfreq 58%

basics

~20 s

A deep RL agent generates its own training data, so random exploration and initialisation change which transitions it ever sees. Bootstrapped value estimates then amplify that early divergence, and the run ends in a completely different policy.

open as a page

Why replace a Q-table with a Q-network when the state space is continuous or huge?

level: middleimportance: must knowfreq 72%

basics

~20 s

A table needs one independent cell per state-action pair, so a continuous state must be binned into a count that explodes with dimensions and stays mostly unvisited. A network shares weights, so one update generalises to similar states.

open as a page

Why is the deadly triad of approximation, bootstrapping and off-policy updates unstable?

level: middleimportance: must knowfreq 64%

basics

~20 s

Each ingredient is safe alone. Together, an approximator's update for one state moves the very targets it is fitting, and off-policy data weights those updates by a distribution the approximation was not fitted under, so the estimates can amplify instead of contract.

open as a page

Why does the max in a DQN target overestimate action values, and how does Double DQN fix it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Taking a max over noisy Q-estimates picks whichever action's error is largest, so targets are biased upward even when true values are equal. Double DQN selects the action with the online network but scores it with the target network.

open as a page

In DQN, why is the regression target computed by a separate, slowly updated copy of the network?

level: middleimportance: must knowfreq 75%

basics

~20 s

The target network is a frozen copy of the Q-network used only to compute the regression target. Without it, every update that raises the prediction also raises the target, so the network chases a label it is moving itself.

open as a page

Why does TD3 train two critics and take the smaller of their target values?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Because a deterministic actor is a dedicated maximiser of its critic, any upward error in the value estimate is actively sought out and then bootstrapped into future targets. Two independently initialised critics rarely err upward in the same place, so taking the smaller target cancels most of that optimism.

open as a page

Why add an entropy bonus to the policy loss when training an actor-critic agent?

level: middleimportance: should knowfreq 55%

basics

~20 s

An entropy bonus rewards a spread-out action distribution, so the policy does not collapse onto one action before the alternatives have been tried. Its coefficient trades exploration against exploitation and is usually small and decayed toward zero.

open as a page

Why does PPO run several gradient epochs over the same batch of trajectories?

level: middleimportance: should knowfreq 54%

basics

~20 s

The importance ratio corrects the objective for data gathered by the previous policy, so one batch stays usable for a few passes. Reuse amortises expensive environment interaction, and the clip keeps that correction trustworthy as the policy drifts.

open as a page

How does SAC's maximum-entropy objective differ from maximising return alone?

level: middleimportance: should knowfreq 50%

basics

~20 s

SAC maximises expected return plus a temperature times the policy's entropy at every visited state, so the optimal policy is stochastic rather than a single best action. The entropy term also enters the critic's bootstrap target, so the values learned are entropy-augmented, not ordinary returns.

open as a page

How does a random-network-distillation bonus give a deep RL agent a novelty signal?

level: middleimportance: should knowfreq 42%

basics

~20 s

A frozen, randomly initialised target network maps each observation to an embedding, and a predictor is trained to reproduce it on visited states. The prediction error is the intrinsic reward: large on unfamiliar observations, shrinking as they recur.

open as a page

Why does a dueling DQN split Q into value and advantage streams, and how are they recombined?

level: middleimportance: should knowfreq 45%

basics

~20 s

A dueling network learns one state value plus per-action advantages, so the state's worth is learned from every transition whatever action was taken. The streams recombine as value plus advantage minus the mean advantage, resolving an ambiguous split.

open as a page

In generalized advantage estimation, what changes as lambda moves from 1.0 down toward 0?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Lambda sets how heavily the advantage estimate leans on the critic. Near 1 it sums many observed rewards: low bias, high variance. Near 0 it is one bootstrapped residual: low variance, but the critic's error passes straight through.

open as a page

What does clipping rewards to their sign cost a deep RL agent's learned policy?

level: seniorimportance: should knowfreq 36%

basics

~10 s

Clipping bounds value targets and stabilises training, but it changes the objective. Once every positive event is worth one, the agent maximises how many good events happen rather than how much they are worth.

open as a page

In off-policy deep RL, what breaks when you raise the update-to-data ratio to 20 gradient steps per environment step?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Efficiency improves at first, then the critic overfits the small pool of early transitions and progress stalls — the primacy bias. Usual fixes are ensembled critics with in-target minimisation, stronger regularisation, or periodically reinitialising the network while keeping the buffer.

open as a page

How should you report deep RL returns across seeds instead of showing the best run's curve?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Run about ten seeds per configuration and report an aggregate that is robust to failed runs, such as the median or interquartile mean, with stratified bootstrap confidence intervals and every per-seed curve visible. Never report the best run.

open as a page

Why can a Q-network's performance on already-mastered states regress as training continues?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Shared weights mean fitting one region of the state space silently moves the values of other regions. Because an agent's own improving policy keeps shifting which states it visits, the regions it stopped visiting are overwritten rather than rehearsed.

open as a page

How does prioritized experience replay choose transitions, and why does it need importance-sampling weights?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Prioritized replay samples stored transitions in proportion to a power of their last temporal-difference error, so surprising transitions are revisited more often. That skews the sample distribution, so each update is multiplied by an importance-sampling weight that undoes the skew.

open as a page

How do you size a replay buffer when its old transitions come from a policy and a world that have both changed?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Size the buffer by how fast the world and the policy change, not by available memory. Data from an older policy in an unchanged environment is merely off-distribution; data from an environment that has since changed is simply wrong and should be evicted.

open as a page

When would you share one encoder trunk between the actor and critic instead of training two networks?

level: principalimportance: should knowfreq 38%

basics

~20 s

Share a trunk when observations are high-dimensional and the encoder dominates the cost: the value head's dense signal also shapes features early. Keep them separate when the state is small or the value loss swamps the policy gradient.

open as a page

When does PPO's clip alone suffice, and when do you add explicit KL control?

level: principalimportance: should knowfreq 36%

basics

~20 s

The clip bounds per-sample ratios, not the overall policy shift, and gives no restoring force once a ratio is already outside its band. Add measured-divergence control when a wrecked policy is expensive or the environment cannot be re-run cheaply.

open as a page

For a bounded physical plant, would you deploy a deterministic actor with injected exploration noise or a maximum-entropy stochastic policy?

level: principalimportance: should knowfreq 40%

basics

~20 s

Prefer the entropy-regularised stochastic learner for training, because its exploration is state-dependent and self-tuning, then ship its median action behind an external rate limiter. Choose the deterministic actor only when a fixed, auditable control law and a reproducible exploration process matter more.

open as a page

With a fixed per-step simulator cost, how do you plan a 10M-step deep RL training budget?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat environment steps and wall-clock time as two separate budgets. Price one step, estimate steps-to-threshold from short pilot runs, reserve most of the budget for search, and pick the algorithm family by whichever resource is scarce.

open as a page

How does TRPO's hard KL constraint differ from PPO's clipped objective?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

TRPO maximises the surrogate subject to an explicit constraint that average KL between old and new policy stays under a threshold, solved by a natural-gradient step plus a line search. PPO approximates that idea first-order, with clipping.

open as a page

Why must a tanh-squashed continuous action correct its log-probability?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Squashing is a nonlinear change of variables, so it compresses probability mass near the bounds. The density of the squashed action equals the pre-squash density divided by the derivative 1 - tanh(u)^2, so the log-probability must have log(1 - tanh(u)^2) subtracted per action coordinate.

open as a page

showing 1–30 of 34