What goes wrong when a Q-learning agent's epsilon decays to near zero after 1% of training?
answer
- commits before the estimates mean anything
- unvisited states keep their initial values
- the policy chooses its own state coverage
- convergence needs every pair visited forever
- decay to a floor, not to zero
basics
~20 sThe agent commits to whatever policy looked best while its value estimates were still mostly noise. States it stopped visiting keep their initial values, so the learning curve flattens early on a suboptimal policy that more exploration would have escaped.
solid answer
~50 sThis is premature exploitation. Early on, the action values are dominated by initialisation and a handful of noisy returns, so the greedy policy at that moment is close to arbitrary — collapsing epsilon locks it in. Because the agent's own actions decide which states it reaches, whole regions of the state space stop being visited at all, and their values are never updated away from their initial guesses; the agent cannot even see the better policy it is missing. The symptom is a return curve that climbs quickly, then goes flat with unusually low variance. Tabular Q-learning's convergence guarantee assumes every state-action pair keeps being updated, and a fast decay breaks exactly that. Fix it by tying the decay horizon to when state coverage stops growing, and by decaying to a small floor rather than to zero.
go deeper
Know the shape of the schedule: start near fully random, reduce as the estimates improve. Be able to say that cutting exploration too early leaves the agent stuck on whatever looked best while it still knew nothing.
Explain the mechanism, not just the symptom. Coverage is the key idea: a greedy agent stops reaching whole regions of the state space, so those values are never updated, and convergence guarantees assume continued visits to every state-action pair.
Demonstrate diagnosis. Say which signals you log — state coverage over time, realised greedy fraction, a periodic greedy-rollout return — and how you tell a too-fast decay apart from a learning-rate or representation problem.
Frame the schedule as a policy about how much reward the organisation is willing to spend learning, and for how long. Argue for a permanent exploration floor and a re-exploration trigger in environments that drift, and own that recurring cost explicitly.
### What decaying epsilon is supposed to achieve Exploration is most valuable when the estimates are worst — at the start, when the value table carries no information. It is least valuable once the estimates are good, because from then on every exploratory step is reward knowingly thrown away. A decay schedule expresses that: high `epsilon` early, low `epsilon` late. A traffic-signal controller should be trying unusual phase timings in its first days and should have stopped once it has learned the rush-hour pattern. The schedule is a **hyperparameter with two failure modes**, and interviewers ask about this because both are easy to produce and easy to misdiagnose. ### Failure mode one: decay too fast (the case asked about) Suppose `epsilon` is multiplied by 0.99 every episode. After 500 episodes it is about 0.007 — effectively greedy. If the task needs 50,000 episodes, the agent has been greedy for 99% of training with respect to a value table that was still noise when it committed. Three things then happen: 1. **Lock-in on an arbitrary policy.** The greedy action in each state is the argmax over estimates that reflect a few samples plus the initialisation. That argmax is close to arbitrary, and the agent now follows it. 2. **Coverage collapse.** This is the part that has no analogue in a one-shot problem: the agent's action choices determine which states it visits next. Once it is greedy, it walks the same trajectory, so states off that trajectory are never entered, and their values are never touched. The controller that settled during a quiet shoulder period never sees, and never learns values for, the congested states rush hour actually produces. 3. **A confident, wrong value function.** The visited part of the table converges nicely and looks self-consistent. Nothing in the training curve shouts "error" — it just plateaus. ### Why the convergence theory says the same thing Tabular Q-learning converges to the optimal action values with probability one under two conditions: the learning rates satisfy the usual stochastic-approximation conditions, and **every state-action pair continues to be updated indefinitely**. That second condition is about *coverage*, and a schedule that reaches zero quickly destroys it. The standard requirement on an exploration schedule is called **GLIE — greedy in the limit with infinite exploration**. It asks for two things simultaneously: `epsilon` tends to zero, so the behaviour becomes greedy in the limit, *and* it does so slowly enough that every state-action pair is still selected infinitely often. A schedule like `epsilon = 1/sqrt(n)` satisfies both, because the exploration mass sums to infinity while the value still tends to zero. Halving `epsilon` every episode does not: the total exploration is finite. A precision point worth having ready: because Q-learning is off-policy, it learns about the greedy policy regardless of how it behaves, so what it strictly needs is continued coverage of every state-action pair — it does not need the behaviour policy to become greedy. On-policy control, which evaluates and improves the very policy it follows, needs the full GLIE condition: the behaviour itself must converge to greedy. ### Failure mode two: decay too slowly The opposite error is less damaging but often misread. If `epsilon` stays high, the agent keeps paying exploration cost after it has learned. Its *behaviour* return sits well below what its greedy policy would earn, and the plot looks like a mediocre agent. Nothing is broken; you are simply measuring the wrong thing. **Always evaluate a separate greedy rollout** alongside the training curve — the gap between behaviour return and greedy return is the exploration bill, and seeing both tells the two failure modes apart at a glance. ### Diagnosing it Do not tune the schedule blind. Log: - **State-action visitation counts**, or a coverage proxy such as the number of distinct states visited per 100 episodes. Under a too-fast decay this stops growing early while the return is still improving — the clearest single signal. - **The fraction of actions taken greedily**, so the realised schedule is visible rather than assumed. - **Greedy-policy return** evaluated periodically with epsilon set to zero. ### Choosing a schedule Sensible defaults: anneal linearly from 1.0 to a floor of 0.01-0.05 over the first 10-50% of training, or use a count-based form such as `epsilon` proportional to `1/sqrt(visits to this state)`, which explores more in states the agent knows least. Decay per episode rather than per environment step when episode lengths vary a lot, otherwise the realised schedule silently depends on how long episodes happen to run. Keep a floor above zero in any environment that can change, and be ready to raise `epsilon` again after a known distribution shift.
- What does greedy-in-the-limit with infinite exploration require of a schedule?Two things at once. Epsilon must tend to zero so the behaviour becomes greedy in the limit, and it must do so slowly enough that every state-action pair is still selected infinitely often. A schedule like `epsilon = 1/sqrt(n)` satisfies both, since the exploration mass still sums to infinity. Halving epsilon each episode fails the second condition, because the total exploration is finite.
- What is the opposite failure, leaving epsilon high for too long?The agent keeps paying for exploration after it has learned. Its behaviour return sits below what its greedy policy would earn, so the training curve makes a good agent look mediocre. Learning usually still succeeds — it just costs more and reports worse. Evaluating a separate greedy rollout with epsilon set to zero separates the exploration bill from an actual learning problem.
- Should epsilon be decayed per environment step or per episode?Per episode is the safer default when episode lengths vary. With a per-step schedule the realised amount of exploration depends on how long episodes happen to run, so an agent that learns to end episodes quickly decays far more slowly in wall-clock terms than one that survives longer. Either way, set the horizon from when state coverage stops growing, not from a round number.
saying these in an interview costs you the question
- Says lower epsilon is always better because greedy maximises reward
- Assumes values of unvisited state-action pairs converge anyway
- Reads the early flat learning curve as the task being solved
- Decays epsilon to exactly zero in a changing environment
- Tunes the schedule without logging state coverage