skip to content

Why does SARSA learn a safe inland route while Q-learning hugs the cliff edge under the same epsilon-greedy exploration?

level: seniorimportance: should knowfreq 45%

answer

  1. exploration is part of the policy being valued
  2. one method evaluates the policy it follows
  3. a random step beside a large penalty
  4. the greedy target never falls in
  5. shrinking exploration makes them agree

basics

~20 s

Q-learning bootstraps from the greedy next action, so its values describe an agent that never steps off the cliff, and the short edge path looks best. SARSA bootstraps from the exploratory action actually taken, so occasional random falls depress the edge cells' values.

solid answer

~50 s

Picture the grid: start and goal on the bottom row, a strip of cliff cells between them, stepping into a cliff cell costs -100 and teleports the agent back to start, every other step costs -1. The shortest route runs along the row beside the cliff. Q-learning's target uses `max_a' Q(s',a')`, so its values describe an agent that never deliberately steps off; nothing about the exploratory falls reaches the edge cells, and it correctly learns that edge route as the optimal policy. SARSA's target uses the action its epsilon-greedy policy actually takes, and beside the cliff that action is sometimes a random step into it, so the -100 propagates back and the safer inland route wins. SARSA is right about the policy being executed — it collects more reward per episode during training — while Q-learning is right about the greedy policy. Decay epsilon to zero and SARSA's policy converges to the edge route too.

go deeper

for a junior

Know that an exploring agent sometimes takes a random action, so a large penalty sitting next to the shortest path can be triggered by accident during training even when the agent knows better.

for a middle

Trace how the exploratory next action enters one method's target and not the other's, and say which route each ends up preferring and what changes as exploration shrinks.

for a senior

Say which learned policy you would ship given how the agent will be executed, and separate reward collected during training from the quality of the final policy — the ranking reverses here.

for a principal

Take a position on the risk this encodes: whether the organisation tolerates catastrophic outcomes during learning, whether exploration continues after launch, and what that implies for the method you standardise on.

## The grid The cliff-walk is a small rectangular gridworld. Start is the bottom-left cell, goal the bottom-right, and the cells between them along the bottom row are **cliff**: entering one yields a reward of -100 and sends the agent straight back to start. Every other move yields -1, so the agent is being asked to reach the goal in as few steps as possible. The agent moves up, down, left or right, and acts epsilon-greedily: the current best action with probability `1 - epsilon`, a uniformly random action otherwise. The genuinely optimal path — the one that maximises return if you can execute your intentions perfectly — is to step up once, walk straight along the row directly above the cliff, and drop into the goal. Any inland detour costs extra -1 steps. ## What each method learns, and why **Q-learning.** Its target is `r + gamma * max_a' Q(s',a')`. Standing in a cell beside the cliff, the maximum is taken over the actions a *greedy* agent would consider best, and a greedy agent never walks into a -100 cell. The occasional exploratory fall does produce a real update — the specific pair (edge cell, step-down) is punished — but the value of the *good* actions in that cell is computed as if the agent will behave perfectly from the next step onward. So the edge cells retain high value, and Q-learning converges to the optimal action values `Q*` and the cliff-edge policy. **SARSA.** Its target is `r + gamma * Q(s',a')`, where `a'` is the action the epsilon-greedy policy actually selects. In an edge cell, with probability roughly `epsilon / 4` that action is the step into the cliff. That means the *value of being in an edge cell at all* absorbs a share of the -100, because the value being bootstrapped is the value of continuing to behave epsilon-greedily. Propagate that back a few cells and the inland route — one or two rows away from the cliff, where a random step is harmless — comes out ahead. SARSA is not being timid by construction; it is correctly evaluating a policy that fumbles a fraction of its moves. ## Neither one is wrong They answer different questions: - Q-learning answers "what is the best possible policy?" Its answer is right, and it is right even though the data was collected by a clumsy agent. - SARSA answers "what is my current, exploring agent worth, and how should it behave?" Its answer is also right, for that agent. The measurable consequence during training is the reverse of what people expect: **SARSA earns more reward per episode online than Q-learning**, because Q-learning keeps walking the edge and keeps falling in. The agent with the better learned policy has the worse learning experience. ## Which one do you ship? The decision hinges on how the agent is executed after training: - If exploration is switched off and you deploy the greedy policy, Q-learning's answer is the one you want, and SARSA's conservative route is needlessly long. In that setup you can also just decay epsilon toward zero during SARSA training: as exploration vanishes, the value being evaluated approaches the greedy one and SARSA converges to the same edge path. - If randomness persists at execution time — actuators that occasionally miss, a policy deliberately kept stochastic, an agent that keeps learning in the field — then SARSA's values are the honest description of what you will actually get, and its route is the better bet. - If a single fall is catastrophic rather than merely -100, the argument stops being about which value is more accurate: you should not be learning against a live cliff at all, and you would take the conservative policy regardless of the mathematics. ## A distinction that catches people out The safe-versus-optimal split is caused by **randomness in action selection**, not by randomness in the environment. If the grid had wind that pushed the agent sideways with some probability, that stochasticity sits inside the MDP's transitions, and *both* methods account for it in expectation — Q-learning's `Q*` would already price in the chance of being blown off the edge. Only the exploration noise is treated differently by the two targets. Candidates who blur these two sources of randomness usually reach the right conclusion for the wrong reason. ## Cheap diagnostics If you suspect this effect in your own agent, compare the greedy policy extracted from the table against the average return actually collected per episode. A method whose learned policy looks excellent while its online return is poor is telling you exactly this story: the values describe a policy the agent is not currently running.

  • Exploration is switched off at deployment. Which learned policy do you ship?
    Q-learning's, because the greedy edge route is genuinely optimal once the agent stops taking random actions, and SARSA's inland detour just pays extra step costs for a risk that no longer exists. If you prefer SARSA for other reasons, decay the exploration rate toward zero during training — its policy then converges to the same edge route.
  • Does the same argument apply if the environment itself is stochastic, say wind that pushes the agent sideways?
    Not in the same way. Transition randomness lives inside the MDP, and both methods account for it in expectation — the optimal action values already price in the chance of being blown into the cliff. The SARSA-versus-Q-learning split is caused specifically by randomness in action selection, so a windy grid changes both answers together rather than separating them.
  • Can SARSA ever converge to the optimal policy?
    Yes, provided exploration is greedy in the limit while every state-action pair is still visited infinitely often — decay the exploration rate to zero slowly enough — and the step sizes decay appropriately. With a fixed exploration rate it converges instead to the values of that fixed epsilon-greedy policy, which is a different and legitimately different fixed point.

A tightrope walker who knows he will occasionally be shoved chooses the wider ledge, even though the narrow one is shorter.

saying these in an interview costs you the question

  • Says SARSA is simply the worse algorithm here
  • Claims Q-learning does not explore during training
  • Blames the difference on the learning rate or discount factor
  • Says the two converge to different policies even as exploration vanishes
  • Assumes the method with the better policy also collects more reward while training

context