In reinforcement learning, how do you add intermediate rewards without changing the optimal policy?
answer
- add a difference, not a bonus
- a function of state, differenced
- gamma times next, minus current
- the terms telescope along a trajectory
- pay for progress, never for position
basics
~20 sUse potential-based shaping: add gamma times a potential of the next state minus the potential of the current state. Because these terms cancel along any trajectory, the optimal policy is provably unchanged while the agent still gets useful intermediate signal.
solid answer
~40 sSparse reward makes learning slow, so the temptation is to hand out partial credit — but arbitrary bonuses change what is optimal. The safe construction is **potential-based shaping**: pick any potential function `Phi(s)` over states and add `F(s, s') = gamma*Phi(s') - Phi(s)` to the reward on every transition. This is provably **policy-invariant**: the shaping terms telescope, so along any trajectory their sum collapses to `-Phi(s0)` plus a terminal term, a constant that cannot reorder policies. For a delivery drone you would take `Phi(s) = -distance_to_drop_point`, which pays the agent for net progress. Contrast the naive version — a per-step bonus for merely *being* close — which invites the drone to hover near the drop point farming the bonus forever instead of landing. Shaping speeds learning; it never fixes a wrong objective.
go deeper
Know that reward is the objective, so adding extra reward can change what the agent decides is best. Recognise that a bonus for sitting near the goal invites the agent to sit there.
Be able to write the potential-based form, gamma times the potential of the next state minus the potential of the current one, and say why the terms cancel along a trajectory.
Show judgment about when to shape at all, how to pick a potential from domain knowledge, and how you would detect that a shaping term is being farmed rather than driving progress.
Own the boundary. Argue when a sparse but honest reward with more compute beats a dense engineered one, and set the review discipline that stops teams from bolting on bonuses whenever learning stalls.
## Why anyone shapes rewards The honest reward for a delivery drone is +1 when the parcel is on the doorstep and 0 otherwise. That specification is correct and nearly unlearnable: the agent must stumble onto a complete successful delivery by chance before it receives any signal at all. **Reward shaping** adds intermediate reward so that partial progress is rewarded and learning starts sooner. The danger is equally obvious. Reward *is* the objective. Change it and you may change what the optimal policy is, and the agent will faithfully optimise your new, wrong objective. ## The naive version and how it fails The reflex is to pay the agent for looking good: a small bonus each step it is within a few metres of the drop point. Now consider what maximises the shaped return. Landing ends the episode and stops the bonus. Hovering nearby collects it forever. If the per-step bonus is large enough relative to the terminal reward, **hovering is strictly optimal under the shaped reward**, and the agent will learn precisely that. The bug is not the agent's; it is that the designer created a reward loop the agent could ride indefinitely. The general pattern: any shaping that rewards a *state* repeatedly, rather than *progress between states*, can create a positive cycle the agent will exploit instead of finishing. ## Potential-based shaping The fix is a specific functional form. Choose any **potential function** `Phi(s)` assigning a number to each state, and add to the reward on every transition: ``` F(s, s') = gamma*Phi(s') - Phi(s) ``` So the agent learns on `r'(s, a, s') = r(s, a, s') + F(s, s')` instead of `r`. The classic result (Ng, Harada and Russell, 1999) is that this transformation is **policy-invariant**: the optimal policy of the shaped MDP is exactly the optimal policy of the original. **Why it works.** Sum the shaping terms along a trajectory, discounting each by the step at which it occurs: ``` sum_t gamma^t * (gamma*Phi(s_{t+1}) - Phi(s_t)) = sum_t (gamma^{t+1}*Phi(s_{t+1}) - gamma^t*Phi(s_t)) ``` Every term cancels against the next — the sum **telescopes** — leaving `-Phi(s_0)` plus a vanishing (or terminal) tail. The shaped return therefore differs from the true return by a quantity that depends only on the starting state, identical for every policy. A constant offset cannot reorder policies, so the optimum is untouched. Equivalently, the shaped optimal action-values satisfy `Q_shaped(s, a) = Q(s, a) - Phi(s)`. The shift depends only on the state, never on the action, so the best action in each state is unchanged. ## Applying it to the drone Take `Phi(s) = -d(s)`, the negative distance to the drop point. Then ``` F(s, s') = gamma*(-d(s')) - (-d(s)) = d(s) - gamma*d(s') ``` With `gamma` near 1 this is essentially the **reduction in distance** achieved this step. Now: - Flying one metre closer pays a small positive amount. - Hovering pays approximately nothing, because `d(s') = d(s)`. - Flying out and back pays approximately nothing net, because the two moves cancel. The hover exploit is gone by construction. This is the practical intuition worth carrying: **pay for progress, not for position.** ## What shaping does and does not buy you **Does**: faster learning, because credit reaches early actions without waiting for a rare terminal success; a denser signal in long-horizon tasks; a clean way to inject domain knowledge (a distance metric, a heuristic estimate of remaining cost) without corrupting the objective. **Does not**: change which policy is best; make a wrong objective right; guarantee an *improvement*. A badly chosen `Phi` — one that points the agent toward a dead end — remains policy-invariant but can make learning slower than no shaping at all. Policy invariance is a safety property, not a performance promise. ## Choosing the potential The ideal `Phi` is the true value function of the optimal policy: shaping with it makes the problem almost trivially easy. You do not have it, of course, which is why any decent approximation helps — negative distance, negative remaining cost, a hand-built heuristic, or a value estimate learned from a simpler related task. For episodic problems, set `Phi` of the terminal state to zero so the telescoping closes cleanly. ## Interview framing The strong answer has three beats: name why sparse reward needs help; give the potential-based form and the telescoping argument for why the optimum survives; and contrast a naive per-step bonus with the failure it invites. Then close with the boundary: shaping is a *learning-speed* intervention. If the agent is doing the wrong thing, the answer is to fix the objective, not to shape harder.
- Does potential-based shaping change the learned values as well as the policy?Yes, the values shift — the shaped optimal action-value equals the original minus the potential of that state. But the shift depends only on the state, not on the action, so the ranking of actions within each state is untouched and the greedy policy is identical. If you need the true values for reporting, add the potential back.
- Can potential-based shaping ever make learning worse?Yes. Policy invariance guarantees the destination, not the speed. A potential that points toward a dead end, or one whose scale swamps the real reward, can send the agent down unproductive paths for a long time. It will still eventually find the same optimum, but slower than with no shaping at all.
- What would you do if no sensible potential function exists?Attack the sparsity a different way: enrich the state so progress is easier to represent, shorten the episode or start it closer to success and gradually extend, or build a curriculum of easier variants. Any of these preserves the objective, which arbitrary bonus rewards do not.
It is like reimbursing an employee for distance travelled rather than paying them per hour spent near the destination. The first rewards arriving; the second rewards loitering.
saying these in an interview costs you the question
- Adds an arbitrary per-step bonus and calls it shaping
- Rewards being near the goal rather than approaching it
- Thinks any shaping preserves the optimal policy
- Claims shaping can repair a wrong objective
- Forgets to zero the potential at terminal states