skip to content

What does the discount factor gamma control in a reinforcement learning agent's return?

level: middleimportance: must knowfreq 64%

answer

  1. weight on a reward k steps away
  2. geometric decay, gamma to the k
  3. one over one minus gamma
  4. makes an endless sum converge
  5. myopic at zero, far-sighted near one

basics

~20 s

Gamma sets how much a reward arriving k steps later is worth now: it is weighted by gamma^k. Low gamma makes the agent short-sighted, high gamma far-sighted, and gamma below 1 keeps the total finite in a task that never ends.

solid answer

~40 s

The return an agent maximises is the discounted sum `G_t = r_{t+1} + gamma*r_{t+2} + gamma^2*r_{t+3} + ...`, so a reward `k` steps away is worth `gamma^k` of its face value now. Gamma therefore sets the effective planning horizon, roughly `1/(1 - gamma)` steps. At `gamma = 0` the agent optimises only the immediate reward; at `gamma = 0.9` a +1 twenty steps out is worth about 0.12, so a warehouse robot would grab the pallet two aisles away and ignore the one twenty aisles away; at `gamma = 0.99` that same distant pallet is still worth about 0.82 and becomes worth fetching. Gamma below 1 is also what makes an infinite sum converge in a continuing task with no terminal state. In genuinely episodic tasks that always end, `gamma = 1` is legitimate.

code

python · 20 lines
python
# What a +1 reward is worth now, depending on how many steps away it is.
def present_value(steps_away, gamma):
    return gamma ** steps_away

for gamma in (0.5, 0.9, 0.99):
    near = present_value(2, gamma)
    far = present_value(20, gamma)
    horizon = 1.0 / (1.0 - gamma)
    print(f"gamma={gamma:<5} near={near:.3f} far={far:.6f} horizon~{horizon:.0f} steps")

# Discounted return of a reward stream, accumulated backwards from the end.
def discounted_return(rewards, gamma):
    total = 0.0
    for r in reversed(rewards):
        total = r + gamma * total
    return total

# A single +1 on the third step is weighted by gamma squared.
print(discounted_return([0, 0, 1], 0.9))   # 0.81
print(discounted_return([1, 1, 1], 0.9))   # 1 + 0.9 + 0.81 = 2.71

go deeper

for a junior

Know the shape of the discounted return and that a reward k steps out is multiplied by gamma to the power k. Be able to say what gamma = 0 and gamma close to 1 do to behaviour.

for a middle

Explain the two reasons discounting exists: it sets the effective horizon of roughly one over one minus gamma, and it makes an infinite sum converge in a task with no end. Do the arithmetic out loud if asked.

for a senior

Diagnose with it. When an agent refuses to make a delayed-payoff move, compare the effective horizon to the delay, and remember that changing the control interval invalidates the old gamma.

for a principal

Argue whether the discount is a modelling convenience or a real business preference. When the domain has a genuine churn hazard or cost of capital, defend deriving gamma from it rather than tuning it as a hyperparameter.

## The return, not the reward An agent in an MDP does not maximise the next reward; it maximises the **return**, the accumulated reward over the rest of the trajectory. The standard form is the **discounted return**: ``` G_t = r_{t+1} + gamma*r_{t+2} + gamma^2*r_{t+3} + ... = sum over k >= 0 of gamma^k * r_{t+k+1} ``` with `gamma` in `[0, 1]`. A reward `k` steps into the future is multiplied by `gamma^k`. That single number does three jobs at once. ## Job one: it sets the planning horizon Because the weights decay geometrically, gamma decides how far ahead the agent effectively looks. A useful rule of thumb is an **effective horizon of about `1/(1 - gamma)` steps**: `gamma = 0.9` gives roughly 10 steps, `gamma = 0.99` roughly 100, `gamma = 0.999` roughly 1000. Rewards well beyond that horizon contribute almost nothing to the return, so the agent behaves as if they did not exist. Concretely, take a warehouse parcel-sorting robot that receives +1 per parcel delivered, and suppose reaching a pallet takes one step per aisle: | gamma | +1 two aisles away | +1 twenty aisles away | |---|---|---| | 0.5 | 0.25 | about 0.000001 | | 0.9 | 0.81 | about 0.12 | | 0.99 | 0.98 | about 0.82 | At `gamma = 0.5` the far pallet is invisible and the robot works only its local aisle. At `gamma = 0.99` the two pallets are nearly equally attractive and the robot will cross the warehouse. **Nothing about the environment changed — only the objective did.** This is why gamma is a modelling decision with real behavioural consequences, not a nuisance hyperparameter. ## Job two: it keeps the return finite Tasks come in two shapes. - **Episodic**: there is a terminal state and every trajectory ends — a board game reaching checkmate, a delivery run completing. Returns are finite sums, and `gamma = 1` is perfectly well defined. - **Continuing**: there is no terminal state and the interaction goes on indefinitely — a data-centre cooling controller that runs forever. Here the undiscounted sum of rewards can diverge, and two policies that both accumulate infinite reward cannot be compared. With `gamma < 1` and bounded rewards (`|r| <= R_max`), the return is bounded by `R_max / (1 - gamma)`, so it always converges and policies stay comparable. That is the formal reason discounting exists. The alternative for continuing tasks is an average-reward formulation, which optimises reward per step instead — a valid but less common choice. ## Job three: it expresses genuine preference Sometimes discounting is not a mathematical convenience but the truth about the problem. Money now really is worth more than money next year; a customer served today really is worth more than one served next month, because they may churn in between. When the domain has a genuine interest rate or survival hazard, gamma should reflect it rather than be tuned. ## Choosing gamma - **Match it to the horizon that matters.** If a decision's consequences play out over about 50 steps, a gamma with an effective horizon of 5 steps will systematically choose wrong. - **Mind the time step.** Gamma is per *step*, so halving the control interval doubles the number of steps in the same wall-clock horizon and the old gamma is now far too impatient. If you change the sampling rate, you must re-derive gamma. - **Higher is not free.** As gamma approaches 1, return estimates have higher variance, credit takes longer to propagate backwards over many steps, and learning slows down markedly. Practitioners often start lower and raise gamma as learning stabilises. - **Never set gamma = 1 for a continuing task.** The objective is then undefined and the algorithm's estimates can grow without bound. ## What gamma is not It is not a learning rate — that scales an update; gamma changes what the agent is trying to achieve at all. It is not a way to make the agent safer or more cautious; a myopic agent is not a conservative agent, it is just one that ignores consequences. And it does not change the environment's dynamics: two agents with different gammas face the same transitions and will still learn different policies, because they are solving different objectives on the same world. ## The diagnostic use When an agent stubbornly refuses to make an investment that pays off later — the robot never crossing the warehouse, the controller never spinning up cooling in advance of a heat spike — the first suspect is a gamma whose effective horizon is shorter than the payoff delay. Compute `1/(1 - gamma)`, compare it to the number of steps between the action and its consequence, and you often have the diagnosis before you look at anything else.

  • When is gamma = 1 a legitimate choice?
    In an episodic task that is guaranteed to terminate — a board game, a delivery run with a fixed end. The return is then a finite sum and needs no discounting to converge. It is never safe in a continuing task such as a cooling controller with no terminal state, where the undiscounted sum can diverge and two policies become impossible to compare.
  • You halve the control interval so the agent acts twice as often. What must change?
    Gamma, because it discounts per step, not per second. The same wall-clock horizon now spans twice as many steps, so the old gamma halves the effective horizon and the agent turns myopic. Roughly, the new gamma should be the square root of the old one to preserve the horizon in real time.
  • Why not just set gamma very close to 1 and always be far-sighted?
    Because it costs stability. Returns over a long horizon have much higher variance, credit has to propagate back over many more steps, and learning slows and becomes less stable. Gamma near 1 also amplifies any bias in the estimates. You pick the smallest horizon that still captures the consequences you care about.

Gamma is an interest rate run backwards: a reward promised twenty steps from now is discounted to its present value, and a low gamma is a punishing rate that makes anything distant worthless today.

saying these in an interview costs you the question

  • Calls gamma a learning rate or a step size
  • Sets gamma to 1 on a task with no terminal state
  • Thinks a low gamma makes the agent safer
  • Keeps gamma fixed after changing the time step
  • Says gamma changes the environment's dynamics

context