skip to content

Why add an entropy bonus to the policy loss when training an actor-critic agent?

level: middleimportance: should knowfreq 55%

answer

  1. keeps the action distribution spread out
  2. on-policy data is self-selected
  3. premature determinism is self-reinforcing
  4. minus sum p log p, added to the objective
  5. small coefficient, decayed toward zero

basics

~20 s

An entropy bonus rewards a spread-out action distribution, so the policy does not collapse onto one action before the alternatives have been tried. Its coefficient trades exploration against exploitation and is usually small and decayed toward zero.

solid answer

~50 s

The policy objective gets an extra term proportional to the entropy of the action distribution, `H = -sum p log p` over the actions in each visited state, scaled by a small coefficient. It exists because an on-policy agent only ever sees the actions it samples: if one action's probability drifts up early, it is sampled more, gets more updates, and the alternatives stop appearing in the data at all. That feedback loop is self-reinforcing, and a policy can lock onto a mediocre action long before it has any evidence. The entropy gradient pushes back toward a spread distribution and keeps the alternatives alive. The coefficient is a real trade: too small and you get early collapse, too large and the policy stays near-uniform and never exploits. A common pattern is a small coefficient decayed toward zero over training, with policy entropy tracked as a first-class diagnostic.

go deeper

for a junior

Know that the term keeps the policy from becoming deterministic too early, and that entropy is high for a spread-out action distribution and low for a near-certain one.

for a middle

Be ready to write the entropy of a discrete action distribution, say which sign it enters the loss with, and explain the self-reinforcing loop that makes premature collapse so hard to undo in on-policy training.

for a senior

Demonstrate that you read the entropy curve during a run and act on it: distinguish an early cliff from a healthy decline, connect a flat high curve to an oversized coefficient, and know that recovery after collapse usually means restarting.

for a principal

Own the schedule as a policy decision. Argue when a permanently stochastic policy is what the product actually wants - adversarial or non-stationary settings - versus when the bonus must decay to zero because deployment needs a decisive agent.

## What the term is For a discrete action set, the entropy of the policy at a state is `H(pi(.|s)) = -sum_a pi(a|s) * log pi(a|s)`. It is maximal - equal to `log(number of actions)` - when the policy is uniform, and it goes to zero as the policy becomes deterministic. The entropy bonus adds `beta * H` to the objective being maximised, equivalently subtracts `beta * H` from the loss being minimised, with `beta` a small positive coefficient. For a continuous policy parameterised as a Gaussian, the same term is a function of the log of the standard deviation, so the bonus resists the spread shrinking to zero. ## Why an on-policy actor-critic needs it The agent learns only from actions it actually took. Suppose one action's probability drifts upward by chance early in training - a couple of lucky rollouts, or simply the initialisation. It is now sampled more often, so it accumulates more updates, so its probability rises further; the other actions appear in fewer and fewer batches, so the advantage estimator has less and less to say about them, so nothing pulls them back up. This is a positive feedback loop, and it can drive a policy to near-determinism in a few thousand updates on evidence that would not survive a second look. A card-game agent shows the failure vividly. With no entropy term, an agent can converge to playing the same card class in every situation: it wins often enough for the advantage signal to stay non-negative, the other plays vanish from the rollouts, and the run flatlines at a mediocre score while the loss curves look perfectly healthy. Add a small entropy bonus and the alternatives keep being sampled long enough for their advantages to be estimated, and the policy escapes. Note that this is not the same as adding random noise to the chosen actions. The policy is already stochastic; the bonus changes the *shape* of the distribution the agent is optimising toward, so exploration and the learned policy are the same object rather than a policy plus a separate exploration rule bolted on. ## Choosing and scheduling the coefficient The coefficient is genuinely two-sided. Too small, and you get the collapse above. Too large, and the entropy term dominates: the policy stays close to uniform, returns plateau at whatever a nearly random policy achieves, and the value function - which is trying to predict returns under an almost random policy - learns something useless for the policy you actually want. A typical starting point for a small discrete action set is around one hundredth of the policy loss scale, tuned by watching entropy rather than by watching returns. Because the need for exploration is highest early, a common pattern is to decay the coefficient from a small starting value toward zero across training. Decaying matters for a second reason: with a non-zero bonus the agent is optimising a modified objective, and the policy it converges to is deliberately more stochastic than the best policy for the raw return. If you want the sharp policy at the end, the bonus has to fade. ## Entropy as a diagnostic Log the mean policy entropy every update; it is one of the two or three most informative curves in an on-policy run. - A steady, gentle decline is healthy: the agent is becoming more decisive as it accumulates evidence. - A cliff in the first fraction of training is collapse. The usual causes are too large a learning rate, an advantage scale that is too big, or too small a coefficient. - A flat line near the maximum, `log(number of actions)`, means the bonus is dominating and the agent is barely committing to anything. - Entropy that *rises* late in training is a warning sign, usually of an unstable value function feeding contradictory advantages to the actor. ## Once it has collapsed Raising the coefficient after the fact rarely rescues a run. Once probabilities are near zero for most actions, both the entropy gradient and the policy gradient at those actions are tiny, and - more importantly - the data no longer contains those actions, so there is nothing to learn from even if the distribution widens slightly. The realistic fixes are structural: restart with a larger coefficient, a smaller learning rate, or a smaller advantage scale. Prevention is much cheaper than recovery, which is why the entropy curve is worth watching from the first update rather than at post-mortem time.

  • What would tell you from the training curves that the entropy coefficient is too high?
    Policy entropy sits flat near its maximum, `log(number of actions)`, instead of declining, while episode return plateaus early at roughly what a near-random policy scores. The agent is being paid more for staying undecided than for choosing well. Lower the coefficient or start decaying it sooner, and check that the advantage scale has not shrunk so far that the entropy term overwhelms it.
  • Should the entropy coefficient stay fixed for a whole run?
    Usually not. Exploration is most valuable early, and a non-zero bonus means you are optimising a deliberately more stochastic policy than the raw return prefers. Decaying it toward zero over training gives exploration when it matters and a sharp policy at the end. Keep it fixed only when the task keeps changing under you and you want the agent to stay adaptable.
  • A policy has already collapsed onto one action. Can raising the coefficient recover it?
    Rarely. With probabilities near zero on every other action, the gradients there are tiny and, more decisively, the rollouts no longer contain those actions, so there is no advantage signal to learn from. The practical fix is a restart with a larger initial coefficient, a smaller learning rate, or a smaller advantage scale, not a mid-run bump.

saying these in an interview costs you the question

  • Describes it as random noise added to chosen actions
  • Treats the coefficient as harmless at any size
  • Confuses policy entropy with randomness in the rewards
  • Says exploration here is handled by an epsilon-greedy rule
  • Calls rising entropy late in training a good sign

context