How does SAC's maximum-entropy objective differ from maximising return alone?
answer
- reward plus a temperature times entropy
- the optimal policy is stochastic on purpose
- entropy sits inside the bootstrap target
- temperature is the reward-per-nat exchange rate
- tuned to hit a target average entropy
basics
~20 sSAC maximises expected return plus a temperature times the policy's entropy at every visited state, so the optimal policy is stochastic rather than a single best action. The entropy term also enters the critic's bootstrap target, so the values learned are entropy-augmented, not ordinary returns.
solid answer
~50 sStandard control maximises the expected discounted sum of rewards; SAC maximises the expected sum of `r + alpha * H(pi(.|s))`, where `H` is the policy's entropy at that state and `alpha` is a temperature. Two consequences matter. First, the optimum is genuinely stochastic — the agent is paid to keep spread wherever spread is cheap in reward, which gives exploration that is state-dependent and automatic instead of an externally scheduled noise process. Second, entropy is folded into the *value*, not bolted onto the actor: the bootstrap target is `r + gamma * (min over twin critics of Q(s', a') - alpha * log pi(a'|s'))` with `a'` freshly sampled from the current policy, so the critic learns soft values. The actor then maximises `Q(s, a) - alpha * log pi(a|s)` through a reparameterised sample. The temperature is normally auto-tuned rather than fixed: it is adjusted so the policy's average entropy tracks a target, rising when the policy is collapsing and falling when it is too random.
go deeper
Recall that this family of agents is rewarded for staying random as well as for reward, so exploration comes from the objective instead of from noise added by hand. Know that the temperature is the knob controlling how much randomness is worth.
Be able to write the objective as reward plus temperature times entropy and to say that the entropy term also appears in the critic's bootstrap target using a freshly sampled next action. Explain the temperature as an exchange rate between reward and nats.
Demonstrate operating judgment: how the temperature is auto-tuned against a target entropy, which direction it moves when the policy collapses, and why a hand-fixed temperature has to be re-tuned as a task's precision demands change.
Own the framing: this is a change of objective, not a trick, and it buys tuning-free state-dependent exploration at the cost of an optimum that is not reward-optimal. Be ready to argue when that trade is right and how you would set or schedule the entropy target across a fleet of tasks.
## The objective Ordinary control maximises `E[ sum_t gamma^t r_t ]`. The maximum-entropy formulation maximises E[ sum_t gamma^t ( r_t + alpha * H(pi(.|s_t)) ) ] where `H(pi(.|s)) = - E_{a ~ pi} [ log pi(a|s) ]` is the policy's entropy at state `s`, and `alpha > 0` is the **temperature**. The agent is paid both for reward and for keeping its action distribution spread out. This is not a tweak of the actor's loss function; it changes what "optimal" means. Under this objective the optimal policy is stochastic by construction: it concentrates only where concentrating actually buys reward, and stays broad where the reward surface is flat. ## Soft values: the entropy is inside the critic The important structural point, and the one candidates most often miss, is that the entropy term appears in the **bootstrap target**: a' ~ pi(.|s') (a fresh sample, not a stored action) y = r + gamma * ( min(Q1_target, Q2_target)(s', a') - alpha * log pi(a'|s') ) critic loss = (Q(s, a) - y)^2 for each critic Because `- log pi(a'|s')` is a single-sample estimate of the entropy at `s'`, the critic is learning the *entropy-augmented* return, sometimes called the soft Q-function. If entropy lived only in the actor's loss, the actor and critic would be optimising different objectives and the bootstrapped values would understate the worth of remaining stochastic. SAC also carries twin critics with a minimum in the target, for the same anti-optimism reason as any modern continuous-control critic. ## The actor's update The actor maximises E_{a ~ pi(.|s)} [ Q(s, a) - alpha * log pi(a|s) ] over states drawn from stored experience. Because the action is sampled, the gradient is taken by reparameterising: draw a fixed-distribution noise variable, push it through the policy's deterministic map to get the action, and differentiate through both the critic and the log-density. The first term pulls toward high-value actions; the second is a pull toward keeping the density low, i.e. spread out. `alpha` is exactly the exchange rate between one unit of reward and one nat of entropy. ## Auto-tuning the temperature A fixed `alpha` is a poor idea, because the reward scale differs per task and per training phase, and a temperature that is right early is usually wrong late. SAC instead treats the objective as a constrained problem — maximise return subject to the policy's average entropy being at least some target `H_target` — and adjusts `alpha` by gradient on the constraint violation. The behaviour is easy to state: - measured entropy **below** target -> `alpha` rises -> more pressure to spread out; - measured entropy **above** target -> `alpha` falls -> more pressure to exploit. `alpha` is optimised in log-space so it stays positive. The common heuristic sets `H_target` to minus the dimensionality of the action vector — one negative nat per action coordinate. That number is worth understanding rather than memorising: for a single coordinate bounded to the interval from -1 to 1, a uniform policy has differential entropy `log 2 ≈ 0.69`, so a target of -1 per coordinate asks for a policy substantially more concentrated than uniform while still meaningfully random. (Differential entropy can be negative; this is not a bug.) On a dexterous in-hand manipulation task — rotating a cube between fingertips — this shows up concretely. Early on the policy needs to be broad enough to discover regrasps at all. Later, the contacts are delicate and a policy that keeps jittering drops the cube; the auto-tuned temperature falls on its own as the return signal grows, whereas a hand-fixed temperature has to be re-tuned by a human at exactly that transition. ## What the entropy term buys, and what it costs Buys: - **State-dependent exploration.** Randomness where the value surface is flat, precision where it is sharp — a fixed injected noise level cannot do this. - **Robustness to critic sharpness.** A policy that must keep spread cannot commit hard to one razor-thin spike in the critic. - **Far less hyperparameter tuning**, since the exploration knob is optimised rather than scheduled. Costs: - The learned optimum is **not** the reward-optimal policy; it is the optimum of a different objective, and with a large `alpha` the gap is visible as a permanently sloppier controller. - Deployment needs a decision about whether to keep sampling or act at the distribution's centre. - The entropy estimate depends on getting the action density right, which is exactly where a bounded-action transform can silently corrupt everything.
- How is the temperature actually adjusted during training?By a gradient step on a constraint: the policy's average entropy is compared with a target, and the temperature rises when entropy falls below it and falls when entropy sits above it. It is optimised in log-space so it stays positive. The effect is a controller that keeps the policy's randomness near a chosen level regardless of the task's reward scale.
- What does a target entropy of minus the action dimensionality actually ask for?One negative nat per action coordinate. For a coordinate bounded to the interval from -1 to 1, a uniform policy has differential entropy about 0.69, so a target of -1 requests a policy noticeably more concentrated than uniform but still clearly stochastic. It is a scale-free heuristic, not a derived optimum, and tightening it yields a sharper, more exploitative policy.
- Why does the entropy term appear in the critic's target rather than only in the actor's loss?Because the quantity being learned is the entropy-augmented return. If the critic bootstrapped ordinary returns while the actor optimised return-plus-entropy, the two would be solving different problems, and the value of remaining stochastic in future states would never propagate backwards. Putting minus alpha times the log-density into the target makes the critic a soft value function.
- Does maximising this objective give you the reward-optimal policy?No, and that is worth saying plainly. The optimum of the entropy-augmented objective is a deliberately spread-out policy, and with a large temperature its reward is measurably below the deterministic optimum. The bet is that better exploration and less tuning during training outweigh the residual gap, and that acting at the distribution's centre at evaluation recovers most of it.
The temperature is an exchange rate between reward and randomness: the agent will only give up spread where the reward it buys is worth more than the entropy it costs.
saying these in an interview costs you the question
- Treats the temperature as a fixed exploration probability
- Says entropy is added only to the actor's loss
- Claims the entropy-augmented optimum equals the reward optimum
- Thinks a larger temperature always improves final performance
- Cannot say which way the temperature moves when entropy drops