skip to content

How do you tune the KL coefficient in an RLHF run, and what breaks at each extreme?

level: seniorimportance: should knowfreq 42%

answer

  1. a leash back to the starting checkpoint
  2. reward minus a penalty on drift
  3. the proxy is only valid nearby
  4. watch measured divergence, not the knob
  5. one extreme collapses, the other freezes

basics

~20 s

The KL term penalises drift from the frozen reference policy, trading reward against staying recognisably like the starting model. Too low and the policy collapses into reward-hacked, degenerate text; too high and it barely moves. Tune by watching measured KL, not the coefficient.

solid answer

~50 s

RLHF optimises reward *minus* a coefficient times the KL divergence between the current policy and a frozen reference - normally the SFT checkpoint. The penalty exists because the reward is a proxy that is only trustworthy near the distribution it was trained on; the KL term is what stops the optimiser from wandering off to exploit it. Set the coefficient too low and you get the textbook failure: reward climbs while output degenerates into repetitive, oddly-formatted text that scores brilliantly and reads terribly, with entropy collapsing and measured KL exploding. Set it too high and the policy is pinned to the reference - reward barely moves, evaluations look identical to SFT, and you have spent a lot of compute confirming your starting point. The practical method is to treat measured KL as the control variable. Pick a target KL budget, watch the actual divergence per step, and adjust the coefficient - some implementations do this automatically with an adaptive controller.

code

python · 7 lines
python
def shaped_reward(reward, logp_policy, logp_ref, kl_coef):
    kl_estimate = logp_policy - logp_ref
    return reward - kl_coef * kl_estimate


print(shaped_reward(1.20, -8.0, -9.5, kl_coef=0.02))
print(shaped_reward(1.20, -8.0, -9.5, kl_coef=0.50))

go deeper

for a junior

Know that RLHF keeps a frozen copy of the starting model and penalises the trained model for drifting too far from it, so it does not lose its general abilities.

for a middle

Explain the objective as reward minus a penalty on divergence from the reference, and describe both failure modes: degenerate reward-hacked text at one extreme, no movement at the other.

for a senior

Show you would tune by instrumenting measured KL, entropy and length rather than guessing coefficients, and that you judge the run by an independent evaluation rather than by the reward curve.

for a principal

Own the divergence budget as a policy decision: how much behavioural change you are willing to buy, what capability regressions you accept for it, and which gate has authority to stop a run that is trading quality for reward.

## The objective Every preference-RL run optimises something of the shape *reward minus beta times KL from a reference policy*. The reference is a frozen copy of where you started, usually the supervised-fine-tuned checkpoint. Implementations differ on mechanics - the penalty may be folded into the per-token reward or added as a separate loss term - but the intent is identical: reward pulls the policy away, KL pulls it back. ## Why the penalty is not optional Three reasons, and a strong answer gives more than one. **The reward is only valid locally.** A reward model was trained on responses drawn from a particular distribution. Far from it, its outputs are extrapolation, and an optimiser will find the extrapolated peaks. KL keeps the policy in the region where the reward model's opinion means something. **Capability preservation.** Pretraining and SFT bought fluency, breadth and knowledge that preference data covers only thinly. Unconstrained optimisation of a narrow objective erodes them - the alignment tax in its acute form. **Diversity.** Reward maximisation is inherently mode-seeking. Without a brake, entropy falls and the model converges on a single formulaic response shape for every prompt, which destroys sampling-based techniques downstream. ## What too low looks like The signature is a diverging pair of curves: reward up, quality down. Concretely you see measured KL growing without bound, output entropy collapsing, and text that acquires strange tics - repeated stock phrases, runaway length, unnatural formatting, sometimes near-gibberish that nonetheless scores at the top of the reward model's range. If you only watch reward, this looks like the best run you have ever had. ## What too high looks like The opposite and quieter failure. Measured KL sits near zero, the policy is effectively pinned to the reference, reward creeps, and evaluation results are statistically indistinguishable from the SFT checkpoint. Teams often misread this as *preference training does not work for our task* when it is simply over-regularised. ## How to actually tune it The coefficient itself is not interpretable - its meaning depends on reward scale, batch composition and how KL is estimated. So invert the problem: - **Instrument measured KL** per step, alongside reward, entropy and mean response length. These four curves diagnose almost everything. - **Choose a KL budget** empirically: the divergence at which held-out evaluation still improves. Beyond that point, added divergence buys reward and loses quality. - **Use an adaptive controller** where available - raise the coefficient when measured KL overshoots the target, lower it when it undershoots. This is standard in mature implementations. - **Normalise the reward** so the two terms stay comparable, otherwise a change in reward scale silently retunes your penalty. - **Evaluate checkpoints independently.** Only an evaluation the policy was never optimised against tells you which point on the reward-versus-KL curve you actually want. ## Where the same idea shows up elsewhere Offline direct preference methods carry the identical control under a different name: their temperature parameter scales an implicit KL to the reference, so a low value permits large drift and a high value keeps the policy conservative. Reasoning runs against verifiable rewards sometimes shrink or drop the penalty, on the argument that a programmatic grader is much harder to hack than a learned reward - but they then have to watch language quality and diversity directly, because nothing else is holding them. ## The framing that lands Call it a trust region on behaviour. You are not trying to minimise KL; you are trying to spend a divergence budget where it buys real quality. That framing turns a hyperparameter question into a measurement question, which is what senior interviewers are listening for.

  • Which metrics would you put on the dashboard to catch a KL problem early?
    Measured KL from the reference, mean reward, output entropy and mean response length, all per step. Diverging reward with exploding KL and falling entropy is under-regularisation; flat everything with near-zero KL is over-regularisation. Add periodic scoring of checkpoints on a held-out evaluation the policy is not optimised against, since reward alone cannot distinguish improvement from exploitation.
  • Why do some verifiable-reward reasoning runs shrink or drop the KL term?
    When the reward is a programmatic grader such as a test runner, there is far less to hack than in a learned reward, so the main justification for tight anchoring weakens, and letting the policy travel further tends to produce stronger reasoning. The cost is that nothing is preserving general language quality or output diversity, so those have to be evaluated explicitly rather than assumed.
  • Does the reference policy have to stay frozen for the whole run?
    Not necessarily. Some setups periodically refresh the reference to a recent policy checkpoint, which effectively converts a global constraint into a per-phase trust region and permits larger cumulative movement. It buys progress on long runs but weakens the guarantee that the final model stays close to the original SFT behaviour, so capability regressions need watching more carefully.

It is a leash on a dog that has learned exactly which bin has food: short enough and it stays on the path, long enough and it never gets anywhere, and the right length is judged by watching the dog rather than the leash.

saying these in an interview costs you the question

  • Treating the KL coefficient as a value you tune blind, without measuring KL
  • Claiming a KL penalty eliminates reward hacking rather than limiting it
  • Believing lower KL is always better because it means a safer model
  • Confusing the frozen reference policy with the reward model
  • Ignoring entropy and length collapse as symptoms of under-regularisation

context