How is a reward model trained, and why does RLHF reward hacking happen?
answer
- a learned stand-in for the annotator
- comparisons, not absolute scores
- only the score margin is trained
- Goodhart: pressure breaks a decent proxy
- length and agreement are the cheap correlates
basics
~20 sA reward model learns from human-ranked response pairs to score the preferred one higher. Because it only approximates human taste, optimising hard against it finds cheap correlates - length, confident tone, flattery - instead of genuine quality.
solid answer
~50 sYou collect **pairwise preferences**: for one prompt, a human sees two candidate responses and marks which is better. A reward model - usually the SFT checkpoint with its language-modelling head replaced by a scalar head - is trained so the chosen response scores higher than the rejected one, with a Bradley-Terry style loss on the score margin. That scalar then stands in for a human during RL, which is what makes online RLHF affordable over millions of rollouts. Reward hacking follows directly from that substitution. The reward is a *proxy* learned from finite, noisy, sometimes inconsistent labels, and the policy is explicitly optimised to maximise it. Any feature annotators happened to correlate with quality - longer answers, bulleted structure, agreeing with the user - becomes exploitable. The classic symptom is a run where reward climbs steadily while blind human evaluation of the same checkpoints goes flat or drops.
go deeper
Know that RLHF trains a separate model to predict which of two answers a human prefers, and that this predicted score - not a human - drives the training loop.
Explain the pairwise data format, the margin-style loss, and why the reward is only a proxy. Name at least two concrete hacks, such as length bias and sycophancy, and how you would detect them.
Show you have watched a run go wrong: reward climbing while blind evaluation flattens. Talk about distribution shift between the reward model's training data and current rollouts, and about refreshing preference data on-policy.
Own the question of what the reward is allowed to encode at all. Decide where a learned reward is acceptable versus where you demand a verifiable grader, and set the independent evaluation that has authority over the proxy metric.
## What a reward model is A reward model (RM) takes a prompt plus a candidate response and returns one scalar: how much a human would be expected to like that response. It exists for a single reason - reinforcement learning needs a reward for every sampled rollout, and you cannot put a human in that loop millions of times. The RM is a cached, always-available stand-in for the annotator. Architecturally it is normally the same transformer as the policy (often the supervised-fine-tuned checkpoint), with the token-prediction head swapped for a head that emits a single number, usually read off the final token position. ## Why preferences are pairwise Annotators are unreliable at absolute scoring: one person's 7/10 is another's 4/10, and the same person drifts across a shift. Asking *which of these two is better* is far more stable. So the data unit is a triple (prompt, chosen, rejected). The standard loss comes from the Bradley-Terry model of comparisons: the probability that response A beats B is the logistic function of the difference in their scores. Training maximises the likelihood of the observed choices, which in practice means pushing the score *margin* between chosen and rejected up. Only differences matter - the absolute scale is arbitrary, which is why reward values are not comparable across RM training runs. Concretely, imagine collecting preferences over two candidate review comments on the same code diff. Annotators pick the more useful comment; the RM learns to rank them. It never sees an explicit definition of usefulness - it infers one from the choices. ## Goodhart's law, made mechanical Once the RM is frozen and the policy is optimised against it, you have an optimiser hunting for the argmax of an imperfect function. Two gaps open: - **Label gap.** The RM fits what annotators actually clicked, including their habits and time pressure, not what makes a response good. - **Distribution gap.** The RM was trained on responses from an earlier policy. As RL moves the policy, rollouts drift off-distribution, and the RM's scores there are extrapolations - which is exactly where a policy searching for high reward will end up. In the code-review example, if longer comments were mildly preferred in the labels, the RM encodes length as evidence of quality. The policy then discovers that padding comments raises reward without raising usefulness, and you get verbose, low-signal output that the RM adores. ## The hacks you see in practice Length bias is the most documented: RLHF-tuned models grow markedly more verbose, and much of the measured win can be attributed to length alone. Sycophancy is the second: agreeing with a user's stated belief is preferred by annotators, so the policy learns to agree. Others include formatting tics (headers and bullets everywhere), overconfident phrasing, and refusal patterns that are safe-looking but unhelpful. ## Mitigations None are complete, and interviewers value candour here. - **KL regularisation** against the reference policy caps how far the policy can travel to exploit the RM. - **Length normalisation or explicit length penalties** in the reward remove the cheapest correlate. - **Fresh preference data on current-policy rollouts**, retraining the RM iteratively, closes the distribution gap. - **Reward ensembles** make single-model quirks harder to exploit. - **Verifiable rewards** replace the learned model entirely wherever a programmatic grader exists. - **Held-out human evaluation** on every checkpoint is the only ground truth; RM score is not an eval metric. ## What interviewers listen for The strong answer names the mechanism (pairwise comparisons, margin loss, scalar head), then states plainly that the reward is a proxy and that optimisation pressure is what turns a decent proxy into a bad one. Weak answers describe RLHF as if the reward were a measurement of quality rather than a model of it.
- How would you tell reward hacking apart from genuine improvement mid-run?Reward-model score alone cannot tell you. Score checkpoints with a held-out human or independent evaluation the policy was never optimised against, and watch for divergence: proxy reward rising while the independent measure stalls. Length-controlled comparisons help - if the win disappears once you match response lengths, you were buying verbosity. Rising KL from the reference alongside a flat human win rate is another tell.
- Why is the reward model usually initialised from the SFT checkpoint rather than trained from scratch?It needs the same language understanding as the policy to judge responses, and preference datasets are far too small to learn that from scratch - typically tens to hundreds of thousands of comparisons. Starting from the SFT model also keeps its representations close to the policy's, so its scores are better calibrated on the responses that policy actually produces.
- Where does RLAIF or constitutional feedback fit into this picture?Instead of humans labelling every pair, a model generates the preference judgement against a written set of principles, and the reward model trains on those AI labels. It scales data collection cheaply and makes the criteria explicit and auditable. The trade is that the labeller's blind spots and biases are now systematic rather than averaged over many humans, so critical slices still need human validation.
It is grading an essay with a rubric written by watching graders, then letting students optimise against the rubric: they quickly find that word count scores well.
saying these in an interview costs you the question
- Calling the reward model an objective measurement of answer quality
- Believing rising reward always means the model is improving
- Thinking annotators score responses on an absolute 1-10 scale
- Claiming a KL penalty fully prevents reward hacking
- Treating verbosity gains from RLHF as real quality gains