skip to content

Preference and RL Post-Training

Post-training shapes refusals, tone and reasoning: GRPO against verifiable rewards today, DPO when preference pairs already exist. Interviewers use it to probe how deeply you read model behavior.

on this pageshow

questions

6

When would you choose offline DPO over an online RL run like GRPO?

level: middleimportance: must knowfreq 58%

answer

  1. the policy is its own implicit reward
  2. one loss over fixed preference triples
  3. no sampling, no reward model
  4. the data describes a model you have left behind
  5. its temperature is the KL knob

basics

~20 s

Choose DPO when good preference pairs already exist and the goal is bounded - tone, format, a known bad habit. It needs no sampling loop, no reward model and no rollout infrastructure. Online RL wins when the reward is verifiable or the model must be pushed well past its current behaviour.

solid answer

~50 s

Direct preference optimisation removes the two heaviest parts of the RLHF pipeline. It reformulates preference learning so the policy is its own implicit reward model, and trains with a classification-style loss on fixed (prompt, chosen, rejected) triples. No reward model is trained, no responses are sampled during training, and the run looks like ordinary supervised fine-tuning with a frozen reference model alongside. That makes it the pragmatic default when preference data already exists and the behaviour change is modest: adopt a house style, stop a specific failure mode, prefer one answer shape over another. It is fast, reproducible and cheap. The weakness is that it is offline. The pairs were produced by some other policy, so as training moves the model, the data increasingly describes a model that no longer exists. Online methods keep sampling from the current policy and grading fresh, which is why verifiable-reward reasoning training and large behavioural shifts still use them. Iterating DPO - generate, label, retrain - recovers part of the gap.

code

python · 10 lines
python
import math


def dpo_loss(logp_chosen, logp_rejected, ref_chosen, ref_rejected, beta=0.1):
    margin = beta * ((logp_chosen - ref_chosen) - (logp_rejected - ref_rejected))
    return -math.log(1.0 / (1.0 + math.exp(-margin)))


print(dpo_loss(-12.0, -14.0, -12.5, -13.0))
print(dpo_loss(-14.0, -12.0, -12.5, -13.0))

go deeper

for a junior

Know that DPO trains directly on pairs of better and worse answers with a single loss, skipping the separate reward model and the sampling loop that classic RLHF needs.

for a middle

Explain that the policy acts as its own implicit reward model, that a frozen reference is still required, and that the temperature parameter plays the role of the KL coefficient.

for a senior

Discuss the offline distribution gap and its consequences - no exploration, stale pairs, margins widened by suppressing the chosen side - and when you would iterate with on-policy regeneration instead.

for a principal

Frame it as an infrastructure and data decision: whether you have preference pairs or a grader, whether the goal is behavioural polish or genuine capability gain, and be candid that published comparisons do not settle the question.

## The idea behind DPO Standard RLHF is two stages: fit a reward model to preferences, then run RL against it with a KL penalty. DPO's derivation observes that the optimal policy for that KL-penalised objective has a closed-form relationship to the reward, which can be inverted - the reward can be written in terms of the policy and the reference. Substituting that expression into the preference likelihood removes the reward model from the picture entirely. What is left is a loss over triples of prompt, chosen response and rejected response. It increases the policy's log-probability of the chosen response relative to the reference, decreases it for the rejected one, and passes the difference through a logistic function scaled by a temperature parameter. Mechanically it is a binary classification loss on a margin. That temperature plays the role the KL coefficient plays in online RLHF: small values allow the policy to move far from the reference, large values keep it conservative. ## What you gain - **No reward model.** One fewer model to train, validate and store, and one fewer thing to be hacked. - **No sampling loop.** Training is a pass over a fixed dataset - no generation infrastructure, no rollout scheduling, no grader service. - **Stability and reproducibility.** It behaves like supervised fine-tuning: same data plus same seed gives the same result, and the failure modes are the familiar ones. - **Cost.** Often an order of magnitude cheaper in wall clock and engineering effort than standing up an online RL stack. ## What you give up **Off-policy data.** The preference pairs were generated by some earlier model. Early in training the loss reflects the current policy reasonably well; as the policy shifts, the gradient is increasingly computed on responses it would never produce. Online methods have no such gap because every rollout comes from the current policy. **No exploration.** DPO can only rank behaviours present in the dataset. It cannot discover a strategy nobody wrote down, which is exactly what RL against a verifiable grader does when it finds a better reasoning path. **Sensitivity to pair quality.** With only a fixed dataset, noisy or inconsistent pairs go straight into the objective with no sampling to average over. Contradictory pairs actively fight each other. **Degenerate solutions.** The loss trains a *margin*, and a margin can be widened by pushing the rejected response's probability down rather than the chosen one's up. Runs can end up reducing the likelihood of both, which shows up as unnatural or degraded generations. Length imbalance between chosen and rejected is a common confound, since log-probabilities scale with token count. ## Variants worth naming **ORPO** removes the reference model too, combining a supervised fine-tuning loss with an odds-ratio penalty on the rejected response so preference alignment happens in a single stage rather than after a separate SFT run. That saves the memory of the frozen reference and one pipeline step, at the cost of the explicit reference anchoring. **Iterative or online DPO** closes the distribution gap partway: sample from the current policy, obtain preferences over those fresh samples from humans or a judge, retrain, repeat. It sits between pure offline and full online RL in both cost and capability. ## Choosing between them A decision that holds up in an interview: - Preference pairs already exist, the target change is a matter of style, tone, format or a specific unwanted behaviour, and you want a result this week - **DPO**. - The reward is programmatically checkable and you want the model to get genuinely better at a task rather than merely to prefer a different phrasing - **online RL**, group-relative. - You have labelling capacity and a real gap after one DPO round - **iterate**: regenerate on-policy, relabel, retrain. Be honest that this is contested ground: reported comparisons vary with data quality, tuning effort and task, and it is entirely reasonable to say the deciding factor is usually what infrastructure and data you already have.

  • DPO training drives down the probability of the chosen responses as well as the rejected ones. Why can that happen?
    The loss optimises the gap between chosen and rejected relative to the reference, and a gap widens just as well by suppressing the rejected side hard. If both fall while the margin grows, the objective is satisfied but generation quality degrades. Mitigations include raising the temperature parameter so the policy stays nearer the reference, adding a supervised term on the chosen responses, and balancing response lengths within pairs.
  • What does DPO's temperature parameter actually control?
    It scales the implicit KL constraint against the reference model. Small values let the policy move far to satisfy preferences, which increases both the achievable behaviour change and the risk of degradation; large values keep it close to the reference and produce a conservative, barely-moved model. It is the offline analogue of the KL coefficient in an online RLHF run and deserves the same care.
  • How does ORPO differ from running SFT and then DPO?
    ORPO merges the two stages: it trains on the chosen responses with a normal supervised loss while adding an odds-ratio term that penalises the rejected ones, and it needs no frozen reference model. That removes a pipeline stage and the reference's memory footprint. The trade is losing explicit anchoring to a reference checkpoint, so preservation of prior behaviour is less directly controlled.

saying these in an interview costs you the question

  • Saying DPO has no KL constraint at all
  • Claiming DPO is strictly better than online RL because it is simpler
  • Thinking DPO can discover behaviours absent from the preference data
  • Forgetting DPO still needs a frozen reference model
  • Ignoring that offline pairs go stale as the policy moves

context

open as a page

Why does GRPO drop the value critic that PPO-style RLHF requires?

level: middleimportance: must knowfreq 62%

basics

~20 s

PPO-style RLHF trains a second network to predict expected return as a baseline. GRPO samples a group of responses per prompt instead and uses the group's own mean reward as that baseline, removing a whole model from memory and from the failure surface.

open as a page

How is a reward model trained, and why does RLHF reward hacking happen?

level: middleimportance: must knowfreq 70%

basics

~20 s

A reward model learns from human-ranked response pairs to score the preferred one higher. Because it only approximates human taste, optimising hard against it finds cheap correlates - length, confident tone, flattery - instead of genuine quality.

open as a page

How do you tune the KL coefficient in an RLHF run, and what breaks at each extreme?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The KL term penalises drift from the frozen reference policy, trading reward against staying recognisably like the starting model. Too low and the policy collapses into reward-hacked, degenerate text; too high and it barely moves. Tune by watching measured KL, not the coefficient.

open as a page

When can a verifiable reward replace a learned reward model in post-training?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Whenever a program can decide correctness - unit tests passing, a checked final answer, a schema validating - the grader becomes the reward directly. This removes the proxy that reward hacking exploits, but only works where correctness is mechanically checkable and the grader itself is hard to game.

open as a page

What is the alignment tax, and how would you measure it across post-training?

level: principalimportance: should knowfreq 36%

basics

~20 s

Alignment tax is the capability lost when a model is made helpful and well-mannered: a post-preference checkpoint can be more usable yet score lower on raw benchmarks than its base. Measuring it means running a fixed capability suite against every checkpoint in the pipeline.

open as a page