When would you choose offline DPO over an online RL run like GRPO?
answer
- the policy is its own implicit reward
- one loss over fixed preference triples
- no sampling, no reward model
- the data describes a model you have left behind
- its temperature is the KL knob
basics
~20 sChoose DPO when good preference pairs already exist and the goal is bounded - tone, format, a known bad habit. It needs no sampling loop, no reward model and no rollout infrastructure. Online RL wins when the reward is verifiable or the model must be pushed well past its current behaviour.
solid answer
~50 sDirect preference optimisation removes the two heaviest parts of the RLHF pipeline. It reformulates preference learning so the policy is its own implicit reward model, and trains with a classification-style loss on fixed (prompt, chosen, rejected) triples. No reward model is trained, no responses are sampled during training, and the run looks like ordinary supervised fine-tuning with a frozen reference model alongside. That makes it the pragmatic default when preference data already exists and the behaviour change is modest: adopt a house style, stop a specific failure mode, prefer one answer shape over another. It is fast, reproducible and cheap. The weakness is that it is offline. The pairs were produced by some other policy, so as training moves the model, the data increasingly describes a model that no longer exists. Online methods keep sampling from the current policy and grading fresh, which is why verifiable-reward reasoning training and large behavioural shifts still use them. Iterating DPO - generate, label, retrain - recovers part of the gap.
code
python · 10 linesimport math
def dpo_loss(logp_chosen, logp_rejected, ref_chosen, ref_rejected, beta=0.1):
margin = beta * ((logp_chosen - ref_chosen) - (logp_rejected - ref_rejected))
return -math.log(1.0 / (1.0 + math.exp(-margin)))
print(dpo_loss(-12.0, -14.0, -12.5, -13.0))
print(dpo_loss(-14.0, -12.0, -12.5, -13.0))go deeper
Know that DPO trains directly on pairs of better and worse answers with a single loss, skipping the separate reward model and the sampling loop that classic RLHF needs.
Explain that the policy acts as its own implicit reward model, that a frozen reference is still required, and that the temperature parameter plays the role of the KL coefficient.
Discuss the offline distribution gap and its consequences - no exploration, stale pairs, margins widened by suppressing the chosen side - and when you would iterate with on-policy regeneration instead.
Frame it as an infrastructure and data decision: whether you have preference pairs or a grader, whether the goal is behavioural polish or genuine capability gain, and be candid that published comparisons do not settle the question.
## The idea behind DPO Standard RLHF is two stages: fit a reward model to preferences, then run RL against it with a KL penalty. DPO's derivation observes that the optimal policy for that KL-penalised objective has a closed-form relationship to the reward, which can be inverted - the reward can be written in terms of the policy and the reference. Substituting that expression into the preference likelihood removes the reward model from the picture entirely. What is left is a loss over triples of prompt, chosen response and rejected response. It increases the policy's log-probability of the chosen response relative to the reference, decreases it for the rejected one, and passes the difference through a logistic function scaled by a temperature parameter. Mechanically it is a binary classification loss on a margin. That temperature plays the role the KL coefficient plays in online RLHF: small values allow the policy to move far from the reference, large values keep it conservative. ## What you gain - **No reward model.** One fewer model to train, validate and store, and one fewer thing to be hacked. - **No sampling loop.** Training is a pass over a fixed dataset - no generation infrastructure, no rollout scheduling, no grader service. - **Stability and reproducibility.** It behaves like supervised fine-tuning: same data plus same seed gives the same result, and the failure modes are the familiar ones. - **Cost.** Often an order of magnitude cheaper in wall clock and engineering effort than standing up an online RL stack. ## What you give up **Off-policy data.** The preference pairs were generated by some earlier model. Early in training the loss reflects the current policy reasonably well; as the policy shifts, the gradient is increasingly computed on responses it would never produce. Online methods have no such gap because every rollout comes from the current policy. **No exploration.** DPO can only rank behaviours present in the dataset. It cannot discover a strategy nobody wrote down, which is exactly what RL against a verifiable grader does when it finds a better reasoning path. **Sensitivity to pair quality.** With only a fixed dataset, noisy or inconsistent pairs go straight into the objective with no sampling to average over. Contradictory pairs actively fight each other. **Degenerate solutions.** The loss trains a *margin*, and a margin can be widened by pushing the rejected response's probability down rather than the chosen one's up. Runs can end up reducing the likelihood of both, which shows up as unnatural or degraded generations. Length imbalance between chosen and rejected is a common confound, since log-probabilities scale with token count. ## Variants worth naming **ORPO** removes the reference model too, combining a supervised fine-tuning loss with an odds-ratio penalty on the rejected response so preference alignment happens in a single stage rather than after a separate SFT run. That saves the memory of the frozen reference and one pipeline step, at the cost of the explicit reference anchoring. **Iterative or online DPO** closes the distribution gap partway: sample from the current policy, obtain preferences over those fresh samples from humans or a judge, retrain, repeat. It sits between pure offline and full online RL in both cost and capability. ## Choosing between them A decision that holds up in an interview: - Preference pairs already exist, the target change is a matter of style, tone, format or a specific unwanted behaviour, and you want a result this week - **DPO**. - The reward is programmatically checkable and you want the model to get genuinely better at a task rather than merely to prefer a different phrasing - **online RL**, group-relative. - You have labelling capacity and a real gap after one DPO round - **iterate**: regenerate on-policy, relabel, retrain. Be honest that this is contested ground: reported comparisons vary with data quality, tuning effort and task, and it is entirely reasonable to say the deciding factor is usually what infrastructure and data you already have.
- DPO training drives down the probability of the chosen responses as well as the rejected ones. Why can that happen?The loss optimises the gap between chosen and rejected relative to the reference, and a gap widens just as well by suppressing the rejected side hard. If both fall while the margin grows, the objective is satisfied but generation quality degrades. Mitigations include raising the temperature parameter so the policy stays nearer the reference, adding a supervised term on the chosen responses, and balancing response lengths within pairs.
- What does DPO's temperature parameter actually control?It scales the implicit KL constraint against the reference model. Small values let the policy move far to satisfy preferences, which increases both the achievable behaviour change and the risk of degradation; large values keep it close to the reference and produce a conservative, barely-moved model. It is the offline analogue of the KL coefficient in an online RLHF run and deserves the same care.
- How does ORPO differ from running SFT and then DPO?ORPO merges the two stages: it trains on the chosen responses with a normal supervised loss while adding an odds-ratio term that penalises the rejected ones, and it needs no frozen reference model. That removes a pipeline stage and the reference's memory footprint. The trade is losing explicit anchoring to a reference checkpoint, so preservation of prior behaviour is less directly controlled.
saying these in an interview costs you the question
- Saying DPO has no KL constraint at all
- Claiming DPO is strictly better than online RL because it is simpler
- Thinking DPO can discover behaviours absent from the preference data
- Forgetting DPO still needs a frozen reference model
- Ignoring that offline pairs go stale as the policy moves