When aligning a fine-tuned Llama, how does DPO differ from PPO-based RLHF?
answer
- one pipeline has a scalar-head model, one does not
- preference pairs versus an online sampling loop
- the policy is its own implicit reward model
- beta plays the KL role
- offline learning cannot explore
basics
~20 sDPO trains directly on chosen/rejected preference pairs with a classification-style loss against a frozen reference model. PPO-based RLHF first trains a separate reward model, then optimizes the policy by sampling completions and scoring them — far more moving parts.
solid answer
~50 sClassic RLHF is a three-stage pipeline: supervised fine-tuning, then a **reward model** trained on human preference comparisons, then **PPO**, which samples completions from the policy, scores them with the reward model, and updates the policy with a KL penalty pulling it toward the SFT reference. It works but is operationally heavy — four models in memory, an online sampling loop, and notorious instability. **DPO** derives a closed-form objective that makes the optimal policy its own implicit reward model, so you skip the reward model and the sampling loop entirely: you feed triples of (prompt, chosen, rejected), compute log-probabilities under the policy and a frozen reference, and minimise a logistic loss with a **beta** parameter playing the role of the KL strength. It is far easier to run, which is why it is the default for Llama alignment; Meta's own Llama 3 post-training used SFT plus rejection sampling and DPO. PPO's advantage is online exploration — it can improve beyond the completions in a static preference dataset.
code
python · 17 linesfrom trl import DPOTrainer, DPOConfig
args = DPOConfig(
beta=0.1,
learning_rate=5e-6,
num_train_epochs=1,
output_dir="llama-dpo",
)
trainer = DPOTrainer(
model=model, # a PEFT-wrapped SFT checkpoint
ref_model=None, # adapter disabled = implicit reference
args=args,
train_dataset=pairs, # columns: prompt, chosen, rejected
processing_class=tokenizer,
)
trainer.train()go deeper
Know that after supervised fine-tuning, preference data of the form chosen-versus-rejected is used to shape behaviour, and that DPO is the simpler modern way to consume it.
Explain the three-stage RLHF pipeline against DPO's single loss, what the reference model is for, and why DPO needs no reward model or generation loop.
Show operational judgment: tuning beta, spotting length bias in preference data, using an implicit reference with a LoRA policy, keeping learning rates low and epochs few, and evaluating generations rather than loss.
Own the strategy — whether to invest in the annotation pipeline and infrastructure PPO demands, when iterative or online preference collection pays for itself, and how alignment choices bind you to a specific base model over time.
## The alignment problem after SFT Supervised fine-tuning teaches a Llama to imitate good answers. It cannot teach it that answer A is *better* than answer B, because imitation has no notion of comparison. Preference optimisation fills that gap: it consumes human (or model) judgements of the form "for this prompt, this response is preferred to that one" and shifts the model's distribution toward the preferred style, helpfulness and refusal behaviour. ## The PPO-based RLHF pipeline 1. **SFT** — fine-tune the base Llama on demonstration data. This becomes the reference policy. 2. **Reward model** — take a copy of the model, replace the language-modelling head with a scalar head, and train it on preference pairs with a ranking loss so that it assigns a higher score to the chosen response. 3. **PPO** — sample completions from the current policy, score them with the reward model, and update the policy with Proximal Policy Optimization, adding a per-token KL penalty against the frozen reference so the policy cannot drift into degenerate text that games the reward model. At peak this holds four models: the policy, the reference, the reward model and (classically) a value/critic model. It requires online generation inside the training loop, which is slow, and it is sensitive to the KL coefficient, the learning rate and reward-model quality. **Reward hacking** — the policy discovering text that scores well but is bad — is the characteristic failure, and it is why the KL term exists at all. ## What DPO changes DPO starts from the observation that the RLHF objective has a closed-form optimal policy, and that this relation can be inverted: the reward implied by any policy relative to a reference is proportional to `beta * log(pi(y|x) / pi_ref(y|x))`. Substituting that into the preference-model likelihood gives a loss you can optimise directly on preference pairs with no reward model and no sampling. Concretely, per training example you compute four log-probability sums — chosen and rejected, under the policy and under the reference — and minimise a logistic loss that wants the policy's margin between chosen and rejected to exceed the reference's. The consequences are practical: - **Two forward passes instead of a generation loop.** Training looks like ordinary supervised training; step time is predictable. - **No reward model to train, tune or debug.** One fewer artifact and one fewer source of reward hacking. - **The reference model can be free.** If you are training a LoRA adapter, disabling the adapter recovers the reference policy, so you do not need a second copy in memory at all — in TRL you pass `ref_model=None` with a PEFT model. - **beta replaces the KL coefficient.** Low beta (0.01-0.05) lets the policy move far from the reference and risks degeneration; high beta (0.3+) keeps it close and may barely change behaviour. 0.1 is the usual starting point. ## Where DPO is weaker DPO is **offline**: it can only learn from the responses present in your preference dataset. PPO generates fresh completions from the current policy and gets them scored, so it can discover behaviour that no annotator wrote down and keeps receiving signal as the policy moves. That is a real advantage on long-horizon or verifiable tasks, and it is why online and iterative variants exist — regenerating preference data from the updated policy between DPO rounds, or using rejection sampling to build a better dataset each iteration, both of which recover part of the exploration benefit at a fraction of PPO's complexity. On tasks with a programmatic correctness signal, reward-based methods that verify outputs rather than compare them are the growing alternative. ## Data and practical setup A DPO dataset is a table of `prompt`, `chosen`, `rejected`. Quality dominates quantity: pairs where chosen and rejected differ on the axis you care about teach that axis, while pairs that differ on length teach the model to be verbose — length bias is the classic DPO artifact, and length-normalised variants exist to counter it. Learning rates are much lower than SFT (around 5e-7 to 5e-6 for full models), and one epoch is often enough; over-training DPO reliably degrades fluency. Always run DPO **after** SFT, not instead of it — the reference policy needs to already be in the right format regime, otherwise the log-probability ratios are dominated by formatting rather than preference. ## What to say in an interview Frame it as a complexity-versus-exploration tradeoff. DPO is the default because it is a supervised-shaped job with one hyperparameter that matters; PPO earns its complexity only when online exploration or a live reward signal genuinely buys you something the static dataset cannot.
- Why can ref_model be None when the policy is a LoRA adapter?Because disabling the adapter recovers the frozen base exactly, and the frozen base is the reference policy. TRL runs the reference forward pass with the adapter switched off instead of loading a second full model, which halves the memory the DPO step needs. With a fully fine-tuned policy you have no such trick and must hold a separate reference copy.
- What does beta control in DPO, and what happens at the extremes?Beta sets how strongly the policy is anchored to the reference — it plays the role PPO's KL coefficient plays. Too low and the policy drifts far from the reference, often degenerating into repetitive or over-confident text; too high and preference signal barely moves the model. Around 0.1 is the usual starting point, tuned against held-out generations rather than loss.
- Your DPO-tuned Llama got noticeably more verbose. What likely happened?Length bias in the preference pairs. If the chosen response is systematically longer than the rejected one, the cheapest way to satisfy the objective is to be longer, so the model learns verbosity rather than quality. Fix it in the data — balance lengths, or use a length-normalised DPO variant — not by lowering the learning rate.
- When is PPO-based RLHF still worth its extra machinery?When online exploration matters: the model must discover behaviours no annotator wrote down, or you have a live reward signal — a verifier, tests, a rubric grader — that can score fresh samples. A static preference set caps what DPO can teach, whereas PPO keeps generating and scoring as the policy moves, which is what long-horizon and verifiable tasks need.
saying these in an interview costs you the question
- Says DPO trains a reward model like RLHF does
- Claims DPO needs an online sampling loop
- Treats beta as a learning rate rather than a KL-strength knob
- Runs preference tuning without doing SFT first
- Assumes DPO can explore behaviour absent from the dataset