What mix of instruction pairs, preference pairs and distilled traces belongs in a fine-tuning set?
answer
- three shapes, three different jobs
- form first, ranking afterwards
- preference volume is far smaller
- sample both candidates on-policy
- teacher mistakes become training targets
basics
~20 sInstruction pairs teach the task's form and are the bulk of the data. Preference pairs come afterwards, in far smaller volume, for qualities expressible only as a comparison. Distilled teacher traces supply reasoning you cannot afford to write by hand.
solid answer
~60 sThree supervision shapes, three jobs. **Instruction pairs** — an input and the one output you want — teach form, structure and task procedure, and they are where most of the volume goes because that is what a fine-tune is actually good at. **Preference pairs** — the same prompt with a chosen and a rejected response, consumed by DPO-family methods — teach ranking between outputs that are both acceptable, which is how you express calibrated hedging or the right amount of caution when no single target sentence captures it. They come second, and typically a fraction of the SFT volume; build them by sampling two responses per prompt from the model you are about to tune, having a judge pick a winner, and discarding ties and near-ties, which carry no gradient signal and plenty of noise. **Distilled traces** from a stronger teacher — input, reasoning, output — buy you process supervision cheaply, at the price of inheriting the teacher's mistakes and accepting its output terms. Where the final answer is verifiable, keep only the traces that got it right.
code
json · 1 line{"prompt": "My cat has not eaten for two days and is hiding under the bed.", "chosen": "Two days without food in a cat is urgent - hepatic lipidosis can develop quickly. Book a same-day appointment and note any vomiting.", "rejected": "Cats often skip meals when stressed. Try warming the food and see how it goes over the next few days."}go deeper
Know the three shapes by name — instruction pairs, preference pairs, distilled traces — and that instruction pairs are the bulk of a typical fine-tuning set.
Explain what each shape teaches and why instruction tuning comes first: preference data moves a model between behaviours it can already produce, so it cannot establish the output form from scratch.
Show the construction judgement — on-policy sampling of candidates, discarding ties, validating the adjudicator, and rejection-sampling distilled traces on a verifiable final answer.
Own the composition as a programme decision with costs attached: added stages, judge error, teacher terms of use and inherited style, and be able to justify each shape by the specific deficiency it exists to fix.
## The three shapes A fine-tuning dataset is not one kind of thing, and the composition question is a genuine design decision rather than a formatting detail. **Instruction pairs (SFT data).** One input, one target output. This is the workhorse. It teaches the model the shape of the task: what fields to emit, in what order, at what length, in what register, following what procedure. Most projects are entirely this, and most projects should be — form is what fine-tuning teaches most reliably. **Preference pairs.** A prompt with two responses, one marked chosen and one rejected. Direct preference optimisation and its descendants consume this shape directly. Preference data expresses something instruction pairs structurally cannot: a ranking over acceptable outputs. When the property you want is "do not over-reassure an anxious owner, but do not alarm them either", nobody can write the single target sentence — but anybody can look at two responses and say which is better. **Distilled traces.** Triples of input, intermediate reasoning, and output, generated by a stronger teacher model. These supply not just the answer but the path to it, which is expensive to author by hand and which is how much of the recent wave of small strong reasoning models was built. ## Ordering, and why it is not negotiable SFT first, preference second. Preference optimisation moves the model between behaviours it can already produce; it is a poor tool for teaching a behaviour from scratch. If the model does not yet emit the right output shape, both members of every pair are wrong and the comparison signal is wasted. Get form right with instruction pairs, then use preference data on the residue — tone, calibration, refusal boundaries, verbosity, the choices between two defensible answers. ## Volumes There is no universal ratio, and anyone offering one confidently is guessing. What holds: - The SFT stage carries the volume, because it is doing the heavy teaching — hundreds to tens of thousands of rows depending on how broad the behaviour is. - The preference stage is typically much smaller. It is a nudge on an already-competent model, and it is expensive per row because each row needs two generations plus an adjudication. - Distilled traces sit inside the SFT volume rather than beside it; they are instruction pairs with reasoning attached. The honest framing for an interview is that you would size the preference stage by whether SFT alone left a gap you can articulate, not by a target ratio. ## Building preference pairs well - **Sample on-policy.** Generate both candidates from the model you are about to tune, at the settings it will run at. Pairs sampled from some other model teach the ranking of *that* model's output distribution, which is not the distribution you are correcting. - **Judge, then discard ties.** An adjudicator — human or model — picks the winner. Ties and near-ties should be dropped: they contribute noise rather than signal, and a pair whose two members are indistinguishable is training on a coin flip. - **Watch the margin.** Pairs where one side is obviously broken are easy and teach little beyond what SFT already covered. The informative pairs are the ones where both responses are plausible. - **Validate the adjudicator.** The judge's verdicts are the labels, so the usual concern applies: check its agreement with an expert on a sample before you trust thousands of its decisions. ## Distilled traces: what they cost Teacher-generated data is the cheapest way to get process supervision, and it comes with three bills. **Error propagation.** The teacher's mistakes become training targets, and a plausible-sounding wrong reasoning chain is worse than an obviously wrong answer because it survives casual review. Where the final answer is checkable — a computed value, a category from a fixed set, code that either passes tests or does not — keep only traces whose final answer is correct. That rejection-sampling step is the single highest-value filter available on distilled data. Where nothing is verifiable, you are relying on a judge and should size your confidence accordingly. **Style transfer you did not ask for.** A distilled student inherits the teacher's voice, its hedging habits and its formatting tics along with its competence. If your product has a defined voice, that is a conflict to resolve deliberately in the data, not to discover after training. **Terms of use.** Whether the teacher's provider permits using its outputs to train a competing model is a real constraint that varies by provider and changes over time. This is a question to answer before building the pipeline, not after, and "we did not check" is a poor answer at a lead level. ## A fourth shape worth naming For tasks with a programmatically verifiable outcome, reinforcement fine-tuning against a grader needs no labelled outputs at all — you supply prompts and a scoring function, and the training loop generates its own attempts. The dataset question becomes "do I have a reliable grader and a good spread of prompts" rather than "can I write the answers". Recognising when a task falls into that category changes what you spend your curation budget on entirely. ## How to decide, in practice Start with instruction pairs and see how far they go. Add preference pairs only for a named deficiency that you can state as a comparison and cannot state as a target. Use distilled traces where reasoning is the thing you lack and where you can filter on a verifiable outcome. And keep the composition documented, because six months later the question "why does this model hedge like that" is answered by the data mix, not by the training code.
- Why sample both preference candidates from the model you are about to tune?Because preference optimisation is correcting that model's output distribution. On-policy pairs show the adjudicator the choices this model actually makes, so the gradient pushes probability mass between behaviours it really produces. Pairs sampled from a stronger model teach the ranking of a distribution you are not training, and often produce pairs with such a large margin that the comparison carries little information beyond what SFT already taught.
- How do you handle the teacher's mistakes in distilled reasoning traces?Filter on the outcome wherever the outcome is checkable — keep only traces whose final answer is verifiably right, discard the rest, and regenerate. That rejection-sampling step removes most of the damage cheaply. Where nothing is verifiable you are down to judge-based filtering with its own error rates, and the plausible-but-wrong chain is the dangerous case, because it reads well and survives casual review.
- When would you skip preference data entirely?When you cannot name the deficiency it is meant to fix. If SFT already produces outputs that domain reviewers accept, preference pairs add cost, a second training stage and a new source of judge error for no stated gain. It earns its place when the remaining problem is a ranking one — calibration, tone, how much to hedge, which of two defensible answers to prefer — that no single target output can express.
- What changes if the task has a programmatic correctness check?Then reinforcement fine-tuning against a grader becomes an option, and the curation problem changes shape: you need a spread of prompts and a reliable scoring function rather than written answers. Budget shifts from authoring outputs to hardening the grader, because a grader that can be satisfied the wrong way will be, and the model will find that path faster than a reviewer will.
saying these in an interview costs you the question
- Preference pairs can replace the instruction-tuning stage
- There is one correct ratio between the three shapes
- Ties in preference pairs should be kept for volume
- Sampling both candidates from a frontier model is best
- Distilled teacher traces need no correctness filtering
- Teacher output terms of use are a legal detail to sort out later