When does distilling a frontier model into a small tuned model beat prompting it?
answer
- a unit-economics move, not a quality move
- student approaches, rarely exceeds, the teacher
- needs a narrow, stable, high-volume task
- serve a cascade with an escalation tier
- re-measure the crossover, prices moved
basics
~20 sDistillation pays when quality is already acceptable and cost or latency is the binding constraint: a stable, high-volume task where a frontier model's per-call price dominates. Its outputs become training data for a small model that serves the routine tier.
solid answer
~50 sDistillation is a unit-economics move, not a quality move. The preconditions are that a large model already solves the task well enough to be a teacher, the task is narrow and stable, and volume is high enough that per-call cost — or tail latency — is what actually hurts. You collect the teacher's inputs, reasoning and outputs on real traffic, filter them, and train a small strong open-weight base on them, usually with a lightweight adapter rather than a full-parameter run. The student typically approaches but does not exceed the teacher, so you keep an escalation path: route the routine tier to the small model and hard or high-stakes cases to the large one. Retrieval stays in place either way. As of mid-2026 the crossover has moved: prompt caching and genuinely cheap small hosted models mean you should re-measure rather than assume distillation wins.
go deeper
Know that distillation means training a small model on a bigger model's outputs, and that the point is cheaper and faster serving rather than better answers.
Explain the preconditions — teacher already good, task narrow and stable, volume high — and why the student approaches but does not exceed the teacher's quality on the task.
Demonstrate the operating view: trace collection and filtering, a stratified eval slice, a cascade with an escalation tier and a confidence signal, and re-deriving the cost crossover against the configuration you would actually ship.
Own the economics and the exit. Justify the spend with a real monthly bill at both price points, decide what the team stops doing to maintain the pipeline, and set a standing re-measurement so a cheaper base model can retire the tune instead of outliving it.
## What distillation means here In this decision, distillation means using a strong model's outputs as the training data for a smaller model that you then serve. You are not trying to make the small model smarter than its teacher; you are trying to buy most of the teacher's task-specific competence at a fraction of the per-call cost and latency. It is the fourth rung of the prompt → retrieval → fine-tune → distill ladder, and it is the rung that most often has a clear, calculable business case. ## The preconditions **The teacher already works.** If a frontier model cannot do the task acceptably, distillation has nothing to copy. Fix quality first; a student trained on mediocre traces inherits the mediocrity and adds its own. **The task is narrow and stable.** A small model gives up general capability. That is only safe when the job is bounded — classify this alert, extract these fields, draft this memo type — and when the specification is not going to change next month, because every change means regenerating traces and retraining. **Volume makes cost the binding constraint.** At a few thousand calls a month, engineering time dwarfs inference cost and distillation is a distraction. At millions of calls a month it can be the difference between a programme that ships and one that does not. Compute the actual monthly bill at both price points before proposing it. **You can legitimately use the teacher's outputs.** Provider terms differ on training competing models from their outputs. This is a real gating condition, not a footnote. ## The shape of the project Collect traces from production or from a representative input set: the input, the teacher's reasoning where it is exposed, and the final output. Filter aggressively — a judge model or programmatic checks to drop wrong, malformed or degenerate rows — and keep an eye on diversity, because traces sampled from real traffic skew toward whatever is most common and the student inherits that skew. Hold out an evaluation slice that never touched training. Then train a small strong open-weight base. In 2026 the usual configuration is a lightweight adapter on a capable small model rather than a full-parameter run; the adapter approach reaches comparable quality at a fraction of the compute, which matters because you will do this again every time the task or the base moves. ## Serving the result The realistic deployment is a cascade, not a replacement. Route the routine tier to the student and keep an escalation path to the teacher for inputs the student is unsure about, for high-stakes decisions, or for anything a confidence signal flags. A KYC operation running three million alerts a month might send the overwhelming bulk of clean, low-risk alerts to the tuned small model and everything with adverse-media or complex-ownership signals to the frontier model. The cost saving comes from the volume tier; the risk protection comes from the escalation tier. Retrieval does not go away. The student is smaller and therefore *more* dependent on being handed the relevant records rather than recalling them. ## Why the crossover keeps moving As of mid-2026 two things have made the naive "small model is cheaper, therefore distil" argument weaker than it was. Caching of repeated prompt prefixes has cut the marginal cost of a long, stable system prompt on a large model substantially. And the price of capable small hosted models has fallen far enough that the gap you are arbitraging is narrower than the headline flagship price suggests. Measure against the actual configuration you would ship — cached prompt, cheapest adequate hosted model — not against the flagship list price. ## What you are signing up for A distilled model is a system you now operate. You own the trace pipeline, the filtering, the eval suite, the serving path for the small model, and the decision of what to do when a materially better base ships — which, at current cadence, is every few months. Many teams find that a new small base plus a good prompt matches last quarter's distilled model with none of the maintenance. Budget a periodic re-measurement rather than assuming the tune stays ahead. ## Answering well Lead with "it is a cost decision, not a quality decision", give the preconditions, describe the cascade rather than the replacement, and mention that you would re-derive the crossover with current prices instead of quoting a rule of thumb.
- How do you decide which requests the small model handles and which escalate?Route by risk and by a confidence signal, not by trying to make one model do everything. Send the high-volume, low-stakes, in-distribution slice to the student; escalate anything the student flags as uncertain, anything above a business risk threshold, and a random sample for continuous comparison. That sample is what tells you whether the student is drifting away from the teacher before customers do.
- What goes wrong if you build the training set purely from sampled production traffic?You inherit the traffic distribution, including its imbalance. Rare-but-important cases are underrepresented, so the student is weakest exactly where the cost of being wrong is highest, and any systematic bias in the teacher's behaviour on those cases is amplified. The fix is to stratify — deliberately oversample rare classes and hard cases — and to keep an evaluation slice that reflects importance rather than frequency.
- A stronger small base model ships six months after you distil. What do you do?Re-measure before retraining. Prompt the new base on your existing evaluation set first; teams regularly find it matches or beats last quarter's distilled model with no training at all, at which point the right move is to retire the tune and reclaim the maintenance. If the tune still wins by a margin that justifies the pipeline, regenerate traces against the current teacher and retrain on the new base.
saying these in an interview costs you the question
- Distillation makes the small model better than the teacher
- Any high-volume task is a distillation candidate regardless of stability
- Assuming the flagship list price when computing the saving
- Replacing the large model entirely with no escalation path
- Dropping retrieval because the student was trained on the domain