skip to content

When is fine-tuning Llama the wrong tool compared with prompting plus RAG?

level: principalimportance: must knowfreq 50%

answer

  1. form versus facts
  2. weights cannot be edited on Tuesday
  3. retrieved answers can cite
  4. the base model will be replaced
  5. a dataset is a maintained asset

basics

~20 s

Fine-tuning teaches form — output shape, tone, task convention, refusal behaviour. It is a poor way to install facts that change, because updating them means retraining. Retrieval handles changing knowledge; fine-tuning handles how the answer is produced.

solid answer

~50 s

The decision rule that survives contact with production: **fine-tune for behaviour, retrieve for knowledge.** If the model already knows the material but keeps formatting it wrong, drifting in tone, ignoring your schema, or answering a specialised task in a generic way, a few thousand well-curated examples fix it and often let you drop to a smaller Llama. If the failure is that the model does not know your documents — or knows a version of them that changed last Tuesday — fine-tuning is the expensive wrong answer: facts baked into weights cannot be updated without another training run, and they cannot be cited. Fine-tuning also carries ongoing cost that a prompt does not: an evaluation set you must maintain, a retraining obligation every time the base model is upgraded, and adapter versioning at serving time. The honest sequence is prompt first, add retrieval, and fine-tune only when you have a measured gap that neither closed — at which point the two combine rather than compete.

go deeper

for a junior

Be able to say that fine-tuning shapes how the model answers while retrieval supplies what it answers with, and that facts which change should be retrieved, not trained in.

for a middle

Give concrete examples of behaviour-shaped failures that tuning fixes — schema violations, tone drift, task conventions — and explain why memorised facts cannot be cited or updated.

for a senior

Show that you sequence the decision from a measured failure mode: baseline prompt, add retrieval, tune only for a persistent behavioural gap, and evaluate for regressions outside the target task.

for a principal

Own it as a commitment, not a task — the curation and evaluation headcount, the retraining obligation at every base-model upgrade, the serving and rollback story, and the cost case for a tuned smaller model against a prompted larger one.

## Why this question is asked Every team that discovers open-weight Llama proposes fine-tuning within a week, usually to "teach it our documentation." The interviewer is testing whether you know that this is the one job fine-tuning is worst at, and whether you can reason about the total cost of an adapter rather than just its training bill. ## What fine-tuning is genuinely good at - **Output format.** Consistently emitting a specific JSON shape, a report template, a diff format, a domain-specific notation. Prompting gets you most of the way; fine-tuning gets you the last few percent and removes the schema instructions from every request. - **Tone and persona.** House voice, register, refusal style, the difference between a clinical summary and a customer-facing one. - **Task shape.** Specialised transformations — classify into this taxonomy, extract these fields, rewrite in this convention — where the instruction is long and the pattern is best shown rather than described. - **Compression.** A tuned smaller Llama frequently matches a prompted larger one on a narrow task, which changes serving cost far more than any prompt optimisation can. - **Latency and token cost.** Behaviour moved into weights is behaviour you stop paying for in every prompt. ## What it is bad at - **Facts, especially changing ones.** Training does not install a lookup table; it shifts a distribution. New facts learned from a small dataset are learned weakly, generalise unpredictably, and are impossible to update without another run. Retrieval puts the fact in the context window where the model can read it, cite it, and be corrected by editing a document. - **Provenance.** A retrieved answer can carry a source; a memorised one cannot. In any regulated or high-stakes setting this alone decides it. - **Recency.** Anything with a change cadence faster than your retraining cadence must be retrieved. - **Long-tail coverage.** Retrieval scales to millions of documents; fine-tuning on the same corpus mostly teaches style, not recall. ## The hidden costs of the fine-tuned path Training an adapter on modest hardware is cheap. What is not cheap: 1. **Data curation.** Quality dominates quantity — a few thousand carefully checked examples beat a hundred thousand scraped ones. This is human time, and it is the largest line item. 2. **Evaluation.** You need a held-out set and a general-capability check, because tuning narrowly can degrade behaviour outside the target task. Without evaluation you cannot tell an improvement from a regression, and "it looks better" does not survive review. 3. **Base-model upgrade tax.** When the next Llama generation lands, your adapter does not transfer. Retraining is another curation-and-evaluation cycle, and teams that skipped writing down their pipeline discover they cannot reproduce it. 4. **Serving complexity.** Adapters to version, route and roll back; or merged checkpoints to store per task. A prompt change is a config deploy; a fine-tune is an artifact release. ## How to sequence the decision 1. **Prompt.** Get a baseline and a measurable failure mode. Most teams stop here and are right to. 2. **Add retrieval** if the failures are knowledge-shaped — wrong facts, missing internal context, stale answers. 3. **Then consider fine-tuning** if the failures are behaviour-shaped and persist under a good prompt: schema violations, tone drift, task-convention errors, or a cost target that a smaller tuned model would hit. 4. **Combine.** The strong production pattern is a fine-tuned model that has learned how to use retrieved context — the right citation format, the right refusal when the context is insufficient, the right synthesis style — with the facts still arriving at inference time. ## Framing it as a principal Treat fine-tuning as taking on a maintained asset, not running a job. The question is not "can we improve the metric" but "are we prepared to own a dataset, an eval harness and a retraining commitment for as long as this feature lives." For a narrow, high-volume, stable task with a clear cost or quality target, yes — the payback is real and often large. For a broad assistant over documents that change weekly, almost never. The answer that impresses is the one that names the failure mode first and picks the tool second, and that is honest about which failures fine-tuning provably does not fix.

  • A team wants to fine-tune Llama on their internal wiki so it "knows the company." What do you tell them?
    That fine-tuning on prose teaches style, not reliable recall, and that anything in a wiki changes faster than they will retrain. Retrieval puts the current page in context, supports citation, and is fixed by editing a document rather than rerunning training. If they still want tuning, scope it to how answers should be formatted and grounded, not to the facts themselves.
  • How would you decide whether a fine-tune actually worked?
    A held-out set for the target task scored before and after, plus a general-capability check to catch regressions outside it, plus a human review of a sample of real production prompts. Training loss is not evidence. Without a pre-registered evaluation you cannot distinguish improvement from overfitting to the examples you happened to curate.
  • What happens to your adapter when Meta ships the next Llama generation?
    It does not transfer — adapters are tied to the exact architecture and weights they were trained against. You repeat curation, training and evaluation on the new base, which is why the durable asset is the dataset and the eval harness, not the adapter file. Teams that treat the adapter as the artifact get stranded on an ageing base model.
  • When do fine-tuning and retrieval genuinely complement each other?
    When retrieval supplies the facts and the fine-tune teaches the model how to use them — citation format, synthesis style, and above all refusing or saying "not in the provided context" when the retrieved passages do not support an answer. That grounding discipline is behaviour, which is exactly what tuning installs well, while the content stays updatable in the index.

Fine-tuning is training a new employee in your house style and procedures; retrieval is giving them access to the filing cabinet. Teaching this quarter's numbers by drilling them into memory means retraining the employee every quarter.

saying these in an interview costs you the question

  • Proposes fine-tuning to teach documents that change weekly
  • Treats fine-tuning as a substitute for retrieval rather than complementary
  • Counts only GPU hours and ignores curation and evaluation cost
  • Assumes an adapter carries over to the next base model
  • Claims a fine-tuned model can cite its sources from memorised facts

context