skip to content

How do you tell that a fine-tune has overfit its training set?

level: seniorimportance: must knowfreq 58%

answer

  1. the curves stop agreeing
  2. training down, held-out up
  3. check the wording, not just the score
  4. long shared n-grams with training answers
  5. reword the question and re-score

basics

~20 s

Two families of signal: curve divergence, where training loss keeps falling while held-out loss turns upward, and output symptoms - answers reciting training phrasings verbatim, and a sharp quality drop when the same question is reworded.

solid answer

~50 s

The quantitative signal is divergence: on a crop-disease fine-tune I have seen training loss still falling at epoch four while held-out loss bottomed out at epoch two and climbed after. That gap is the cleanest evidence the model is fitting the dataset rather than the task. But loss curves alone are not enough, so I also read the outputs. Three symptoms matter. **Verbatim recitation** - measure long n-gram overlap between generated answers and the training corpus; a jump in shared eight-grams means memorised phrasing rather than learned behaviour. **Paraphrase brittleness** - rewrite held-out prompts into different wording and re-score; a model that only performs on the training set's sentence shapes collapses here. **Diversity collapse** - the model produces one templated answer structure regardless of input, and stops saying it does not know. Any of these means the useful checkpoint was earlier than the one you kept.

code

python · 19 lines
python
def ngrams(text, n=8):
    words = text.split()
    return {tuple(words[i:i + n]) for i in range(len(words) - n + 1)}


def memorization_rate(generated, training_completions, n=8):
    corpus = set()
    for row in training_completions:
        corpus |= ngrams(row, n)
    hits = sum(1 for out in generated if ngrams(out, n) & corpus)
    return hits / len(generated)


train = ["apply a protectant fungicide before the next rain event to limit spread"]
outs = [
    "apply a protectant fungicide before the next rain event to limit spread",
    "spray a protectant product ahead of the coming rainfall to slow it down",
]
print(memorization_rate(outs, train))  # 0.5

go deeper

for a junior

Know the basic picture: training loss going down while performance on unseen data gets worse means the model is memorising, and the model to keep is an earlier checkpoint.

for a middle

Be able to describe divergence between training and held-out loss and to name output symptoms - verbatim training phrasings and worse answers when the question is reworded.

for a senior

Show you measure the symptoms rather than eyeball them: n-gram overlap against the training corpus across checkpoints, a paraphrased eval slice, output-diversity and abstention rates, all read together.

for a principal

Own the consequences: memorisation is also a data-leakage risk, and lost abstention turns uncertainty into confident advice. Set the policy for which signals block a release and who reviews them.

## What overfitting means for a fine-tuned language model Overfitting is the model reducing loss by absorbing properties of the specific training examples that do not transfer - the annotators' phrasing, the exact entities that appeared, the layout of the reference answers - rather than the underlying behaviour you wanted to teach. Fine-tuning is unusually prone to it because the datasets are small (hundreds to tens of thousands of rows against a base model with billions of parameters) and the examples are often stylistically homogeneous, having come from a handful of authors or a single synthetic generator. The distinctive thing about overfitting in a language model is that it does not look like noise. It looks like *confidence*. An overfit model produces well-formed, domain-shaped, plausible text, and only fails when the input drifts away from the training distribution - which is exactly what production traffic does. ## Signal one: curve divergence The classic evidence is the gap between the two loss curves. Training loss falls monotonically because that is what the optimiser is doing. Held-out loss, computed on examples never used for updates, falls while the model is learning transferable structure and turns upward once it starts memorising. In a concrete case - a crop-disease advisory fine-tune - held-out loss bottomed at epoch two and rose steadily while training loss kept falling through epoch four. The best checkpoint was the one at the held-out minimum, and everything after it was strictly worse for deployment despite looking better on the training curve. Two cautions. First, the divergence point is a property of *this* dataset and *this* configuration, not a universal epoch count, so it must be measured per run. Second, loss is only a proxy; on some runs the graded task metric peaks a little before or after the loss minimum, and when they disagree the task metric wins. ## Signal two: verbatim recitation The most legible output symptom is the model reproducing training-set phrasing word for word. It is directly measurable: take the generated answers for your held-out inputs, and compute the fraction that share a long n-gram (eight tokens is a common threshold) with any training completion. Track that number across checkpoints. A sharp rise means the model is retrieving stored strings rather than composing an answer, and the strings it retrieves will be wrong whenever the retrieved case does not match the new input. This also has a privacy dimension. If the training data contained customer names, addresses or identifiers, a memorising model can surface them in unrelated answers, which turns an eval finding into an incident. ## Signal three: paraphrase brittleness A model that learned the task should be robust to how the question is asked. A model that learned the dataset is not. The test is cheap: take a slice of held-out inputs, rewrite each into different wording that preserves meaning - reorder the facts, change register from clinical to conversational, translate the units, vary the length - and re-score. A model whose score holds within noise has learned behaviour. A model that drops sharply has learned surface form. This is one of the most informative checks available and it is routinely skipped, because it requires building a paraphrased variant of the eval set. A related probe is *distribution shift by construction*: hold out inputs that come from a period, region or entity absent from training, and compare the gap against the in-distribution slice. A widening gap across checkpoints tracks overfitting directly. ## Signal four: diversity and calibration collapse Late-stage overfit models become templated. Every answer follows the same three-paragraph shape, opens with the same clause, and reaches for the training set's most frequent conclusion. Two symptoms are worth measuring: **output diversity** - the number of distinct answer templates or the type-token ratio across a batch of varied inputs - and **abstention behaviour**. If the training data contained no examples of the model declining or expressing uncertainty, the fine-tune learns that a confident answer is always correct, and abstention rates fall to zero even on inputs where the right answer is "insufficient information." On a diagnostic advisory task, that is the most dangerous failure of the four, because it converts uncertainty into an authoritative recommendation. ## Reading the signals together No single signal is decisive. The pattern that confirms overfitting is *coherent movement across several of them at the same checkpoint*: held-out loss turning up, n-gram overlap climbing, the paraphrase slice diverging from the literal slice, output diversity falling. When they move together, the diagnosis is settled and the remedy is a smaller effective training budget, a more diverse dataset, or a lighter-touch adapter - and the deployable artifact is an earlier checkpoint, not the last one. What should not be accepted as evidence either way is a good score on the training set, or on any eval set whose items resemble training items closely. Overfitting is precisely the condition under which those numbers look excellent.

  • How would you measure verbatim memorisation concretely rather than eyeballing outputs?
    Index the training completions by n-gram, then for each generated answer on held-out inputs compute the longest shared n-gram and the fraction of answers sharing at least an eight-gram with any training completion. Track that fraction across checkpoints. A rising curve is memorisation; combine it with a spot check that the shared spans are substantive content rather than boilerplate like a standard disclaimer.
  • Held-out loss is flat but the model still recites training phrasings. What does that suggest?
    That the held-out set is too close to training - near-duplicates or the same entities on both sides - so it is not measuring generalisation. Rebuild the split so the held-out items come from a period, source or entity the training data never contained, and re-measure. Flat loss with visible memorisation almost always means a leaky split rather than a healthy model.
  • Does a low-rank adapter make overfitting impossible?
    No. Limiting trainable parameters reduces capacity to memorise, but it does not eliminate it - a small adapter trained for enough epochs on a few hundred homogeneous examples will still absorb their phrasing. The signals and the checkpoint discipline apply regardless of whether you trained all parameters or an adapter.

saying these in an interview costs you the question

  • Judging overfitting from the training loss curve alone
  • Assuming a small adapter cannot memorise its training data
  • Treating high scores on a leaky held-out set as generalisation
  • Never re-testing with reworded versions of the same questions
  • Shipping the final checkpoint because training loss was lowest there

context