skip to content

In LLM training, what is the mid-training stage and what goes into it?

level: seniorimportance: should knowfreq 32%

answer

  1. still next-token prediction, changed mixture
  2. raise sequence length late
  3. code, maths, long documents, transcripts
  4. learning rate decaying to zero
  5. last tokens leave the deepest mark

basics

~20 s

Mid-training is a named stage between bulk pretraining and instruction tuning where the data mixture changes: longer documents to extend usable context, plus heavier code, maths, curated and synthetic text, usually with longer sequences and a decaying learning rate.

solid answer

~50 s

The modern pipeline is described in four stages — data, pretraining, mid-training, post-training — and mid-training is the deliberate re-mixing at the end of the pretraining run. Two things typically change. The sequence length is raised and long documents are up-weighted so the model actually learns to use a long window, which is far cheaper than doing the whole run at that length. And the mixture tilts toward what you want the final weights to be good at: code, mathematics, scientific text, tool-call transcripts, curated and synthetic material. It usually runs while the learning rate decays toward zero, which is why it matters disproportionately — the model finishes on this distribution, so late-stage data leaves a heavier imprint per token than early-stage data. That leverage cuts both ways: contamination or an over-weighted domain here does more damage than the same mistake early in the run.

go deeper

for a junior

Know the pipeline has more than two steps: a big general pretraining run, a mid-training phase that changes the data mixture, and then the post-training that makes a model chat.

for a middle

Explain what changes and what does not: same next-token objective, different mixture, longer sequences, decaying learning rate, and still a base checkpoint at the end.

for a senior

Argue the why: long-context and domain capability bought cheaply in the tail, and outsized imprint per token while annealing. Then name the risks you would manage — contamination in the late mixture and regression from over-specialisation.

for a principal

Own the allocation. Decide what capability is worth buying with the final, highest-leverage tokens, how mid-training variants are branched from a shared bulk checkpoint, and what evidence gates a mixture change into a run that cannot be repeated cheaply.

## Where the stage sits Until recently the pipeline was described in two moves: pretrain on a big general corpus, then post-train it into an assistant. Practice has settled on an explicit third stage between them. The bulk of pretraining runs on a broad, mostly-web mixture at a moderate sequence length. Then, before any instruction tuning or preference work, the run continues with a changed mixture and changed settings. That continuation is mid-training. It is still next-token prediction on documents — the objective does not change. What changes is *which* documents and at *what* sequence length, and that the learning rate is on its way down. ## What typically goes in - **Long documents, at a longer sequence length.** Full scientific papers, books, large codebases, extended transcripts. Training the whole run at maximum length would be prohibitively expensive, so the length is raised late and long-form data is up-weighted then, which is how models learn to use context rather than merely accept it. - **High-value domains.** Code, mathematics with worked derivations, scientific and technical writing — material that is scarce as a fraction of the crawl but heavily represented in what users actually ask for. - **Tool-call and structured transcripts.** Sequences showing structured calls and their results, so that the base model is already familiar with the shape before any agent-style post-training. - **Curated and synthetic text.** Rephrased or generated material at higher density than raw crawl. - **Recency refresh.** Newer crawl data late in the run, which is one reason a model's practical knowledge horizon does not line up neatly with the date of the main crawl. ## Why a separate stage rather than one uniform mixture **Cost.** Long-context training is expensive per token. Confining it to the tail of the run buys most of the capability for a small fraction of the cost. **Recency of imprint.** As the learning rate decays, updates get smaller but the model's final position in parameter space is dominated by where it has just been. Data seen while annealing has an outsized effect per token compared with the same data seen early. That makes the tail the cheapest, highest-leverage place to buy a capability. **Separation of concerns.** The expensive bulk run can use cheap, plentiful, generic tokens; the scarce, expensive, curated tokens are spent where they count. Teams can also re-run mid-training variants from a shared bulk checkpoint to produce differently specialised base models without repeating the full run. **It precedes instruction tuning.** This is the boundary candidates most often get wrong. Mid-training is still document prediction. It does not produce an assistant, it does not use instruction-response pairs or preference data, and its output is still a base checkpoint. ## The risks that come with the leverage - **Contamination is worse here.** A benchmark leaking into the annealing mixture inflates scores far more than the same leak early in the run, because the model finishes on it. Decontamination on the mid-training mixture deserves more scrutiny, not less. - **Over-specialisation.** Push the mixture too far toward one domain and general capability regresses; the model is spending its final, most influential updates narrowing itself. - **Long-context is not free just because the window grew.** Raising the sequence length and adding long documents improves usable context, but models still degrade well before their advertised limit, and that has to be measured on multi-fact, long-range evaluations rather than assumed from the configured length. - **Mixture opacity.** Because this stage is where much of the differentiated capability is bought, it is also the least documented part of most model releases, which makes reproducing or diagnosing behaviour harder. ## What interviewers are checking Mostly that you have a current mental model of the pipeline rather than the two-stage version, and that you can explain *why* the split exists in cost and leverage terms instead of just naming it. A strong answer places mid-training precisely — after the bulk of next-token pretraining, before any instruction tuning, still producing a base checkpoint — names what goes into it, and then volunteers the risk: the last tokens the model sees are the ones that matter most, so this is exactly where contamination and over-specialisation do their worst damage.

  • Why extend the sequence length late rather than train the whole run long?
    Cost. Attention and memory pressure make long-sequence steps far more expensive per token, so a full run at maximum length is prohibitive. Most of what long-context ability requires is learning to attend over distant material, which can be bought in a comparatively short tail phase with long documents up-weighted. The bulk of the run then stays cheap at a moderate length.
  • Why is contamination in the mid-training mixture more damaging than early contamination?
    Because the model finishes on that distribution while the learning rate decays, so late data leaves a heavier imprint per token. A benchmark item seen early may be substantially washed out by trillions of subsequent tokens; the same item seen during annealing sits close to the final weights and inflates the score directly. Decontamination should therefore be strictest on the late mixture.
  • Does mid-training produce something you can chat with?
    No. It is still next-token prediction over documents, so the output is a base checkpoint that continues text rather than answers. It may make the model much better at code, maths or long inputs, and familiar with structured transcripts, but instruction-following and turn-taking behaviour come from the post-training stages that follow.

saying these in an interview costs you the question

  • Describing mid-training as instruction tuning by another name
  • Thinking the training objective changes at this stage
  • Placing mid-training after preference optimisation
  • Assuming a longer configured window means usable long context
  • Treating late-stage mixture choices as low-stakes

context