skip to content

Pretraining Objectives and Data

One enormous next-token run over filtered web text, then a mid-training stage for long context and domain mixes. This is what explains knowledge cutoffs, contamination and why base models do not chat.

on this pageshow

questions

6

Why does a base LLM checkpoint continue your prompt instead of answering it?

level: juniorimportance: must knowfreq 62%

answer

  1. it completes documents, not conversations
  2. FAQ pages are full of more questions
  3. no marker says the prompt ended
  4. role turns and an end-of-turn token
  5. instruction data comes after pretraining

basics

~20 s

A base checkpoint is trained only to continue documents, so a question is most plausibly followed by more document text rather than an answer. Answering, and stopping when finished, come from the post-training stages that produce an instruct checkpoint.

solid answer

~50 s

Pretraining optimises one thing: the probability of the next token in real text. In real text, a line like `Q: What is a patent claim? A:` most often sits inside an FAQ page, so the highest-probability continuation is a short answer followed immediately by the *next* question — and a base model will happily produce five more of them, then a footer, then a cookie notice. It never learned that the text before it is an instruction addressed to it, and it has no notion of a reply being complete. An instruct checkpoint is that same base model after post-training on instruction-response data, which teaches the mapping from request to answer, plus a chat template with explicit role turns and an end-of-turn token that gives it somewhere to stop. Base checkpoints still shine at few-shot continuation and are the starting point for further training; instruct checkpoints are what an API usually serves.

go deeper

for a junior

Be able to say that a base model only continues text, that an instruct model has been trained afterwards on request-and-response data, and that this is why the API you call is normally the instruct one.

for a middle

Explain the mechanism rather than the label: the likely continuation of an FAQ line is another FAQ line, and chat behaviour comes from a role-based template plus a learned end-of-turn token.

for a senior

Show judgment about which checkpoint to start from — base for further training and completion-shaped work, instruct for anything user-facing — and know that template mismatch quietly degrades an instruct model.

for a principal

Own the consequence for build strategy: whether the organisation trains from base and inherits full control of behaviour, or builds on a vendor instruct checkpoint and inherits its persona, refusal profile and template as constraints.

## The behaviour, concretely Send the string `Q: What is a patent claim? A:` to a base checkpoint and you are likely to get something like: a plausible one-line answer, then `Q: How long does a patent last? A:`, then another, then a navigation menu. Send it to an instruct checkpoint of the same model and you get one answer that ends. Nothing is broken in the first case. The base model did exactly what it was trained to do. ## Why continuation is the correct behaviour for a base model Pretraining maximises the likelihood of the token that actually follows in the corpus. The model has seen millions of documents where a question-and-answer pattern appears, and in almost all of them the pattern repeats — that is what an FAQ page is. So the statistically correct continuation of one question-answer pair is another question-answer pair. There is no marker in the training data that says "the text ends here and now a different agent should respond", because in a document there is no different agent. Everything is one stream. Three consequences follow directly: - **No addressee.** The model has no representation of "you asked me something". Your prompt is a prefix, not a request. - **No stopping rule.** Documents do not end after one paragraph, so the model keeps going until it hits a length limit. Practitioners using base models supply stop sequences manually. - **No persona or refusal behaviour.** Helpfulness, tone and safety responses are not properties of text prediction; they are added later. ## What turns a base checkpoint into an instruct checkpoint Post-training. The model is trained further on data where a request is followed by the desired response, which teaches the request-to-response mapping, and further stages shape which of many possible responses it prefers. Two mechanical pieces matter as much as the training itself: **A chat template.** Conversations are serialised into a fixed format with explicit role markers — a system portion, user turns, assistant turns. The model learns that its job is to produce the assistant span. This is a *formatting convention baked into the weights*, which is why feeding a chat-formatted prompt to a base model does nothing useful, and why using the wrong template with an instruct model degrades it noticeably. **An end-of-turn token.** Post-training teaches the model to emit a special token when the reply is done, which the serving layer turns into a stop. That is the entire mechanism behind "it knows when to stop". ## When you would still want the base checkpoint - **Further training.** Any fine-tuning or continued pretraining of your own generally starts from base, because you are not fighting an existing conversational style. - **Pure completion tasks.** Code infilling, text continuation and template completion match what a base model natively does. - **Few-shot pattern induction.** Base models follow a demonstrated pattern in the prompt very literally, which is sometimes cleaner than an instruct model that tries to be conversational about it. - **Studying pretraining effects.** Measuring what pretraining alone produced requires a model that post-training has not reshaped. Conversely, anything user-facing wants the instruct checkpoint: it follows instructions, respects a system prompt, terminates, and has whatever safety behaviour the provider trained in. ## Common confusions worth pre-empting *"Base models are just smaller or older."* No — base and instruct are usually the same weights at different pipeline stages, from the same run. *"The base model is dumber."* It has the same knowledge. Post-training reshapes how that knowledge is expressed and adds a great deal of usability; measured raw capability is broadly comparable, while conversational usefulness is not. *"Prompting harder fixes it."* Partly, and this is worth knowing: you can coax reasonable behaviour out of a base model by making the *document* look like one where a good answer follows, for instance by writing a few complete examples first. You are not instructing it; you are making the answer the likely continuation. *"It refused me, so it must be instruct."* Refusals, hedging and self-description as an assistant are all post-training artefacts. Their presence is a reliable tell that you are not on a base checkpoint. ## What interviewers are checking That you can reason from objective to behaviour without hand-waving: pretraining predicts document continuations, therefore a base model continues documents, therefore chat behaviour and stopping must be added afterwards. Naming the chat template and the end-of-turn token turns a decent answer into a confident one.

  • What mechanically makes an instruct model stop, when a base model does not?
    Post-training teaches it to emit a special end-of-turn token at the end of its reply, and the serving layer stops generation when that token appears. A base model never learned to produce such a marker, because documents do not contain one, so it keeps predicting until a length limit or a manually supplied stop sequence cuts it off.
  • When would you deliberately start from a base checkpoint rather than an instruct one?
    When you are training further yourself, since a base model carries no conversational style you would have to override, and when the task is genuinely completion-shaped — code infilling, template filling, continuing a document. Base models also follow a demonstrated few-shot pattern very literally, which is sometimes cleaner than an instruct model that insists on conversational framing.
  • Does a base model know less than the instruct version of the same model?
    No. They usually come from the same pretraining run and carry the same knowledge; the instruct checkpoint is that base model trained further. What changes is expression and usability — instruction-following, a system-prompt-respecting persona, termination and safety behaviour — not the facts stored in the weights.

A base model is autocomplete for the whole web: type a question and it writes what usually comes after such a question on a page, which is often more questions. An instruct model has been taught that the text is addressed to it and that a reply has an end.

saying these in an interview costs you the question

  • Calling the base model broken or undertrained
  • Believing base and instruct are different sizes or generations
  • Assuming a base model understands a system prompt
  • Thinking pretraining alone teaches question answering
  • Expecting a base model to stop without a stop sequence

context

open as a page

What is an LLM's knowledge cutoff, and why is it fuzzy rather than a hard date?

level: middleimportance: must knowfreq 52%

basics

~20 s

A knowledge cutoff is the point after which no training text was collected. It is fuzzy because crawls trail real events, coverage of recent months is thin, later training stages can add newer data, and the model has no reliable sense of its own horizon.

open as a page

What does next-token cross-entropy actually optimise during LLM pretraining?

level: middleimportance: must knowfreq 72%

basics

~20 s

Next-token cross-entropy maximises the probability the model assigns to each real next token in the corpus. It optimises fit to the data distribution, including that data's errors and style, never truthfulness, helpfulness or task success.

open as a page

When building a pretraining corpus, why is deduplication worth more than extra raw volume?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Repeated text teaches memorisation and burns compute on tokens the model has already fitted, so removing near-duplicates often improves a model more than adding volume. Quality filtering helps too, but every filter narrows the distribution the model can represent.

open as a page

In LLM training, what is the mid-training stage and what goes into it?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Mid-training is a named stage between bulk pretraining and instruction tuning where the data mixture changes: longer documents to extend usable context, plus heavier code, maths, curated and synthetic text, usually with longer sequences and a decaying learning rate.

open as a page

When does synthetic or teacher-distilled text belong in a pretraining mix?

level: principalimportance: should knowfreq 35%

basics

~20 s

Generated text buys coverage where real text is scarce or badly written, and it is now standard practice. It also inherits the generator's blind spots, so ground it in real source documents, cap its share, and measure diversity rather than loss alone.

open as a page