skip to content

What does next-token cross-entropy actually optimise during LLM pretraining?

level: middleimportance: must knowfreq 72%

answer

  1. one loss, over every token position
  2. negative log of the observed token
  3. exponentiate the mean for perplexity
  4. rewards likelihood, not correctness
  5. corpus properties become model properties

basics

~20 s

Next-token cross-entropy maximises the probability the model assigns to each real next token in the corpus. It optimises fit to the data distribution, including that data's errors and style, never truthfulness, helpfulness or task success.

solid answer

~50 s

At every position in the training text the model outputs a probability distribution over the vocabulary for the next token, and the loss is the negative log of the probability it gave to the token that actually followed. Averaged over trillions of positions, gradient descent simply makes the observed corpus more likely; exponentiating that average loss gives perplexity. The bet is that predicting text well *requires* learning grammar, then meaning, then world facts, then latent skills like arithmetic and translation — which is why one objective yields such general capability. But it optimises likelihood, not correctness: a falsehood repeated across the corpus becomes a confident prediction, every token is weighted equally whether it is a load-bearing digit or filler, and nothing in the loss says "answer the user" or "stop when finished". Those come from later stages. In practice the corpus *is* the specification.

code

python · 11 lines
python
import math

# Probability the model assigned to the token that actually came next.
probs_of_true_next_token = [0.5, 0.25, 0.1, 0.9]

losses = [-math.log(p) for p in probs_of_true_next_token]
mean_loss = sum(losses) / len(losses)

print(f"per-token losses: {[round(x, 3) for x in losses]}")
print(f"mean cross-entropy: {mean_loss:.3f} nats")
print(f"perplexity: {math.exp(mean_loss):.2f}")

go deeper

for a junior

Be able to say plainly that pretraining predicts the next token over huge amounts of text, with no human labels, and that being good at prediction is not the same as being right.

for a middle

Explain the mechanics: a distribution over the vocabulary at each position, negative log probability of the true token as the loss, perplexity as its exponential, and teacher forcing making all positions trainable in parallel.

for a senior

Show that you reason from objective to failure mode — confidently repeated corpus errors, no stopping behaviour, equal weight on trivial and load-bearing tokens — and that you treat data curation as the real lever because the corpus is the specification.

for a principal

Own the framing that choosing the objective and mixture is choosing what the organisation is optimising. Be ready to argue where extra signal should enter — data curation, later training stages, or verifiers — rather than complicating the pretraining loss.

## The objective, stated precisely Pretraining is self-supervised: there are no labels, only text. The corpus is tokenized into a long stream of integer token IDs. At each position the model consumes the tokens before it and emits a probability distribution over the entire vocabulary for the token that comes next. The loss at that position is the negative logarithm of the probability the model assigned to the token that *actually* came next — cross-entropy between the model's predicted distribution and the one-hot truth. Training averages that over every position in every batch and takes a gradient step, so each update makes the real continuations of the real corpus a little more likely. Two derived numbers matter in conversation. **Perplexity** is the exponential of the mean cross-entropy: a perplexity of 10 means the model is, on average, as uncertain as if it were choosing uniformly among 10 tokens. **Bits per byte/character** normalises the same quantity by text length so corpora with different tokenizers can be compared. ## Teacher forcing, and one consequence Because the whole sequence is known in advance, all positions are trained in parallel and the model always conditions on the *true* prefix, never on its own sampled output. This is called teacher forcing, and it is what makes pretraining efficient. It also means the model never practises recovering from its own mistakes during pretraining — a mismatch between training and generation (often called exposure bias) that contributes to drift in long free-running generations. ## Why one dumb objective produces general ability Prediction is compression. To assign high probability to the next token in a physics paper you need syntax, then topic, then the actual physics; to predict the next line of a diff you need to model the code. Because there is no task label, any text can be training data, and any regularity in that text is something the loss rewards you for internalising. This is why capability appears to arrive "for free" as scale grows: skills such as translation, arithmetic and code completion are latent in the prediction task rather than taught explicitly. Note the contrast with the other classic self-supervised objective, masked language modelling (BERT-style), where random tokens are blanked and predicted from both sides. Decoder-only LLMs use the **causal** next-token form, which is what makes them natural generators. ## What the objective does not optimise - **Truth.** The loss rewards likelihood under the corpus. A misconception written on ten thousand pages is, to this objective, simply a high-probability continuation. Lower validation loss does not mean fewer factual errors. - **Helpfulness or instruction-following.** Continuing a document is not answering a question. Nothing in the loss distinguishes "reply to the user" from "write the next paragraph of the web page". - **Knowing when to stop.** There is no reward for producing a complete, terminated answer; end-of-turn behaviour is imposed later by the chat format and post-training. - **Error importance.** Every token contributes equally. Getting a decimal digit wrong in a dosage costs exactly what getting a filler word wrong costs. The loss has no notion of which mistakes are consequential. - **Safety, calibration or refusal.** None of these are expressible as "predict the observed token". ## The corpus is the specification Since the only thing being optimised is agreement with the training text, every property of that text becomes a property of the model: the language mix, the share of code, the writing registers present, how often documents repeat, and anything accidentally included such as a public benchmark. This is precisely why the expensive engineering in pretraining is *data* engineering — deduplication, quality filtering, mixture weights — rather than clever loss functions. A change to the mixture is a change to the model's specification. ## What interviewers are checking They want to hear that you can separate the mathematical objective (maximise corpus likelihood) from the product behaviour (a helpful assistant), and that you know the second is not implied by the first. Strong answers name cross-entropy and perplexity precisely, explain teacher forcing in a sentence, give the compression intuition for why general skills emerge, and then draw the hard line: likelihood is not truth, and everything conversational about a deployed model is added after this stage.

  • If validation loss drops, can you conclude the model became more factual?
    No. Lower loss means the model assigns higher probability to the held-out text, which measures fit to the data distribution. If that data contains errors, outdated claims or a skewed register, better fit reproduces them more confidently. Factuality has to be measured with task evaluations against ground truth, not inferred from perplexity.
  • Why does the loss weight a critical digit the same as a filler word?
    Cross-entropy is defined per token position with no notion of semantics or consequence, so every position contributes its own negative log probability equally. Importance-weighting would require knowing which tokens matter, which is exactly the labelling that self-supervision avoids. Downstream stages and verifiers, not the pretraining loss, are where consequence-aware signals enter.
  • How does causal next-token prediction differ from masked language modelling?
    Masked language modelling blanks random tokens and predicts them using both left and right context, which suits encoding and classification. Causal next-token prediction conditions only on the preceding tokens, so the trained model can generate text autoregressively one token at a time. Modern decoder-only LLMs use the causal form for exactly that reason.

saying these in an interview costs you the question

  • Saying the model is trained to answer questions correctly
  • Claiming lower pretraining loss means fewer hallucinations
  • Describing pretraining as supervised learning with labels
  • Thinking the loss is computed only on the final answer
  • Confusing masked-token prediction with causal next-token prediction

context