skip to content

What did the Chinchilla result change about splitting a fixed LLM training budget?

level: middleimportance: must knowfreq 68%

answer

  1. one budget, two knobs
  2. isoFLOP curves, not intuition
  3. parameters and tokens grow together
  4. roughly twenty tokens per parameter
  5. the giant models were undertrained

basics

~20 s

Chinchilla showed that for a fixed training-compute budget, parameters and training tokens should grow together — roughly 20 tokens per parameter — rather than spending almost the whole budget on more parameters, as earlier Kaplan-era guidance implied.

solid answer

~50 s

Both results fit power laws for loss against model size, data and compute, but they disagreed on how to spend a fixed budget. The earlier Kaplan-style guidance said the dominant lever is parameter count: make the model much bigger and train it on comparatively few tokens. The Chinchilla work re-ran the experiment with isoFLOP sweeps — for each compute budget, train many (size, tokens) combinations and read off the minimum — and found parameters and tokens should scale in roughly equal proportion, about 20 tokens per parameter. The discrepancy is largely attributed to learning-rate schedules that were not matched to each run's token count. The practical consequence was that the giant models of that era were *undertrained*: a 70B model trained on 1.4T tokens beat several models three to seven times its size. Today the 20:1 figure is a reference point, not a target — frontier models are deliberately trained far past it.

code

python · 8 lines
python
def compute_optimal(flops_budget, tokens_per_param=20.0):
    # C ~= 6 * N * D  and  D = ratio * N  =>  N = sqrt(C / (6 * ratio))
    n = (flops_budget / (6.0 * tokens_per_param)) ** 0.5
    return n, tokens_per_param * n

for c in (1e21, 2e21, 1e23):
    n, d = compute_optimal(c)
    print(f"{c:.0e} FLOPs -> {n/1e9:.1f}B params, {d/1e12:.2f}T tokens")

go deeper

for a junior

Know that model size and training data are two separate dials, that they trade off against one fixed compute budget, and that Chinchilla's finding was to grow both rather than only the model.

for a middle

Be able to explain the isoFLOP experiment and the C is roughly 6 times N times D relation, state the ~20 tokens-per-parameter result, and say why the huge models of that era were called undertrained.

for a senior

Expect to be pushed on the limits: the law predicts pretraining loss, not downstream capability, and it optimises training cost while ignoring serving cost. Explain why production ratios sit far past 20:1.

for a principal

Own the framing that a scaling law is an empirical fit carrying its own recipe assumptions, and that a revised result changed a whole industry's model sizes. Be ready to argue what would make you re-fit rather than trust a published curve.

## The problem being solved Pretraining a large language model consumes a fixed, expensive budget of floating-point operations (FLOPs). That budget can be spent in two very different ways: on a bigger model (more parameters, N) or on more training data (more tokens, D). A useful approximation ties the three together: **C ≈ 6 × N × D** where C is training FLOPs. The factor of ~6 comes from roughly 2 FLOPs per parameter for the forward pass and ~4 for the backward pass, per token. The consequence is that N and D trade off directly: at a fixed C, doubling the parameter count halves the number of tokens you can afford. "Scaling laws" are the empirical answer to how that trade should be made. ## Kaplan-era guidance The 2020 scaling-law work fitted smooth power laws relating validation loss to N, D and C, and reported that loss improves predictably as each grows. Its allocation advice was strongly parameter-biased: given more compute, grow the model aggressively, feed it comparatively few tokens, and stop training early rather than run to convergence. The models built on that advice were very large and, by later standards, thin on data. ## The isoFLOP correction The 2022 Chinchilla work asked the allocation question directly rather than inferring it. For each of many fixed compute budgets, it trained a family of models spanning small-and-data-rich to large-and-data-poor, and plotted final loss against parameter count. Each such curve — an **isoFLOP curve** — is U-shaped: too small a model underfits the budget, too large a model never sees enough tokens, and the minimum sits in between. Connecting the minima across budgets gives the compute-optimal frontier. The answer was that N and D should each scale as roughly the square root of compute. Double the FLOP budget and you want about 1.4× the parameters *and* 1.4× the tokens — not 2× the parameters. At the scales studied, the optimum sat near **20 training tokens per parameter**. The headline demonstration: a 70B-parameter model trained on 1.4T tokens outperformed contemporaries several times larger that had been trained on a few hundred billion tokens. The extra parameters had not been wasted so much as starved. ## Why the two results disagreed The discrepancy is now mostly attributed to methodology rather than to a change in the underlying physics. The earlier study used a learning-rate schedule that was not re-tuned to each run's actual token count, which systematically penalises the longer, data-rich runs, and counted parameters in a way that shifted the small-model end of the fit. Re-analyses that correct for these bring the two families of results substantially closer. This is a useful lesson in itself: a scaling law is an empirical fit, and it inherits every training-recipe assumption baked into the runs it was fitted on. ## What the law does and does not tell you It predicts **pretraining loss** — average next-token cross-entropy on held-out text. It does not directly predict downstream task accuracy, instruction-following quality, or reasoning ability, all of which are mediated by post-training. It also says nothing about *serving* cost: it minimises training FLOPs for a target loss, treating inference as free. That single omission is why the compute-optimal point is usually the wrong place to build a production model. The law is also fitted to a particular data distribution and architecture family. Change the data quality, the tokenizer, or the architecture and the coefficients move; the *shape* (smooth power law, U-shaped isoFLOP curves) has proved robust, the constants have not. ## Status of 20:1 today As of mid-2026, ~20 tokens per parameter is a historical reference point, cited to explain the correction, not a target anyone hits. Production models are routinely trained at hundreds to thousands of tokens per parameter — an open-weight 8B model trained on 15T tokens sits near 1,900:1. Two forces pushed practice past the frontier: inference economics, which reward small models even at extra training cost, and the fact that tokens rather than FLOPs are increasingly the binding constraint. Meanwhile the single-curve story has itself been superseded: pretraining is now discussed alongside reinforcement-learning post-training compute and test-time compute as three separate scaling regimes with different returns. ## What an interviewer is checking That you know the allocation question exists, that the answer is empirical and was revised, that ~20:1 was the revised answer at the time, and — most importantly — that you can say why nobody builds at that ratio any more.

  • If you doubled the training FLOP budget, how would compute-optimal scaling have you spend it?
    Split it across both axes rather than one. Since C is roughly proportional to N times D and the optimum keeps them in fixed ratio, each grows as the square root of compute: about 1.4x the parameters and about 1.4x the tokens. Putting the whole doubling into parameters lands you off the frontier on the undertrained side of the isoFLOP curve.
  • Chinchilla predicts pretraining loss. Why is that a weaker guarantee than it sounds for a product team?
    Loss is average next-token cross-entropy on held-out text. It correlates with capability but does not directly predict the things a product cares about — task accuracy, instruction-following, tool use, refusal behaviour — all of which are shaped by post-training. Two models at identical pretraining loss can differ enormously after alignment. Treat the law as a budgeting tool for the pretraining stage, not as a capability forecast.
  • Why is nobody building at 20 tokens per parameter today?
    Two reasons. The law minimises training FLOPs and treats inference as free, which is wrong for anything serving real traffic — a smaller, over-trained model is cheaper on every request forever. And high-quality tokens, not FLOPs, are increasingly the binding constraint, so the interesting question shifted from how to split a compute budget to how to get more usable data at all.

Baking with a fixed amount of fuel: the earlier advice was to buy an ever-larger oven, and the correction was that a mid-size oven left running long enough beats a huge one switched off early.

saying these in an interview costs you the question

  • Says bigger models are always better regardless of training data
  • Thinks Chinchilla recommends training on fewer tokens
  • Quotes 20 tokens per parameter as a rule current models follow
  • Assumes compute-optimal also means cheapest to serve
  • Treats scaling laws as predicting specific downstream skills

context