Why deliberately train a 3B LLM far past its compute-optimal token count?
answer
- training paid once, inference paid forever
- which term scales with request count
- parameters set serving cost, tokens do not
- the flat tail of the loss curve is purchasable
- deployed models sit deliberately off the frontier
basics
~20 sBecause inference, not training, is the bill. Over-training a small model costs more once, but every request afterwards runs on fewer parameters — cheaper, faster and denser to batch across hundreds of millions of monthly calls.
solid answer
~50 sCompute-optimal scaling minimises the FLOPs needed to reach a target loss, and it implicitly treats serving as free. For anything with real traffic that assumption is wrong. If a model handles 400M requests a month, the lifetime inference bill dwarfs the one-time training bill, and inference cost is driven mainly by parameter count — it sets weight memory, arithmetic per token, KV-cache pressure and therefore how many concurrent requests fit on a GPU. So you pick the parameter count you can afford to *serve*, then spend the extra training compute pushing that fixed-size model further along its data curve. The loss-versus-tokens curve keeps improving well past the compute-optimal point, just with diminishing returns, so you are buying quality at a price your users never pay again. The result is intentionally off the training-optimal frontier: too small for its budget, exactly right for its deployment.
code
python · 9 linesdef lifetime_flops(n_params, train_tokens, monthly_requests, out_tokens, months):
train = 6.0 * n_params * train_tokens
serve = 2.0 * n_params * out_tokens * monthly_requests * months
return train, serve
# compute-optimal 30B vs deliberately over-trained 3B, 400M req/month for 24 months
for name, n, d in (("30B @ 600B tok", 30e9, 600e9), ("3B @ 6T tok", 3e9, 6e12)):
t, s = lifetime_flops(n, d, 400e6, 500, 24)
print(f"{name}: train {t:.2e}, serve {s:.2e}, total {t + s:.2e}")go deeper
Know that a smaller model is cheaper and faster to run, and that training a small model on lots of data is a deliberate way to get a cheap-to-serve model rather than an accident.
Explain that training cost scales with parameters times tokens and is paid once, while inference cost scales with parameters alone and is paid per request — so training tokens are free at serving time.
Show you can size from a traffic forecast: estimate lifetime serving FLOPs against one-time training FLOPs, and name the serving channels parameter count drives — weight memory, per-token arithmetic, KV-cache room and batch density.
Own the capital-versus-operating framing across a portfolio: which workloads justify a bespoke over-trained model, which should rent a larger one, and how a traffic forecast that turns out wrong by an order of magnitude changes the answer.
## Two different optimisation problems The compute-optimal frontier answers: *given C training FLOPs, what model reaches the lowest loss?* A deployed system asks something else: *given a quality bar and a traffic forecast, what total cost of ownership do I pay?* Those are different objectives and they have different optima. Write the lifetime cost roughly as: **total ≈ training_flops + requests × tokens_per_request × inference_flops_per_token** Training FLOPs are approximately 6 × N × D, paid once. Inference is approximately 2 × N FLOPs per generated token (forward pass only), paid on every token of every request, forever. The moment `requests × tokens_per_request` grows large, the second term dominates — and note that it depends on N but *not* on D. Training tokens are free at serving time. Parameters are not. ## The concrete shape of the decision Consider a feature that will serve on the order of 400M requests a month. Two candidates reach a similar quality bar: a larger model sitting on the compute-optimal frontier, and a 3B model trained on far more tokens than its budget "deserves". The compute-optimal one is cheaper to train and more expensive on every single one of those 400M requests. Over a couple of years of traffic the extra training spend on the 3B model is recovered many times over. Parameter count drives serving cost through several channels at once, which is why it dominates: - **Weight memory.** Fewer parameters means the weights fit in less accelerator memory, leaving room for larger batches and longer KV caches — or fitting on cheaper hardware entirely. - **Arithmetic per token.** Decode cost scales with active parameters, so latency per token falls. - **Batch density.** More concurrent requests per accelerator directly divides the fixed cost of the machine across more users. - **Deployment surface.** Small enough, and the model runs on-device or at the edge, where the marginal serving cost approaches zero. ## Why over-training works at all The loss-versus-tokens curve at fixed parameter count does not have a cliff. Past the compute-optimal ratio it keeps descending, just more slowly — you are on the flat part of a power law. That means over-training is a smooth purchase: you can decide how much extra training compute you are willing to convert into serving savings, rather than facing a hard wall. Practice has run a very long way down that curve; open-weight models in the 7-8B range trained on 15T tokens sit near 1,900 tokens per parameter, roughly a hundred times the classic reference ratio. The limits are real, though. Returns per extra token shrink, so at some point the marginal training FLOP would have been better spent on a slightly larger model, on post-training, or on data quality. And you cannot over-train past the data you have — which is exactly where the data-availability constraint starts to bind. ## Related levers that serve the same goal Over-training is one way to get a small, strong model, and it is usually combined with others: distilling from a larger teacher, quantising weights to 4-bit formats for serving, and architectures where only a fraction of parameters are active per token so that serving cost tracks *active* rather than total parameters. When someone says "inference-optimal sizing", they usually mean the whole bundle, with over-training as the training-side component. ## How to reason about it in an interview The strong answer is not "small models are cheaper". It is the shape of the trade: training cost is capital expenditure paid once and scales with N × D; inference cost is operating expenditure paid per request and scales with N alone. Therefore the right question is never "what is the compute-optimal model" but "at my traffic volume, where does total cost bottom out" — and at high volume that answer is always smaller and longer-trained than the frontier says. The cases where you should *not* over-train are equally worth naming: low-traffic internal tools, research models whose only job is to establish a capability, offline batch jobs where latency is irrelevant and you can amortise a big model over a scheduled window, and anything where the quality ceiling of the small model simply does not clear the product bar. Sizing is a function of the traffic curve, not a universal preference.
- At what point does over-training stop paying?When the marginal training FLOP buys less quality than spending it elsewhere. Returns along the token axis decay as a power law, so each additional trillion tokens moves loss less than the last; past some point a slightly larger model, better post-training, or higher-quality data is the better purchase. You also stop when you run out of tokens worth training on. The stopping point is set by your traffic volume and quality bar, not by a fixed ratio.
- Would you make the same call for a low-traffic internal tool?No. The whole argument rests on the serving term dominating, which requires volume. For an internal tool handling a few thousand calls a day, the lifetime inference cost is trivial and the sensible move is to take the best model you can get for the least engineering effort — often an off-the-shelf larger one. Over-training is a decision you earn by having traffic.
- Does this argument change for architectures where only some parameters are active per token?The structure holds but the relevant quantity shifts. Serving arithmetic tracks the parameters actually activated per token, while weight memory still tracks the total. So such a model can be cheap in FLOPs and expensive in memory footprint at the same time, and the sizing decision becomes two-dimensional: total parameters for memory and residency, active parameters for compute and latency.
Compute-optimal sizing is like buying the cheapest car on the lot; inference-optimal sizing is buying the one with the best fuel economy because you are about to drive it a million miles.
saying these in an interview costs you the question
- Assumes compute-optimal sizing is also cost-optimal to deploy
- Thinks extra training tokens make inference slower
- Ignores request volume when choosing a model size
- Believes loss stops improving past the compute-optimal token count
- Treats over-training as always right, regardless of traffic