What is the data wall in LLM pretraining, and what options remain when it binds?
answer
- compute grows faster than human writing
- which input becomes the binding constraint
- a few epochs, then diminishing returns
- generated tokens need a check or a stronger source
- recursive self-training degrades the distribution
basics
~20 sThe data wall is the point where high-quality text runs out before compute does, so extra FLOPs no longer buy extra fresh tokens. The remaining levers are repeating data across epochs, generating synthetic data, and pulling in other modalities.
solid answer
~50 sScaling laws assume you can always buy more tokens to match more compute. That assumption expires: the stock of high-quality human-written text on the open web is large but finite, and frontier runs already consume a meaningful fraction of it. Once tokens rather than FLOPs are the binding constraint, the scaling question changes shape — adding accelerators no longer moves loss, because there is nothing new to feed them. Three levers remain. You can **repeat** data: a handful of epochs over the same corpus is close to as valuable as fresh tokens, with returns decaying quickly beyond that. You can **generate** data, which is now the dominant industrial answer — rephrasing existing text, diversifying it across many personas, and distilling outputs from stronger models. And you can **broaden** the input: code, other languages, licensed proprietary corpora, and non-text modalities. All three raise the ceiling; none of them makes it disappear.
go deeper
Know that pretraining needs enormous amounts of text, that the supply of good human-written text is finite, and that synthetic and repeated data are the usual responses.
Explain the inversion — compute keeps growing while high-quality text does not — and describe what each lever buys: a few epochs of repetition, generation anchored to a stronger source, and expansion into code and other modalities.
Show judgement about generated data: name model collapse, insist that synthetic corpora come from a stronger generator or carry a verifiable check, and connect the wall to why compute is being redirected toward post-training and inference-time regimes.
Treat data access as strategy, not procurement: licensing, proprietary corpora and generation pipelines are the durable advantage once FLOPs are purchasable by anyone. Be ready to argue where you would place a multi-year bet under that constraint.
## What the wall actually is Every scaling law is a statement about a joint budget: reach a target loss by spending FLOPs on parameters and tokens together. The laws are silent about where the tokens come from, because when they were fitted, data felt effectively unlimited relative to the compute anyone could afford. That relationship has inverted. Compute keeps growing with capital investment and hardware generations. The stock of high-quality, human-written public text does not — it grows at roughly the rate humans write, which is slow, and much of what exists is duplicated, low-quality, machine-generated, or behind licences. Frontier pretraining runs now consume tens of trillions of tokens, which is a serious fraction of the usable public web. The **data wall** is the regime where you can afford more compute than you have distinct, useful tokens to spend it on. The practical symptom is that the two axes stop being interchangeable. Below the wall, "we have more compute" implies "we can train a better model". At the wall, more compute buys you a larger model on the same data — which pushes you toward the undertrained side of the isoFLOP curve, or toward repeating what you already have. ## Lever one: repeat the data The cheapest response is more epochs. Empirical work on data-constrained training found that repeating a corpus for a small number of epochs is nearly as good as an equal quantity of fresh tokens, with the value of each additional pass dropping off sharply after a handful. Push far past that and you get the classic failure: the model memorises rather than generalises, and validation loss stops tracking capability. So repetition buys you a factor, not an order of magnitude. ## Lever two: generate the data This is where the industry actually went. Synthetic generation covers several distinct techniques with different risk profiles: - **Rephrasing and augmentation** — rewriting existing documents into other styles, difficulty levels or formats, which multiplies effective tokens while keeping the underlying facts anchored to real source text. - **Persona- or prompt-diversified generation** — generating large volumes of varied text by systematically varying who is speaking and about what, to avoid the collapse into a single register that naive generation produces. - **Distillation from stronger models** — using a capable model to produce training text, including worked reasoning, for a smaller or newer one. The well-documented failure mode is model collapse: train recursively on your own unfiltered output across generations and the distribution's tails vanish, variance shrinks, and quality degrades. What makes synthetic data work in practice is that it is not recursive and not unfiltered — it is generated by a stronger model or anchored to real documents, then filtered and verified before it enters the corpus. Synthetic data that can be *checked* (code that compiles, maths with a known answer, claims with a source) is far more valuable than synthetic data that can only be judged. ## Lever three: broaden what counts as data Beyond public English prose there is code, many other languages, licensed and proprietary corpora, transcribed audio and video, and images. Each expands the token pool, and some of it is qualitatively different rather than merely additional — code in particular carries structure and verifiability that plain prose does not. Acquiring it is as much a legal and commercial problem as a technical one, which is part of why data access has become a strategic asset rather than a commodity input. ## The deeper consequence The data wall is one of the reasons the field stopped telling a single-curve scaling story. When fresh pretraining tokens are the scarce input, the marginal return on pretraining compute falls, and other regimes — reinforcement learning after pretraining, and spending compute at inference time on a single hard request — start to look like better places to put the next accelerator. Those regimes have their own scarce inputs (verifiable tasks and reward design; latency and per-request budget), but they are not gated on the supply of human prose. ## What a good answer sounds like Name the inversion — compute grows, high-quality text does not — say that the binding constraint has moved from FLOPs to tokens, list repetition, generation and modality expansion with an honest account of what each buys, and note that generation is only safe when the generator is stronger than the student or the output can be verified. Avoid the two lazy poles: "we have already run out" and "synthetic data solves it". Neither is true.
- What actually goes wrong if you train recursively on unfiltered synthetic data?The distribution narrows across generations. Each pass over-samples the modes of the previous model and loses the tails, so rare facts, unusual phrasings and minority patterns disappear; variance collapses and the model becomes fluent but impoverished. This is model collapse. The practical defence is that production synthetic data is not recursive — it comes from a stronger generator or is anchored to real documents — and is filtered or verified before use.
- Does the data wall apply equally to reinforcement-learning post-training?Not in the same form. RL after pretraining consumes tasks with checkable outcomes rather than raw prose, so it is gated on the supply of verifiable problems, reward design and rollout compute rather than on the stock of human text. That is one reason it became an attractive place to spend compute. It has its own ceiling — writing tasks whose success can be judged automatically is hard — but it is a different ceiling.
- How does data quality interact with the wall?Filtering effectively converts a large low-quality pool into a smaller high-quality one, and better tokens improve loss faster per FLOP, so curation raises the ceiling without adding a single new document. But it is subtractive: aggressive filtering shrinks the pool you are already short of. In the data-constrained regime the interesting work is turning marginal text into useful text rather than discarding it.
saying these in an interview costs you the question
- Claims we have literally run out of text to train on
- Assumes synthetic data is unlimited free training data
- Thinks unlimited epochs over one corpus works fine
- Believes more compute always translates into a better model
- Confuses the data wall with the context-window limit