When does synthetic or teacher-distilled text belong in a pretraining mix?
answer
- scarcity, not compute, is the constraint
- a generator cannot invent knowledge
- anchoring and verification break the loop
- diversity loss hides from the loss curve
- hold out fresh human text
basics
~20 sGenerated text buys coverage where real text is scarce or badly written, and it is now standard practice. It also inherits the generator's blind spots, so ground it in real source documents, cap its share, and measure diversity rather than loss alone.
solid answer
~50 sBy the mid-2020s the binding constraint on pretraining stopped being compute and became high-quality text, so labs generate their own: rephrasing existing documents into cleaner or more varied styles, expanding scarce domains, and distilling a stronger teacher model's outputs into a smaller student's corpus. Where the generation is *anchored* — a real source document rewritten, or a derivation checked by an external verifier — it adds genuine signal, because the grounding carries information the generator did not invent. Where it is unanchored, generated text mostly resamples what the generator already believes: the distribution narrows toward its typical phrasing, tail phenomena disappear, and its factual and stylistic biases are amplified rather than corrected. This is contested ground as of mid-2026, with no settled optimal share. The defensible position is procedural: anchor generations on real inputs, keep a substantial real-data share, hold out freshly collected human text for evaluation, and track diversity metrics alongside loss.
go deeper
Know that a lot of modern training text is generated by other models, and that the obvious catch is a model cannot teach what it does not already know.
Distinguish the forms — rephrasing real documents, expanding thin domains, distilling a stronger teacher — and explain why anchoring generation on a real source or a verifier is safer than free sampling.
Show you can detect the failure: diversity collapse and inherited bias do not appear in held-out loss, so name the ablations, tail-sensitive evaluations and provenance tracking you would run before committing a full budget.
Own the policy. Set the real-data floor, the provenance and rollback requirements, the independence of the evaluation set from the generator, and say plainly that the optimal share is unsettled and must be earned experimentally for your workload.
## Why this question exists at all For most of the field's history the assumption was that text is abundant. That assumption has broken down: the volume of high-quality, permissively usable, well-written text is finite, while appetite for training tokens is not. So generated text moved from a curiosity to a normal ingredient of pretraining mixes. The interview question is never "is synthetic data allowed" — it is "under what discipline", because the failure mode is slow and does not show up in the loss curve. ## The three things people mean by synthetic pretraining data **Rephrasing and restyling.** Take a real document and have a model rewrite it — as clean prose, as a question-and-answer pair, as an explanation for a different audience. The facts come from the source; only the presentation is generated. This is the safest form, because the information content is anchored. **Coverage expansion.** Generate text for a domain the crawl barely covers, or vary a scarce pattern across many personas, entities or settings so the model sees the underlying structure more often. Value depends entirely on whether the generator actually knows the domain. **Teacher distillation.** Run a stronger model to produce text — including worked reasoning — and train a smaller model on it. This is standard production practice for making small models punch above their size, and it transfers the teacher's *behaviour* efficiently. ## What generated text can and cannot add Information-theoretically, a generator cannot create knowledge it does not have. If you sample from a model and train on the samples, you are re-presenting its own distribution, and repeated rounds of that shrink variance: rare constructions, unusual phrasings and long-tail facts appear less and less, until outputs converge on the generator's central tendencies. This degeneration is why unanchored, recursively generated corpora are dangerous. What generated text *can* add is real: - **Restructured access to knowledge already in real data.** Presenting the same fact in several forms genuinely improves whether a model can use it, even though no new fact entered. - **Signal injected from outside the generator.** A grounding document, a compiler, a test suite, a solver or a symbolic checker supplies information the generator did not have. Generation plus verification is not a closed loop and does not degenerate the same way. - **Density.** Real crawl text is mostly filler; generated text can be uniformly on-topic, which raises the useful-signal-per-token of a fixed budget. ## The risks a principal is expected to name - **Distribution narrowing.** Diversity loss is invisible in average loss and shows up as blandness, register collapse and poor tail coverage. - **Inherited blind spots.** The student cannot exceed the teacher on anything the teacher gets systematically wrong, and it copies the teacher's stylistic tics and refusal patterns as if they were facts about language. - **Error amplification.** A confident mistake generated a million times becomes a strongly learned prior, and it is far harder to trace than a bad source you can delete. - **Contamination.** If generation prompts are derived from evaluation-adjacent material, benchmark content re-enters the corpus laundered through the generator, where string-matching decontamination will not catch it. - **Provenance and legal exposure.** Teacher outputs carry terms of use, and "a model wrote it" does not clear the rights on what the model reproduced. - **Evaluation circularity.** Judging a synthetically trained model with the same family of models that generated the data hides exactly the failures you care about. ## The discipline that makes it defensible Anchor generation on real inputs wherever possible, so each generated document traces back to a source. Prefer pipelines with an external verifier over pipelines that only sample. Keep a substantial floor of real human text in the mixture rather than letting the synthetic share drift upward run over run. Track provenance per document so a bad generator can be rolled back. Measure diversity explicitly — vocabulary and n-gram coverage, register spread, performance on tail domains — not just held-out loss. And reserve an evaluation set of freshly collected human text that no generator has seen, judged by humans or by a model family unrelated to the generator. ## Being honest about the state of the art As of mid-2026 there is no published consensus optimal ratio, and the strongest evidence is that quality of the pipeline dominates the ratio: well-anchored, verified generation at a high share can beat poorly grounded generation at a low one. A candidate who states a confident universal number is guessing. The right answer describes the experiment that would settle it for your workload — mixture ablations at small scale, with tail-sensitive evaluations — and the monitoring that would catch degeneration before a full run is wasted.
- Why does grounding a generation on a real document change the risk profile?Because the information then comes from outside the generator. Rephrasing a real specification carries that specification's facts; the model only supplies presentation. Unanchored sampling instead resamples the generator's own distribution, so repetition narrows it. The same logic covers external verifiers — a compiler, test suite or solver injects signal the generator did not have, which is why generate-and-verify pipelines do not degenerate in the same way.
- How would you detect degeneration before spending a full training run?Run small-scale mixture ablations with identical budgets and compare on tail-sensitive evaluations, not just held-out loss, which is insensitive to diversity collapse. Track lexical and n-gram coverage, register and domain spread, and performance on minority languages and niche domains. Evaluate with human-collected data the generator never saw, judged by an unrelated model family or by people, to avoid circularity.
- What makes teacher distillation into a pretraining corpus different from ordinary data collection?Provenance and ceiling. The student inherits the teacher's systematic errors, stylistic tics and refusal behaviour as if they were properties of language, so it rarely exceeds the teacher on those axes. Teacher outputs also carry licensing terms, and the fact that a model produced the text does not clear rights on what it reproduced. Both need tracking per document.
saying these in an interview costs you the question
- Claiming synthetic data always causes model collapse
- Assuming generated text adds knowledge the generator lacks
- Quoting a universal optimal synthetic-to-real ratio
- Evaluating a distilled model with its own teacher family
- Ignoring that generation can launder benchmark text into the corpus