Why do synthetically expanded fine-tuning sets collapse in diversity, and how do you prevent it?
answer
- generators sample their favourite scenarios
- row count hides a narrow distribution
- per-row quality scores stay high throughout
- generate against a coverage grid with quotas
- cluster embeddings before you train
basics
~20 sFree-running generation keeps sampling the generator's favourite mode, so thousands of rows re-tell a handful of scenarios. Prevent it by generating against a structured grid of case attributes with a quota per cell, and measure diversity by clustering row embeddings before you train.
solid answer
~50 sLanguage models are mode-seeking: ask one for a thousand example cases and you get its ten favourite stories a hundred times over, in different words. Expanding 200 veterinary triage seeds Self-Instruct style into 8,000 rows, we found roughly 40% were some variant of "a dog ate something it shouldn't" — a real collapse, invisible in the row count and invisible to a per-row quality judge, because each individual row was fine. The fix is to stop letting the generator choose the scenario. Enumerate the axes that matter — species, body system, urgency, how articulate the owner is — take the cross product, and give each cell a quota, so the plan constrains the sampling. Then vary within the cell with Evol-Instruct style depth and breadth rewrites for the harder multi-symptom cases. Measure before training: cluster the row embeddings and look at cluster count, cell fill rates and mean pairwise distance. Raising temperature is not a fix — it changes the wording, not the mode.
code
python · 13 linesspecies = ["dog", "cat", "rabbit", "parrot"]
systems = ["gastrointestinal", "respiratory", "dermatological", "neurological"]
urgency = ["emergency", "same-day", "routine"]
cells = [(s, y, u) for s in species for y in systems for u in urgency]
quota = 8000 // len(cells)
for s, y, u in cells:
prompt = (
f"Write {quota} distinct owner-reported cases: a {s} with a "
f"{y} problem, triage urgency '{u}'. Vary owner detail level."
)
generate(prompt)go deeper
Know that generating training data from a model tends to produce many near-copies of a few scenarios, so a big synthetic file can be far less varied than its row count suggests.
Explain the mechanism — mode-seeking decoding, self-conditioning on accepted outputs, seed bias — and the fix of generating per cell against an explicit grid of case attributes with quotas.
Demonstrate that you would measure the distribution before training with embedding clustering and cell fill rates, and that you would treat an eval set drawn from the same generator as unable to reveal the problem.
Own the coverage specification as a first-class artefact of the programme, budgeted and reviewed like a schema, so that generation is constrained by a plan rather than by whatever the generator finds typical.
## The failure, concretely You need roughly 8,000 instruction pairs for a veterinary triage assistant. You write 200 seed cases by hand, expand them Self-Instruct style — sample a few seeds into the prompt, ask for new cases in the same spirit, add the good ones back to the pool — and you get your 8,000 rows. Every row reads well. A judge scores them highly. Then somebody clusters them and finds that about 40% describe a dog that ate something it should not have, and that four or five scenarios account for the majority of the file. The dataset is not 8,000 examples. It is a few hundred examples, each repeated with different names and adjectives, and the model you train on it will be excellent at the dog-ate-something case and thin everywhere else. ## Why generators collapse toward a mode - **Mode-seeking sampling.** Decoding favours high-probability continuations. Asked for "a case", the model returns the most typical case, over and over, because typicality is exactly what it optimises for. - **Self-conditioning.** Pipelines that feed accepted outputs back into the prompt pool amplify whatever the early rounds happened to produce. Bias compounds; the distribution narrows with each round. - **The seed set defines the horizon.** If 30 of your 200 seeds are ingestion cases, the generator reads ingestion as the theme and expands it hardest. Human seed-writing is itself biased toward the vivid and the recent. - **Low-entropy prompts.** One generation prompt reused thousands of times gives the model no reason to move. Diversity has to come from what varies in the prompt, and if nothing varies, nothing does. ## Why it is easy to miss Every signal you would naturally look at is blind to it. Row count is high. Per-row judge scores are high, because each row is individually correct and well-formed. Loss curves look normal. And if the evaluation set was drawn from the same generator, it collapses in exactly the same way, so the fine-tuned model scores beautifully on the very cases it is overfitted to. Diversity collapse is a distribution defect, and only distribution-level measurement finds it. ## Prevention: constrain generation with a coverage plan The reliable fix is to move the choice of *what* to generate out of the model and into an explicit plan. - **Build the grid.** Enumerate the axes that make cases genuinely different — for triage: species, body system, urgency band, owner's level of detail, presence of a chronic condition. Take the cross product; that grid is your specification. - **Quota each cell.** Divide the target volume by the number of cells and generate per cell with the attributes injected into the prompt. The generator now varies wording within a cell you chose, instead of choosing the scenario itself. - **Seed from the grid, not from a favourites list.** Human-written seeds should be sampled to fill the grid too, or they will reintroduce the same skew. - **Vary the surrounding context.** Different personas, registers and lengths in the generation prompt add real variance; a first-time rabbit owner writes nothing like a breeder. - **Deepen deliberately.** Evol-Instruct style rewrites — take a simple case and make it harder by adding a constraint, a co-occurring symptom, a contradictory detail — produce the difficult multi-symptom rows that free-running generation almost never emits, because difficult cases are not typical cases. - **Ground in real material where you can.** Sampling anonymised real intake records as scenario skeletons anchors the distribution to production instead of to the model's priors. ## Measurement, before training - Embed every row and cluster: report cluster count, the size of the largest clusters, and mean pairwise distance. A handful of clusters holding most of the mass is the collapse signature. - Report **cell fill rates** against the grid — empty and overflowing cells are actionable in a way that an aggregate score is not. - Cheap lexical proxies (distinct n-gram ratio, self-similarity between rows) catch the crude cases and cost nothing to run on every generation batch. - Keep a human-written check set that the generator never touched, so you always have one view of quality that cannot inherit the generator's blind spots. ## Repairing a collapsed set Semantic deduplication plus rebalancing — collapse near-duplicates, downsample the dominant clusters, then top up the empty cells with targeted generation. Expect the honest row count to fall a lot. That is the correct outcome: you are trading a large fake dataset for a smaller real one, and the smaller real one trains a better model. ## What does not work Raising the temperature. It perturbs word choice while leaving the scenario distribution intact, and past a point it degrades correctness while the collapse survives. Likewise, simply generating more rows: the extra rows arrive in the same proportions, so absolute counts in the tail grow slowly while the skew stays exactly where it was.
- Your per-row judge scores are high across the whole synthetic set. Why does that not rule out collapse?A judge scores each row on its own merits — correctness, specificity, format — and a thousand well-written variants of one scenario all pass. Collapse is a property of the set, not of any row in it, so it can only be seen with distribution-level measurement: embedding clusters, cell fill rates against a coverage plan, or lexical self-similarity across rows.
- Would raising the generation temperature fix this?No. Temperature changes wording, not the underlying scenario distribution, so you get more varied prose about the same handful of cases, and past a point correctness degrades while the collapse survives. Diversity has to be injected through what varies in the prompt — attributes drawn from a coverage grid, different personas, deliberate difficulty rewrites — not through the decoding parameters.
- How does collapse in the training set show up after deployment?The model is strong on the dominant scenarios and noticeably weak everywhere else, and it tends to force unusual inputs into the shapes it saw most — reading an unfamiliar presentation as the case it over-learned. If the eval set came from the same generator it will not show this at all, which is why a human-written check set covering the thin cells is worth keeping.
Ask one person to invent a thousand example cases and you will get their ten favourite stories a hundred times over. Hand them a checklist of case types with a quota for each and they are forced off autopilot.
saying these in an interview costs you the question
- Higher temperature will fix synthetic diversity
- High per-row judge scores prove the set is diverse
- Generating more rows fixes the skew
- Deduplication alone rebalances a collapsed dataset
- The seed examples do not shape what gets generated