How many training examples does a narrow supervised fine-tune actually need?
answer
- quality and coverage over raw count
- hundreds for a narrow task
- tens of thousands for broad behaviour
- inconsistent targets cap the ceiling
- duplicates only reweight one example
basics
~20 sA few hundred to a few thousand consistent examples usually move a narrow, well-defined task; broad behaviour change needs tens of thousands. Coverage of the real input distribution and consistent labelling matter far more than raw row count.
solid answer
~50 sVolume is the wrong first question — the useful one is how many distinct behaviours you are teaching and whether each is covered. For a narrow task with one output shape, such as turning a pet-owner's description into a triage urgency band plus a short rationale, a few hundred to a couple of thousand carefully written rows will usually get you most of the gain. Broad behaviour change across many task types is where you get into tens of thousands. What kills a fine-tune far more often than a small row count is **inconsistency**: if half the examples hedge and half give a flat verdict, the model learns to do both at random. Duplicating rows to inflate the count adds nothing. Spend the next thousand rows on slices the set does not yet cover, not on more of the type you already have.
go deeper
Be ready to say that a narrow task often needs only hundreds to a couple of thousand clean examples, and that consistent, correct target outputs matter more than the row count.
Explain why the model imitates the variance in the targets, so mixed formats and label noise put a ceiling on quality, and why duplicated rows only reweight one example instead of adding information.
Show how you would allocate a data budget: enumerate the input axes, find the empty slices, run a data-scaling curve, and argue for cleaning versus collecting based on where the errors actually sit.
Own the framing that dataset size is a cost decision against alternatives — prompting, retrieval, a bigger base model — and that curation effort, not row count, is the line item that decides whether the programme pays for itself.
## What the row count is really measuring A supervised fine-tuning dataset is a set of pairs: an input the model will see in production, and the exact output you want it to produce. When someone asks "how much data do I need", the honest answer is that the count is a proxy for two things that actually determine the result — how many distinct behaviours you are asking the model to learn, and how well the examples cover the range of inputs it will meet. A thousand rows that all show the same behaviour on the same kind of input is one behaviour taught a thousand times. ## Rough volumes people actually work with - **Narrow, single-format tasks** — classify an inbound veterinary triage message into an urgency band and emit a fixed set of fields; rewrite an internal ticket into a standard summary. A few hundred to roughly two thousand clean rows is typically enough, and the curve often flattens well before the top of that range. - **Broader behaviour** — a whole assistant persona, multi-turn conversational style, many task types under one model. Tens of thousands is the honest order of magnitude, which is why teams reach for synthetic expansion rather than hand-writing at that scale. These are starting points, not laws. The number that matters is rows *per distinct slice*, and a set of 800 rows spread evenly over eight real input types is usually worth more than 5,000 rows concentrated on the two most common ones. ## Why consistency dominates volume A fine-tune imitates the distribution of the outputs it is shown, variance included. If some examples answer in two sentences and some in six, some open with a caveat and some do not, some emit the fields in one order and some in another, the model does not average those into a canonical form — it learns that all of them are acceptable and picks one at random per request. Label noise behaves the same way: a small fraction of wrong targets puts a ceiling on the quality the fine-tune can reach, and no amount of extra rows lifts that ceiling. This is why a written spec for the target output, applied by whoever (or whatever) produces the rows, is worth more effort than another data source. If the rows are produced by several people or several generation prompts, review a sample specifically for cross-row consistency, not just per-row correctness. ## Coverage: matching the real input distribution Before counting rows, enumerate the axes your inputs vary along, then check the set against them. For a triage assistant that might be species, body system, urgency, and how much detail the owner gives. Empty cells in that grid are where the fine-tuned model will behave worst, and they are invisible if you only look at aggregate counts. The cheapest quality win in most projects is filling the sparse cells rather than growing the dense ones. Coverage should reflect production traffic, but not slavishly: rare-but-costly inputs deserve more representation than their natural frequency, because the loss you care about is not uniform across cases. ## Duplication is not data Copying a row to reach a bigger number simply reweights that example — the model sees the same gradient signal repeatedly and drifts toward that row's specifics. Near-duplicates do the same thing more subtly, which is why deduplication is part of preparing the set rather than a nicety. "8,000 rows" that dedupe down to 3,000 distinct behaviours is a 3,000-row dataset with a skew problem. ## Deciding where the next budget goes A practical technique is a data-scaling curve: train on 25%, 50% and 100% of the set and compare quality at each point. If the curve is still climbing steeply at the top, more of the same data is a reasonable buy. If it has flattened, extra volume is wasted and the remaining errors are concentrated in slices the data does not represent — go and write rows for those, or make the existing rows more consistent. ## When more data is not the answer at all If the model's failures are about *knowing* things — facts it was never trained on, information that changes weekly — a bigger fine-tuning set does not fix it, and training on unfamiliar facts tends to increase confident errors. Fine-tuning reliably teaches form, style and task procedure; supplying knowledge is a retrieval problem. Recognising that distinction in an interview is worth more than any specific row count.
- How would you tell whether more rows are still buying you anything?Build a data-scaling curve: train on 25%, 50% and 100% of the set and compare quality at each point. A curve still climbing steeply at the top says more of the same data is worth buying. A flat curve says the remaining errors live in slices the data does not cover, so the next budget should go to new slices or to cleaning inconsistent targets rather than to volume.
- You only have 40 real examples. What do you do with them?Do not train on 40 rows and hope. Use them as seeds — they define the shape and the voice — and expand synthetically against a coverage plan, keeping a handful of the originals untouched as a human-written check set. If the task is narrow enough that prompting already half-works, it is often cheaper to stay with prompting until you have enough real traffic to curate from.
- Does the same volume guidance hold for adapter-based tuning as for full-parameter training?Broadly yes on the data side: the size and coverage of the set is driven by how many behaviours you are teaching, not by which parameters you touch. What changes is the failure mode at the top end — a small adapter has limited capacity, so a very large and diverse dataset can exceed what it can absorb. That is a capacity question, though, not a reason to curate differently.
saying these in an interview costs you the question
- More rows always make a fine-tune better
- Scraping 100k noisy rows beats 500 curated ones
- You need millions of examples, like pretraining
- Duplicating examples counts as extra data
- Volume compensates for inconsistent target style
- Adding rows will teach the model missing facts