How do you choose calibration data for post-training quantization?
answer
- a few hundred samples, not millions
- representative beats abundant
- render it through the real prompt template
- narrow deployment, narrow calibration corpus
- never calibrate on your eval set
basics
~20 sUse a few hundred samples that look like production traffic: same domain, same prompt format, realistic lengths. Calibration only sets ranges and error statistics, so a corpus that misses the real distribution produces a model tuned for text it will never see.
solid answer
~50 sCalibration is the only place a post-training method learns anything about your workload, so treat the set as part of the artifact, not as boilerplate. Four things matter. **Size** is usually a few hundred sequences — enough to estimate channel statistics, and beyond that the returns fade while overfitting risk for error-compensating methods grows. **Domain match** matters most when deployment is narrow: quantizing an on-prem defect classifier for a factory edge box using five hundred real line-inspection transcripts is meaningfully better than using generic web text. **Format match** is the one teams forget — an instruction-tuned model in production always sees a chat template with a system prompt, so calibrating on raw prose calibrates the wrong activation distribution. **Length** should span the sequence lengths you actually serve. One modern wrinkle: with mixture-of-experts models, a small or narrow set may never route tokens to some experts, leaving those experts' scales estimated from almost nothing. Diversity is not optional there.
go deeper
Know that post-training quantization needs a small sample of representative text — a few hundred sequences — and that it is used to measure ranges, not to train the model.
Explain what the calibration pass actually collects and why matching the production domain, prompt format and sequence length changes the resulting scales. Know the rough set size and why it is small.
Bring the failure mode and the diagnosis: a model that scores fine on generic benchmarks and poorly on your suite, held-out task evals kept disjoint from calibration, saturation monitoring for static ranges, and expert-coverage checks on mixture-of-experts models.
Own calibration data as a governed asset — provenance, retention, whether regulated content may be processed inside the quantization pipeline, and a policy that every quantized artifact ships with a record of the corpus and eval it was validated against.
## Why the set matters at all Post-training quantization has exactly one channel through which your specific deployment can influence the result: the calibration pass. Whatever the algorithm — plain range estimation, error compensation, salience scaling, importance weighting — it looks at the activations produced by those samples and nothing else. The quantized checkpoint is, in a real sense, fitted to that corpus. Teams routinely spend a week choosing between quantization algorithms and thirty seconds choosing the data those algorithms will see, which is backwards. ## Size The practical range is small: on the order of 128 to 512 sequences at the model's typical serving length. The quantities being estimated — per-channel magnitudes, per-layer output error, curvature — converge quickly, and past a few hundred sequences the marginal sample changes almost nothing. There is also a downside to piling on more: methods that fit a correction to the calibration set (the error-compensating family especially) can overfit an estimate built from a finite sample, and a huge but unrepresentative set makes that worse rather than better. More data is not the lever; more *representative* data is. ## Domain match The narrower the deployment, the more domain match pays. Consider quantizing a defect-classification model to run on an on-prem factory edge box with no internet egress. The traffic it will serve is inspection reports and line telemetry in a house vocabulary, with recurring part numbers and terse operator shorthand. Calibrating with five hundred samples drawn from actual line-inspection logs gives the algorithm the activation distribution it will really encounter; calibrating on a generic web corpus gives it a plausible-looking distribution from a different world. Conversely, for a general assistant serving open-ended traffic, a broad and diverse corpus is the honest choice — matching *the* distribution means matching a wide one. The air-gapped case adds a constraint worth naming in an interview: the calibration data cannot leave the plant, so the quantization run happens inside the perimeter, and the resulting scales encode statistics of proprietary data. That is usually fine, but it is a data-governance fact, not a neutral build step. ## Format match This is the most common unforced error. An instruction-tuned model in production never sees bare prose: it sees a chat template, a system prompt, tool definitions, sometimes retrieved documents pasted into a long context. Those structural tokens produce their own characteristic activation patterns. Calibrating on raw text from a generic dataset therefore calibrates a distribution the served model does not visit. The fix is trivial — render your calibration samples through the exact same prompt template the application uses — and the omission is easy to make. Length deserves the same treatment. If production requests routinely run tens of thousands of tokens, calibrating exclusively on short snippets misses the activation behaviour of long-context positions. ## Mixture-of-experts coverage Almost all current large open-weight models are mixture-of-experts, and this changes the calibration calculus. Each token is routed to a small subset of experts, so a calibration set of a few hundred narrow sequences may leave many experts touched by very few tokens — or none. Their scales then rest on a near-empty sample. Practical responses are to enlarge and diversify the set specifically to spread routing, to check per-expert token counts during the calibration pass as a diagnostic, and to be suspicious of a run where the coverage histogram has a long empty tail. ## Contamination and leakage Two hygiene rules. Do not calibrate on your evaluation set — the quantized model is fitted to those activations, and your measured degradation will be optimistically biased. And treat calibration samples as data you have processed: if they contain personal or regulated content, the pipeline that touches them inherits that classification even though the output is only a set of scales. ## Detecting a mismatch after the fact The characteristic symptom is a quantized model that looks fine on generic perplexity or a public benchmark and degrades on your own task suite — precisely because the generic measurement matches the calibration corpus and your traffic does not. The defence is to hold out a task-specific eval set that was never used for calibration, and to compare the full-precision and quantized models on it rather than on a generic proxy. For activation quantization with static ranges, monitoring saturation rates in production is the direct signal that the calibrated ranges were too narrow. ## What a strong answer sounds like Size, domain, format, length, and — for current architectures — expert coverage; plus the discipline of keeping calibration and evaluation data disjoint. Ending on the failure symptom, rather than the recipe, is what marks the answer as coming from someone who has shipped one of these.
- Why not simply use as much calibration data as you can afford?Because the statistics being estimated converge within a few hundred sequences, so extra samples buy almost nothing, while the run gets slower. Error-compensating methods can also fit their corrections to whatever set they are given, so a large but unrepresentative corpus actively steers the result toward the wrong distribution. The lever is representativeness, not volume.
- How would you notice, after deployment, that the calibration set was mismatched?The signature is a quantized model that holds up on generic perplexity or public benchmarks but degrades on your own task suite — because the generic measure resembles the calibration corpus and your traffic does not. Hold out a task-specific eval set that never touched calibration and compare full-precision against quantized on it. With static activation ranges, rising clip or saturation rates in production are the direct signal.
- Does calibration data choice matter differently for a mixture-of-experts model?Yes, and it is the newest wrinkle. Routing sends each token to a small subset of experts, so a few hundred narrow sequences can leave many experts with almost no tokens and therefore scales estimated from nearly nothing. Deliberately diversify the set to spread routing, and inspect per-expert token counts during the calibration pass as a coverage check before trusting the artifact.
saying these in an interview costs you the question
- Says any generic text corpus works equally well
- Assumes more calibration data is always better
- Calibrates on raw prose for a chat-templated model
- Reuses the evaluation set as calibration data
- Treats calibration as retraining rather than statistics gathering