skip to content

Why must a TF-IDF vocabulary and its idf values be fitted on the training fold only?

level: middleimportance: must knowfreq 55%

answer

  1. preprocessing that learns is not preprocessing
  2. which columns exist is learned from data
  3. N and df are counted over something
  4. split before anything is counted
  5. held-out documents voted on their own weights

basics

~20 s

The vocabulary and the idf values are learned parameters. Building them over the whole corpus before splitting lets held-out documents shape the features that describe them, so validation scores come out optimistically biased and overstate what production will do.

solid answer

~50 s

The counting step learns two things from data: which terms become columns, and the document frequency behind each term's idf. Both are fitted parameters, exactly like a scaler's mean. If you build them over the full corpus and then split, every held-out document has already voted on the column layout and on the idf denominators, so information from the evaluation set has leaked into the representation of the evaluation set. The score you measure is then better than the score you will get on genuinely new text. The fix is mechanical: inside each cross-validation fold, fit the vocabulary and idf on the training part, then transform the validation part with those frozen statistics — unseen terms simply fall outside the vocabulary and are dropped. The same frozen statistics ship with the model and are applied to production traffic.

go deeper

for a junior

Remember the one-line rule: split first, then count. Be able to name the two things the counting step learns — the list of columns and the idf numbers — and say that both come from training data only.

for a middle

Explain why this counts as leakage without any labels involved: held-out documents influenced the feature space that describes them. Walk through the correct order of operations inside a cross-validation fold.

for a senior

Show the operational side: frozen statistics ship with the model artefact, unknown terms are dropped at serving time, and the out-of-vocabulary rate is worth monitoring as a drift signal.

for a principal

Own the guardrail rather than the instance. Argue for a pipeline structure where fitting anything outside a fold is impossible by construction, and for review habits that catch the same mistake in imputation and scaling.

## The vectoriser is a model It is easy to think of turning text into a matrix as preprocessing — plumbing that happens before the modelling starts. It is not. Two things are estimated from data: 1. **The vocabulary.** Which distinct terms get a column, and in which order. This is decided by scanning a corpus. 2. **The idf values.** One number per term, `log(N / df(t))`, where `N` is the number of documents scanned and `df(t)` is how many of them contained the term. Both are *fitted parameters*. They must be estimated on training data, frozen, and then applied unchanged to anything you want an honest estimate for — validation folds, the test set, tomorrow's traffic. ## What the leak actually is Suppose you have 20,000 support tickets that you want to route to the right queue. The tempting workflow is: build the TF-IDF matrix over all 20,000 tickets, then split into train and test, then fit a classifier. Two channels of information now flow backwards from the test set: - **Column selection.** A term that occurs only in test tickets got a column because those tickets existed. If you also applied a minimum-document-frequency cut-off, test documents helped a term clear the threshold — or failed to save one that the training data alone would have dropped. The feature space has been shaped by data you are pretending not to have seen. - **The idf values.** `N` and every `df(t)` were counted over all 20,000 tickets. So the weight given to `refund` in a test ticket depends partly on how often `refund` appears in other test tickets. In effect the test set has told you how distinctive its own words are. The result is a validation score that is optimistically biased: it measures performance on documents whose features were partly designed around them. ## How big is the damage? Usually modest, sometimes not. The leak is worst when: - **The corpus is small.** With 500 documents, one held-out document moves `df` noticeably and can single-handedly put a term into the vocabulary. - **Rare terms carry the signal.** Idf deliberately hands the biggest multipliers to low-`df` terms, which are exactly the terms whose statistics are least stable and most influenced by individual held-out documents. - **You aggressively prune.** Any document-frequency threshold makes the leak structural rather than numerical: a whole column exists, or doesn't, because of the test set. It is rarely a catastrophic ten-point inflation. It is typically a small, systematic optimism — which is worse in one sense, because it is invisible: the model still looks plausible, so nothing prompts you to investigate. And the reason to be strict is not the size of this particular leak; it is that the same habit applied to target encoding, imputation or scaling produces much larger ones. ## Doing it correctly Inside a `k`-fold loop, for each fold: 1. Split first. Nothing derived from data crosses the split before it is made. 2. Fit the vocabulary and idf on the training part of that fold only. 3. Transform the training part and the held-out part with those same frozen statistics. 4. Train and score. The vocabulary therefore differs slightly from fold to fold, which is correct — it is part of the pipeline whose variability you are trying to measure. Averaging over folds then estimates the variance of the *whole* procedure, not just of the classifier on a fixed feature space. ## Terms the training fold never saw A held-out document will contain terms absent from the fitted vocabulary. They have no column, so they are silently dropped: the document is represented only by the terms the training fold knew about. This is the honest outcome — a model deployed tomorrow faces exactly this. Two consequences worth stating in an interview: - You cannot invent a column at transform time, because the model's weight vector has a fixed length. Growing the vocabulary means retraining. - If a large share of a held-out document's terms are unknown, that is a signal in its own right: the training corpus is unrepresentative, or the domain has drifted. Tracking the out-of-vocabulary rate on live traffic is a cheap drift monitor. A hashing-style representation avoids the fixed-vocabulary problem by mapping terms to a fixed number of columns arithmetically, at the cost of collisions — but note that the idf statistics, if you use any, are still fitted quantities and still belong to the training fold. ## Shipping it Whatever you fitted has to travel with the model: the term-to-column mapping and the idf vector. A model artefact without them is useless, and a mismatch between the vectoriser used at training time and the one used at serving time silently scrambles every column. Treat them as part of the model's version, not as a script that gets re-run.

  • What happens to a term in a held-out document that the training fold never saw?
    It has no column, so it is dropped and the document is represented only by known terms. That is the honest behaviour, because a deployed model faces exactly this. You cannot add a column at transform time — the trained weight vector has a fixed length — so absorbing new vocabulary means retraining. A rising out-of-vocabulary rate on live traffic is a useful drift alarm.
  • Does the same training-fold rule apply to a minimum-document-frequency cut-off?
    Yes, and more strongly. A document-frequency threshold decides which terms exist as columns at all, so counting it over the full corpus lets held-out documents create or destroy features. That is a structural leak rather than a numerical one, and it is the version of this mistake that does the most damage on small corpora.
  • Roughly how large is this leak in practice, and why be strict about it anyway?
    Usually small — a systematic optimism rather than a collapse — and largest on small corpora where rare, high-idf terms carry the signal. Be strict anyway for two reasons: the bias is invisible, so nothing prompts you to look for it, and the same shortcut applied to imputation, scaling or target statistics produces leaks that are anything but small.

saying these in an interview costs you the question

  • Calls vectorising preprocessing, so exempt from the split
  • Fits idf on all data, then splits for training
  • Says the leak is harmless since no labels were used
  • Rebuilds the vocabulary at serving time from live traffic
  • Ships a model without the idf vector and column mapping

context