skip to content

How do exact and near-duplicate rows distort a training set and its class balance?

level: juniorimportance: should knowfreq 50%

answer

  1. the same information, counted twice
  2. the loss sums over rows
  3. effective sample size below row count
  4. class frequencies become fiction
  5. exact match misses paraphrases

basics

~10 s

Duplicated rows count the same information several times. They upweight those examples in the loss, inflate one class's apparent frequency, and make the real sample size much smaller than the row count suggests.

solid answer

~50 s

A duplicate is not new evidence, but every training procedure treats it as if it were. If a scraped product catalogue lists the same item forty times under forty different titles, that item contributes forty times the gradient of a unique item, so the model is pulled toward whatever is easy about it. The damage spreads to every decision made from counts: if the duplicated rows all carry the same class, the class balance you measure is fiction, so resampling ratios and class weights are computed from numbers that do not describe the world. The effective sample size also collapses -- 100,000 rows containing 20,000 distinct records give you the statistical power of roughly 20,000, so any variance estimate that assumes independent rows is too optimistic. Exact-match dedup catches only the trivial case; near-duplicates such as copy-pasted reviews need normalisation plus a similarity threshold you tune by eyeballing a sample.

go deeper

for a junior

Be ready to say plainly that a duplicated row is counted again in the loss and again in the class counts, and that exact-match checks miss near-duplicates. Naming normalisation as the cheap first step is enough at this level.

for a middle

Explain the mechanics: duplication equals sample weighting, effective sample size falls below the row count, and every count-derived setting such as class weights inherits the error. Describe a concrete near-duplicate detection ladder.

for a senior

Show you would audit for duplication before trusting any dataset statistic, and that you can tell an ingestion artefact from real repeated behaviour. Explain how you choose and validate a similarity threshold on real records.

for a principal

Own the tradeoff between aggressive deduplication, which quietly deletes real minority records, and permissive rules, which leave the base rate wrong. Decide where deduplication belongs so every team reads the same prevalence numbers.

## What a duplicate actually is Two kinds of repetition get lumped together and they are not the same thing. - **Record duplication** -- the same underlying entity or event has been written into the dataset more than once, usually through a join fan-out, a re-run of an ingestion job, or a scrape that pulled the same page from several URLs. This is an artefact. - **Legitimate repetition** -- two different customers really did buy the same product, on different days, and both rows are real observations. This is signal. Everything below is about the first kind. Deleting the second kind is itself a bug, because it destroys the true frequency of an event. ## Why the loss cares Almost every learner minimises a sum (or mean) over rows: `loss = (1/n) * sum_i L(y_i, f(x_i))`. A row that appears `k` times enters that sum `k` times, which is exactly the same as giving one copy a weight of `k`. The model does not know the copies are copies -- it only sees that this region of feature space is very densely populated and very consistently labelled, so it will bend the decision boundary to be right there, at the cost of regions represented by a single honest row. A tree will happily create a pure leaf around the duplicated block; a linear model will let those rows dominate the coefficient estimate. ## Why the counts you report are wrong Suppose 8% of your rows are labelled `fraud`, and you configure a resampling ratio or a class weight from that 8%. If a batch job duplicated a slice of the fraud rows five times, the real prevalence might be closer to 2%. Every downstream decision made from the 8% -- the sampling ratio, the class weight, the prior you assume when calibrating, the base rate against which you judge a lift -- is now anchored to a number produced by a pipeline defect. The same applies to a class that looks large and easy: near-duplicate reviews copied across a sentiment dataset make that sentiment class look both more common and more separable than it is, because the model keeps meeting text it has already seen. ## Effective sample size Statistical claims about a dataset assume independent observations. Duplication breaks that assumption in the most direct way possible: the extra copies carry zero additional information. A table of 100,000 rows built from 20,000 distinct records behaves, for the purpose of estimating any quantity, roughly like 20,000 rows. Confidence intervals, standard errors and any variance estimate computed with `n = 100,000` will be too narrow, so you will believe differences are real that are not. When someone quotes a dataset size in an interview, a good instinct is to ask how many *distinct* entities it covers. ## Finding the near-duplicates Exact equality is the easy 20%. Real catalogues, review corpora and CRM extracts are full of rows that differ by whitespace, casing, a trailing SKU code, a reordered address, or a paraphrase. A workable ladder: 1. **Normalise first** -- lowercase, strip punctuation and whitespace, unify units and date formats, sort multi-valued fields. A large share of near-duplicates become exact duplicates after this step alone. 2. **Build a canonical key** where the domain gives you one -- a normalised barcode, a stripped email, a phone number in a single format. 3. **Use a similarity measure** for free text: overlap of character or word n-grams (shingles) between two records, with a threshold. To avoid comparing every pair against every other pair, group candidates first by a cheap key (first token, a hash bucket, a coarse category) and only compare within groups. 4. **Hand-check a sample at the threshold.** The threshold is a business decision, not a mathematical one: too loose and you delete genuinely distinct short records, too tight and you keep the copies. ## Dedup or reweight? Dropping all but one copy is the default and usually right for record duplication. Two nuances are worth voicing. First, if you are unsure whether repetition is artefact or signal, downweighting the group to a total weight of one is gentler than deletion and reversible. Second, whatever you do, do it *before* you compute class balance, prevalence, sampling ratios and any headline dataset statistic -- otherwise you clean the data but keep making decisions from the dirty numbers.

  • How would you catch near-duplicates that an exact-match check misses?
    Normalise aggressively first -- casing, whitespace, punctuation, units, field order -- which converts many near-duplicates into exact ones. Then build a canonical key where the domain offers one, and for free text compare character or word n-gram overlap within cheap candidate groups so you avoid all-pairs comparison. Pick the similarity threshold by hand-checking a sample of borderline pairs, because it trades deleted real rows against retained copies.
  • Is removing duplicates always the right fix?
    No. Distinguish a duplicated record from a genuinely repeated event: if a hundred customers really bought the same product, those rows are real observations and collapsing them destroys the true frequency your model needs. Deduplicate artefacts of ingestion and joins; keep real repetition. When you are unsure, downweight the group so it totals one observation instead of deleting rows outright.
  • What goes wrong if you compute class weights before deduplicating?
    The weights are derived from inflated counts, so you correct an imbalance that does not exist -- possibly down-weighting a class that is not actually the majority. Every count-derived setting has this problem: resampling ratios, the assumed prior when calibrating probabilities, and the base rate you compare a lift against. Deduplicate first, then compute every dataset statistic from the cleaned table.

Photocopying one witness statement forty times does not give you forty witnesses; it gives you one very loud one, and a case file that looks much thicker than the evidence in it.

saying these in an interview costs you the question

  • Says duplicates are harmless because more data always helps
  • Checks only exact string equality and declares the data clean
  • Deletes every similar-looking row without inspecting a sample
  • Computes class weights or prevalence before deduplicating
  • Confuses a genuinely repeated event with a duplicated record
  • Quotes row count as dataset size without counting distinct entities

context