skip to content

Before dropping repeated rows from a delivery, why would you mark or count them first?

level: middleimportance: should knowfreq 44%

answer

  1. clean table, no evidence
  2. measure before you change anything
  3. set-size distribution is the diagnosis
  4. marking keeps the losing rows visible
  5. drop last, once you can say why

basics

~20 s

Dropping is destructive and unmeasured: it produces a clean table and no evidence. Counting tells you whether this is a handful of stray copies or a uniformly doubled delivery, and marking keeps every row so a later step, or a person, can still decide.

solid answer

~50 s

Dropping repeats answers the question "how do I make this table clean" and deletes the evidence needed for "why was it dirty". Counting comes first because the shape of the repetition is the diagnosis: if every repeat set has exactly two members, the delivery mechanism duplicated the whole file; if most sets have one member and a few have hundreds, your idea of what identifies a row is wrong. Those two findings have different owners and different fixes, and the count distinguishes them in one pass. Marking — keeping every row and adding a flag saying which are repeats and which was chosen — costs a column and preserves the ability to reconstruct what was removed, which matters when the choice of survivor is contestable. Drop last, once you can say what you are removing and why.

code

pseudocode · 21 lines
pseudocode
repeat_columns = [account_id, effective_date]
precedence     = descending by source_rank, then by revision_number

sets = group rows by repeat_columns

# 1. COUNT - change nothing, learn the shape
surplus_rows  = row_count(rows) - count(sets)
size_histogram = count of sets grouped by size(set)
#   {2: 48000}                -> the whole delivery was doubled
#   {2: 31, 3: 4}             -> scattered corrections
#   {1: 90000, 412: 1}        -> repeat_columns do not identify a record

# 2. MARK - keep every row, add the evidence
for each set:
    ordered = order rows of set by precedence
    for each row at position p in ordered:
        row.in_repeat_set = size(set) > 1
        row.is_chosen     = (p == 1)

# 3. DROP - only now, and only by filtering the mark
reduced = rows where is_chosen

go deeper

for a junior

Know that removing repeats is not the only option: you can count them without changing anything, or flag them and keep every row. Reach for counting first so you know what you are about to remove.

for a middle

Explain what the distribution of repeat-set sizes tells you — a uniform factor of two means the delivery was doubled, a few large sets mean the columns you chose do not identify a record — and why a single rows-removed total cannot distinguish those.

for a senior

Show that you weigh preserving the losing rows against the cost of carrying marker columns, and that you treat quietly cleaning a producer's defect inside your transform as a decision with an owner, not as tidiness.

for a principal

The standing question is whether repeats should be removed by transforms at all or should fail the delivery, and what a team gives up in incident visibility every time a pipeline absorbs an upstream problem without reporting it.

## Three dispositions, not one operation Finding repeats and removing them get treated as one act. They are three separate dispositions, and the order in which you reach for them is most of the skill. - **Count them.** Do not change the table. Produce the number of repeat sets, the number of surplus rows, and the distribution of set sizes. - **Mark them.** Keep every row and add columns that say which rows belong to a repeat set and which member the stated rule would choose. - **Drop them.** Reduce to one row per set and discard the rest. Each answers a different question, and only the third changes the table. ## What counting tells you that dropping hides The headline number — "we removed 4,000 rows" — is nearly useless on its own. The distribution of set sizes is the diagnosis: - **Every set has exactly two members, and almost every row belongs to one.** The delivery itself was duplicated: a file processed twice, a producer that re-emitted after a restart. The fix is upstream and your table was never really wrong, just doubled. - **A small number of sets, each with two or three members, scattered through the data.** Ordinary late corrections or overlapping source systems. This is a data question, not a plumbing one, and it needs the precedence rule. - **Most sets have one member and a handful have hundreds.** The columns you nominated as defining a repeat do not identify a record at all. A table of daily balances compared on the account alone produces exactly this shape, and dropping would destroy the history. - **The number of surplus rows equals the row count of a known input piece.** Something was combined twice. All four look identical after a drop, because a drop reports one number and destroys the shape that distinguished them. ## What marking buys Marking is the disposition people skip, and it is the one that survives contact with a disagreement. Keep every row; add a flag for "this row belongs to a set of more than one" and, if you have stated a precedence rule, a second flag for "this is the member the rule chose". Then: - A consumer that wants the reduced table filters on the chosen flag and gets exactly what a drop would have given. - A consumer investigating a wrong figure can see the rows that lost, which a dropped table cannot show at any price. - Changing the precedence rule is re-running one cheap step rather than re-reading the source. - The repeats remain visible to whoever is supposed to fix the producer, instead of being quietly absorbed by your transform. The cost is one or two columns and the discipline of making consumers filter. That is a small price where the survivor is contestable, and pointless where the repeats are identical in every column. ## The three compared | | Count | Mark | Drop | |---|---|---|---| | Changes the table | No | Adds columns only | Yes, destructively | | Answers | How bad is it, and of which kind | Which rows lost, and to what rule | What does the clean table look like | | Survivor rule needed | No | Yes, to flag the chosen member | Yes, stated or inherited | | Recoverable afterwards | Nothing to recover | Fully | Only by re-reading the source | | Right when | Always, first | The choice is contestable or disputed downstream | The repeats are exact copies, or the rule is settled | ## When dropping first is the right answer This is not an argument for never dropping. Where the repeats are identical in every column, there is no losing row to preserve and no rule to contest, so counting once to confirm that is the case and then dropping is proportionate. Where the reduction feeds a large downstream job, carrying marker columns through it has a real cost. The rule of thumb is that the more the survivor is a judgment, the more marking earns its keep, and counting is cheap enough to do unconditionally. ## The habit to demonstrate An interviewer asking this is checking whether you treat a clean table as the goal or as a by-product. The strong answer is that you never learn anything from a table that is already clean, so you measure before you clean, you keep the evidence where the choice could be argued with, and the number you report is not "rows removed" but the shape of the repetition and what it implies about where the defect lives.

  • Every repeat set has exactly two members and nearly every row is in one. What is your conclusion?
    That the delivery was processed twice rather than that the data is wrong. A uniform factor across essentially the whole table is a property of the transport, not of individual records, and the fix belongs to whoever loads or emits the file. Confirm it by checking that the surplus row count equals the row count of one input piece before you touch the data.
  • What does it cost a downstream consumer if you drop repeats inside your transform without saying so?
    They see a table that has always looked clean, so the producer's defect never surfaces and nobody funds fixing it. They also cannot reconcile your row count against the source, which turns every future discrepancy into an investigation of your transform. Reporting the count alongside the clean table costs nothing and removes both problems.
  • When is marking not worth the two extra columns?
    When the repeats are identical in every column, since there is no losing row worth preserving and no rule anyone could contest. Also when the marked table feeds a job large enough that carrying two more columns through it has a measurable cost, and the precedence rule has already been agreed and is not under review.

saying these in an interview costs you the question

  • Goes straight to removal and reports only the number of rows removed.
  • Treats marking and dropping as the same answer with different syntax.
  • Assumes a clean removal report proves the input had few repeats.
  • Reads a single surplus-row total as if it identified the cause.
  • Cleans the table inside the transform and never tells the producer.