skip to content

A de-duplication keeps one row from each set of repeats; on what ordering is that survivor defined?

level: middleimportance: must knowfreq 62%

answer

  1. first needs an order to exist
  2. the arrangement came from the previous step
  3. defaults differ: earliest, latest, none, unspecified
  4. order deliberately, then reduce
  5. no precedence column, no implementable recency rule

basics

~20 s

On whatever order the rows happen to be sitting in when the step runs, which is a by-product of the previous operation rather than anything you stated. Put the rows in a deliberate order on a column that encodes precedence, or the survivor is not reproducible.

solid answer

~50 s

"Keep the first one" is meaningless until something has defined an order, and nothing in a table defines one by itself. The survivor is decided by the order the rows are in at that moment, which came from whatever step ran before — a read, a match, a grouping — and can change when any of those change. Tool behaviour varies more than people expect: some keep the earliest occurrence in the current order by default, some keep the latest, some refuse to default and make you choose, and some guarantee no ordering at all, so the survivor is not determined by anything you wrote. The portable habit is the same in every case: put the rows in an order yourself, on a column that genuinely encodes precedence, and only then reduce to one per repeat set. If no column encodes precedence, "keep the latest" is not implementable, and saying so is the finding.

go deeper

for a junior

Remember that keeping the first of a set of repeats only means something once an order exists, and that a table does not carry one by itself. Ask what defines first before answering.

for a middle

Explain where the incoming arrangement came from, that defaults for which occurrence survives differ between tools, and that the fix is to order on a precedence column deliberately and then reduce, rather than relying on any default.

for a senior

Show that you would recognise an unchanged output row count with moved totals as an inherited-ordering symptom, and that you would say plainly when no column in the record can decide recency instead of producing a plausible number.

for a principal

The judgment call is whether the pipeline should refuse to reduce when no precedence column exists, or reduce arbitrarily and record that it did — and which of those a downstream consumer can actually live with.

## "First" is not a property of the data A table is a set of rows with a current arrangement, and that arrangement is not part of what the data means. Nothing about account `4471` makes one of its two rows earlier than the other. So an instruction like "keep the first occurrence" is not yet a rule — it is a rule waiting for an order, and if you did not supply one, the operation used whatever arrangement it found. That arrangement has a history. It is the output of whichever step ran immediately before: the order a reader produced walking a file, the arrangement left behind by a match between two tables, the arrangement a grouping imposed on its results. Change the reader, add a filter, change which side of a match you started from, and the arrangement can change without any of your logic changing. The de-duplication then keeps a different row and every number downstream moves, with nothing in the code to point at. ## The behaviours differ across tools, and the difference is not cosmetic This is the part candidates most often state as universal. There is no single rule in this family: | Behaviour | What the survivor is | What it costs you | |---|---|---| | Keeps the earliest in the current arrangement, by default | Whichever row the previous step left first | Reproducible only if the previous step is | | Keeps the latest in the current arrangement, by default | Whichever row the previous step left last | Same exposure, opposite row — silently different results from the same code moved between tools | | Requires the choice explicitly | Whatever you asked for | Forces the question, which is the safest of the four | | Guarantees no ordering at all | Not determined by anything you wrote | The same input can yield different survivors on different runs | So a sentence like "de-duplication keeps the first row" is a statement about one tool, not about the operation. What **is** universal is the arithmetic: each set of repeats contributes exactly one row to the result, and the removed rows are gone without a record unless you arranged otherwise. ## The repair, in order 1. **Name the columns that define a repeat.** Two rows are the same row on those columns and only those. Everything else is a value that may differ. 2. **Name the column that decides precedence** — an event timestamp, an effective date, a version number, a source rank. This has to be a column whose values actually mean "later" or "more authoritative", not a column that merely correlates with arrival. 3. **Put the rows in an order on that column, deliberately, as its own step**, so the arrangement the reduction sees is one you wrote rather than one you inherited. 4. **Reduce to one row per repeat set**, and if your tool lets you say which end of the order to keep, say it rather than relying on a default. Written that way, the step is reproducible: the same input yields the same survivor regardless of which reader produced it, which filter ran before it, or which tool executes it. ## When no column can decide The interesting case, and the one interviewers press on, is that step 2 sometimes has no answer. The two rows disagree about the balance and there is no timestamp, no version, no source marker — nothing in the record says which claim is the later or better one. Then: - **"Keep the latest" is not implementable.** Any rule you write is picking arbitrarily and dressing it as recency. - **The honest options are to pick arbitrarily and say so in writing, to keep both and push the choice downstream, or to go back to the producer and ask for the column that would decide it.** Which of these is right depends on what the number is used for. - **A file's row order is not a precedence column.** Neither is the position a row happens to occupy, nor the order rows come back from a match — in several designs that last one is explicitly unspecified and can differ between operations in the same tool. Stating that the data cannot answer the question is a stronger answer than producing a number, because the number produced by an arbitrary rule looks exactly like a number produced by a correct one. ## What this looks like when it goes wrong in production The classic report is "the same job gave a different total than yesterday and nothing changed". Something did change — the arrangement of the rows entering the reduction — and because the reduction's rule was inherited rather than written, the survivor changed with it. The tell is that the row count of the output is identical and only the values moved: the same number of repeat sets, each resolved to a different member. A reduction whose ordering is stated cannot produce that symptom.

  • The same job produced a different total than yesterday, and the output row count is unchanged. What do you suspect first?
    That a reduction to one row per repeat set resolved its sets differently. An unchanged output row count means the same number of repeat sets survived, so the sets are stable and only their chosen members moved — the signature of a survivor decided by an inherited arrangement rather than a stated order. Look for what changed upstream of it: the reader, a filter, or the side a match started from.
  • Why is a column that correlates with arrival, such as a load batch number, a weak precedence column?
    Because it records when a row reached you, not when the fact it states became true. A correction loaded in the same batch as the original, or an older record backfilled later, both invert the relationship. It is usable only when you can say the producer never backfills and never reorders, which is a claim about their system rather than about your data.
  • If the tool guarantees no ordering for the survivor, is it wrong to use it at all?
    No, but only where the repeats are identical in every column, since then the survivor is indistinguishable whichever it is. The moment the rows disagree anywhere, an unordered reduction is picking arbitrarily, and the result is not reproducible even on the same input.

saying these in an interview costs you the question

  • States that de-duplication keeps the first occurrence, as though every tool did.
  • Assumes the incoming row arrangement is stable across runs without having stated one.
  • Treats the position a row occupies as evidence of when it arrived.
  • Says keep the latest without naming a column that says which is later.
  • Believes the order rows come back from a match is dependable enough to reduce on.
  • Expects some diagnostic when the survivor was decided arbitrarily.