skip to content

Two readings share one timestamp in a table labelled by time — what changes when you ask for that label?

level: middleimportance: should knowfreq 44%

answer

  1. two sensors, one instant
  2. the result's shape follows the data
  3. a lookup returns a sub-table
  4. range selection keeps working fine
  5. row count against distinct label count

basics

~20 s

Asking for a repeated label returns every row carrying it, so the result is a sub-table rather than a single row. The shape now depends on the data rather than on the code, and in several designs nothing raises.

solid answer

~50 s

A repeated stamp means the label no longer identifies a record. The same expression that returned one row yesterday returns three today, so the *shape* of the result is a property of the data rather than of the code. That is the part that hurts: downstream code written against a single record either fails somewhere far from the cause, or quietly folds several readings into one. Designs vary — some reject repeats when the column is promoted, some allow them freely, some allow them and refuse specific operations. What does not vary is that repeated stamps are normal in time data: two sensors report at the same instant, a batch is written with one stamp, or the clock is coarser than the arrival rate. So the fix is not to delete duplicates but to decide what identifies a reading before you promote anything.

go deeper

for a junior

Know that a timestamp is not automatically unique, and that asking for one that repeats gives you every matching row rather than an error.

for a middle

Explain that the result's shape now depends on the data, not the code, and name why repeats are normal in time data rather than a fault to delete.

for a senior

Trace the damage past the lookup: the case where nothing fails and two readings become one average, and the check that would have caught it where the stamp was promoted.

for a principal

Make what identifies a reading an explicit decision recorded before the pipeline is written, because every consumer downstream inherits it and none of them can recover it from the data alone.

## Repeated stamps are the normal case, not the accident In most tables a duplicated key is a defect. In time data it is the expected condition, and for several independent reasons: - **Several sources report at once.** Ten sensors sampling on the same schedule produce ten readings at each moment, and that is the data working correctly. - **The clock is coarser than the arrivals.** A stamp recorded to the second cannot separate events arriving four times a second; the repeats are a property of the recording, not of the world. - **A batch shares one stamp.** A writer that stamps a whole load with the moment the load ran gives every row in it the same value. - **Two readings genuinely coincide.** Nothing forbids it, and for a busy feed it happens constantly. So a design that refused to let you promote a repeating stamp would be refusing a very ordinary table. Several designs do not refuse. ## What the tool does instead of complaining | What you do | With unique labels | With repeats | |---|---|---| | Ask for one label | One row | Every matching row, as a sub-table | | Ask for a range | A block of rows | A block of rows — unchanged | | Line two objects up on shared labels | One-to-one | No longer one-to-one; the row count can grow | | Promote the column | Accepted | Accepted in some designs, rejected in others | The second row of that table is why the problem survives so long. **Range selection carries on working perfectly**, so the pipeline runs green for months; the single-label lookup that sits in one branch of one function is the thing that changes shape, and it changes shape only on the days the data happens to contain a repeat. ## What it does downstream The damaging part is not the lookup itself, it is that **the shape of the result stopped being a property of the code**. A function that took a reading and returned a number now sometimes receives a small table. Three outcomes are common, in rising order of unpleasantness: 1. **A loud failure near the cause** — the next operation rejects the shape and you find it immediately. This is the good case. 2. **A loud failure far from the cause**, three steps later, where the shape has already been reshaped once and nothing points back at the duplicate stamp. 3. **No failure at all.** The next step aggregates, so two readings quietly become their mean or their sum, and the number is wrong in a direction nobody can see. The row counts stay plausible throughout. ## Deciding what a repeat means Before promoting a stamp, answer one question: **what identifies a reading in this data?** 1. **If it is the moment together with the source**, then the moment alone was never an identifier, and code that looks a moment up must expect several rows. Say so explicitly rather than discovering it. 2. **If the repeats are the same measurement recorded twice**, they are genuinely redundant and the pipeline should collapse them at a point you chose, with a rule you wrote down. 3. **If the repeats come from a clock too coarse to separate real events**, the stamp is a bucket rather than a moment, and treating it as a moment will keep producing surprises. 4. **If you cannot tell**, that is the finding. Do not promote until you can. ## Finding out before it finds you The check is cheap: compare the number of rows with the number of distinct labels. Equal means unique; a gap is the count of extra rows, and where the gap is small relative to the table it is the case most likely to slip through testing, because a sample of the data is unlikely to contain one. Do the check at the point the stamp is promoted, and record which answer you got as an expectation the pipeline asserts. **A repeat that appears for the first time in production is not a new defect — it is the day the data finally exercised the assumption the code has always made.**

  • Why does a pipeline with repeated stamps often run for months before anyone notices?
    Because the operations used most keep working. Range selection over a block of rows is unaffected by repeats, and so is most whole-column arithmetic. Only a single-label lookup changes shape, and only when the data happens to contain a repeat, so the failure arrives on an ordinary Tuesday with no code change behind it.
  • What is the cheapest check for repeats?
    Compare the row count with the number of distinct labels. Equal means the label identifies a row; the difference is exactly how many extra rows there are. Run it where the stamp is promoted and keep the answer as an assertion, so a first repeat is reported rather than absorbed.

saying these in an interview costs you the question

  • Assumes promoting a repeating stamp raises an error
  • Treats every duplicate timestamp as dirty data to be deleted
  • Expects a label lookup to return one row whatever the data holds
  • Thinks a finer clock resolution removes the possibility of ties
  • Believes repeats only matter when you look a single label up