skip to content

What makes a write key safe for a re-run to repeat, and which common key choices quietly break that?

level: middleimportance: must knowfreq 64%

answer

  1. pure function of the record
  2. nothing from the clock or the attempt
  3. replace, never accumulate
  4. the destination must enforce it
  5. same record, same key, every attempt

basics

~20 s

A safe write key is computed only from fields the input record already carries, so a re-run derives the same key and the destination replaces the earlier row. Keys drawn from randomness, wall-clock time, attempt identity or a destination-assigned sequence all differ on the second run.

solid answer

~50 s

The property you need is that the key is a pure function of the input record: same record in, same key out, on every attempt, on any machine, in any order. Then a repeat of the write lands on the row the first attempt created and replaces it, so the destination ends in the state it would have reached had the work run once. Keys break this the moment anything outside the record leaks in — a randomly generated identifier, the time the write started, the attempt or run identity, the number of the piece the record happened to land in, or a sequence the destination assigns on arrival. Two further conditions are easy to miss. The *operation* must replace rather than accumulate: adding to a running total repeats faithfully even under a perfect key. And the destination must actually enforce the key, with a uniqueness constraint or a match-and-replace statement; a key the destination ignores leaves the second write a plain insert.

code

sql · 14 lines
sql
-- Safe to repeat: the key is the aggregate's own grouping fields,
-- and the measure is written whole rather than added to.
MERGE INTO daily_region_totals t
USING (SELECT :day AS day, :region AS region, :total AS total) s
   ON t.day = s.day AND t.region = s.region
WHEN MATCHED THEN
  UPDATE SET total = s.total
WHEN NOT MATCHED THEN
  INSERT (day, region, total) VALUES (s.day, s.region, s.total);

-- Not safe to repeat, despite the identical key:
-- UPDATE daily_region_totals
--    SET total = total + :total
--  WHERE day = :day AND region = :region;

go deeper

for a junior

Recall the one-line test: could a second run compute this same key from the same input record, without looking at the clock, the machine or a counter? If not, the repeat inserts a new row.

for a middle

Explain the full set of conditions — a key that is a pure function of the record, an operation that replaces rather than accumulates, and a destination that actually enforces uniqueness on that key.

for a senior

Bring the awkward cases: records with no natural identifier, values that are themselves non-deterministic so the replacing row differs from the original, and destinations that cannot enforce a key at all.

for a principal

Treat the key as a contract with every downstream reader. Choosing it fixes what 'the same record' means for the whole platform, and changing it later invalidates every deduplication and every reconciliation built on top of it.

## The property, stated precisely A **deterministically chosen write key** is a key computed only from the input record, so that a re-run computes the same key and the destination overwrites the earlier row instead of adding a second one. The test has three clauses, and all three matter: 1. **Same input, same key** — the function reads nothing but the record's own fields. 2. **On any machine, in any order** — it does not depend on which worker ran the unit, which piece the record landed in, or what was processed before it. 3. **On every attempt** — the second attempt after a crash derives the identical value. When those hold, the repeat is not prevented; it is made not to matter. The second write arrives, finds a row with that key, and replaces it. That is the whole trick, and it is why this is usually the cheapest of the three ways to survive a replay. ## Where keys come from, and where they go wrong | Key source | Safe to repeat? | Why | |---|---|---| | A natural identifier carried by the record | Yes | Present in the input; identical on every attempt | | A hash of the record's identifying fields | Yes | A pure function of the same input | | The grouping fields of an aggregate, such as day and region | Yes | The recomputed group writes the same key | | A randomly generated identifier made at write time | No | A fresh value on every attempt, so the repeat inserts | | Wall-clock time when the unit started | No | Differs between the crashed run and the restarted one | | Attempt or run identity, or the piece number | No | Deliberately different on the retry, which is the point of them | | A sequence the destination assigns on arrival | No | The destination has no idea the second row is the same work | The last row is the one candidates miss most often: a destination-assigned identifier looks like a key and guarantees nothing, because it is minted *after* the decision that should have been made by the key. ## The two conditions hiding behind the key **The operation must replace, not accumulate.** A statement of the shape `total = total + amount` is faithful arithmetic and a perfect duplicate generator: a stable key tells the destination which row to touch, and the repeat then adds the amount a second time. The safe form recomputes the value and writes it whole — `total = :recomputed_total` — which is why aggregates keyed on their grouping fields survive replay so naturally, and why running counters maintained by increment do not. **The destination must act on the key.** Computing a key and sending it as an ordinary column changes nothing. The destination needs either a uniqueness constraint on that key plus a statement that replaces on conflict, or an explicit match-and-replace statement. Where the destination cannot enforce uniqueness at all — an append-only store, an object destination, a plain file — a key is still worth writing, but it has become the input to deduplication at the reader rather than a guarantee at the write. ## When the key is stable and the value is not One subtlety is worth holding. If the value being written is itself derived non-deterministically — a sampled subset, a value stamped with the current time, an aggregate whose result depends on arrival order — then the re-run writes a *different* row under the same key. Replacing is still the right behaviour and the destination still ends with one row, but the row's contents now depend on which attempt won, which may or may not be acceptable for the numbers downstream. If instead the key itself is a hash of a non-deterministically derived field, the property collapses outright: the second attempt computes a different key and inserts. ## Where this sits among the alternatives A deterministically chosen write key is the first thing to reach for because it needs nothing from the runtime — no coordination, no shared transaction, no destination integration. It asks only that the destination can replace by key. When the destination cannot, the other two moves are committing the output and the recorded read position (the marker saying how far into the input the job has durably got) as one unit, or keeping a deterministic identifier on every row and collapsing duplicates when the data is read. Runtimes differ sharply in how much help they give here: some ship destination integrations that arrange the coupling for you, others leave the whole question to the author, and no runtime can supply a key the record does not contain.

  • The records carry no natural identifier. What do you key on?
    A hash over the fields that identify the record — typically the source identifier plus the event moment plus the fields that distinguish two otherwise-identical records. It must be computed from the record alone, never from arrival time or the piece it landed in. If genuinely identical records can legitimately occur, the hash collapses them, so you need a field that distinguishes them, or you accept deduplication at the reader instead.
  • Does a stable key help when the destination is append-only?
    Not at the write, but it is still the thing worth writing. An append-only destination stores both copies; a deterministic identifier on each row lets every reader keep one copy per identifier. You have moved the cost from the writing side to every reader, permanently, which is a real trade and not a free one.
  • Why is a key generated by the job at write time not enough, even if the job stores it?
    Because the restarted attempt generates a new one. Storing it durably before the write turns it into a write-ahead record of intent, which is a different mechanism with its own gap. A key that must survive a crash to be useful is not deterministic; a key recomputed from the record needs to survive nothing.

saying these in an interview costs you the question

  • Generates a fresh random identifier for each row at write time
  • Uses the current time or the run identity inside the key
  • Thinks a key alone fixes a statement that adds to a running total
  • Keys on the piece number or the worker that handled the record
  • Sends the key as an ordinary column the destination never enforces
  • Relies on an identifier the destination assigns on arrival