skip to content

Your team wants records arriving after a published daily total to correct it rather than be dropped — what must every downstream destination be able to do first?

level: seniorimportance: should knowfreq 54%

answer

  1. can the destination change its mind
  2. overwrite by group key
  3. append-only means double counting
  4. some effects cannot be recalled
  5. derived figures must restate too

basics

~20 s

Each destination must be able to replace a value it already published for that group, keyed by the group's identity. An append-only destination turns a correction into a second value beside the old one, which reads as a double count.

solid answer

~40 s

A corrected restatement republishes a group's result with a new value, so the requirement is that every place the old value landed can express *this group's answer changed* — an overwrite keyed by the group, or an append-only store whose readers collapse to the newest value per key. A destination that only appends, and whose readers sum everything, silently double counts. The shape of the correction varies by runtime: some emit the group's whole new value, some a withdrawal of the previous value followed by its replacement, some only a delta — the destination has to understand whichever yours produces. Two things cannot be restated at all: effects already released into the world, and derived figures whose owners do not restate in turn, where the correction stops one hop downstream.

go deeper

for a junior

Grasp the basic requirement: publishing a corrected total only helps if whatever received the first total can replace it rather than keep both.

for a middle

Explain why an append-only destination inflates rather than fixes, and that the correction may arrive as a full value, as a withdrawal plus replacement, or as a delta.

for a senior

Inventory the destinations before promising restatement: which can overwrite by group key, which only append, which are effects already released, and which derived figures inherit the obligation.

for a principal

Recognise that restatement is a commitment binding teams you do not own, and that the honest alternative is a stable published figure plus a preserved record of what was missed.

## What a restatement actually emits A **corrected restatement** is the job publishing a group's result again with a new value, after a record arrived that the **completeness claim** — the timestamp the job carries asserting nothing older is still expected — had already passed. It sounds like one mechanism and it is at least three, and which one you get depends on the runtime: - **The whole new value for the group.** The consumer sees `(group key, 4,213)` after having seen `(group key, 4,197)` and is expected to take the newer one. - **A withdrawal of the previous value followed by its replacement.** The consumer sees the old value marked as no longer holding, then the new one — a stream of changes rather than a stream of facts. - **A delta only.** The consumer sees `+16` and is expected to add it to what it already holds. Every one of these requires the destination to understand the convention. Handing a delta to a consumer that treats each message as a complete answer, or a complete answer to one that adds, are two different ways to be wrong by roughly the size of the correction. ## The requirement on each destination The test is not *can this destination accept another write*. It is: **can the value a reader sees for this group become the new value, without the old one continuing to count?** That is possible in three shapes and impossible in a fourth. | Destination | Absorbs a restatement? | What happens if you try | |---|---|---| | Keyed store, written by group key | Yes | The new value replaces the old; readers see one answer | | Append-only files or log, readers collapse to newest per key | Yes | Correct, provided every reader collapses — including the ad-hoc one | | Append-only table, readers sum all rows | No | Old and new both count; the total is inflated by the old value | | An effect already released — a message sent, a payment made | No | Nothing to overwrite; the correction can only be compensated | The third row is the trap, because it does not fail. The write succeeds, the pipeline reports health, and a total that was slightly short becomes substantially long. ## Effects you cannot take back Some of what a pipeline publishes is not a value at all. A notification already delivered, an invoice already issued, a file another team already copied, a figure already read aloud in a meeting — these have left the system's control. For anything in this class, restating upstream does not undo the downstream act; it can only be followed by a separate, deliberate compensating one. Knowing which of your destinations are values and which are events is the first inventory to make, and it is usually shorter and more alarming than people expect. ## The correction has to propagate A restatement that only the first destination understands stops one hop downstream: 1. The job republishes the hour's total. 2. The keyed store the job writes replaces it correctly. 3. A daily figure built by summing those hours was computed an hour ago and is not recomputed — so the day is still wrong. 4. A weekly figure built from the daily one inherits the error, now two hops from anything that would notice. So the requirement is recursive: **every figure derived from a restatable figure must itself be restatable, or be recomputed after corrections settle.** In practice this is what makes restatement an organisational commitment rather than a pipeline option — it obliges teams you do not own to re-read. ## What varies between engines - Whether the runtime can restate at all depends on whether the group's retained state is still alive; a grouping that released it at close has nothing to fold the record into. - A grouping that **republishes an updated running result on every input** restates continuously, which means its destinations were already obliged to overwrite and nobody may have noticed. - A **single pass over a finished bounded input** expresses the same idea as replacing a previous output wholesale rather than as a per-group correction. - Where the runtime emits withdrawals, a destination that ignores them will hold both values; where it emits full values, a destination that appends will hold both as well. The failure mode is the same from two opposite mechanisms. ## When the answer is that you cannot If the destinations cannot overwrite and cannot be changed, restating is not available and saying so is the senior answer. The fallback is one of the other two promises: publish the figure as final and divert the late records to a separate repair channel, so the discrepancy is preserved, quantified and reconciled deliberately, rather than pushed into a destination that will quietly add it to the old value.

  • A destination only appends, but every reader takes the newest row per key. Is restatement safe?
    Yes for those readers, and that is a real and common design. The risk is the reader you did not enumerate — an ad-hoc query, an export, a copy another team made — that sums everything. The collapsing rule has to be a property of the data as published, not a convention each reader remembers.
  • How do you handle a destination that cannot overwrite and cannot be changed?
    Do not restate into it. Publish its figure as final and route late records to a separate repair channel, so the gap is preserved and quantifiable. A different destination on the same pipeline can still be restated; the promise is per-figure, not necessarily per-job.
  • Why does a correction often stop one hop downstream?
    Because the derived figure was computed from the old value and nothing recomputes it. Hourly totals get corrected, the day built from them does not, and the week built from the day inherits the error. Restatability has to hold for every figure in the chain, or corrections have to be followed by a recomputation of what depends on them.

saying these in an interview costs you the question

  • Assumes writing the corrected value again is enough, whatever the destination.
  • Thinks a successful write proves the correction was absorbed.
  • Ignores figures derived downstream, so the correction stops one hop away.
  • Treats an already-sent notification as something a restatement can undo.
  • Believes every runtime expresses a correction the same way.