A record turns up after its group's answer was already published — what three things can a pipeline do with it?
answer
- three options, not three settings
- drop, divert, restate
- each is a promise downstream
- final figures versus accurate figures
- restating requires an overwritable destination
basics
~20 sDrop it, divert it to a separate repair channel, or reopen the group and publish a corrected answer. These are three different promises to everyone downstream about whether a published number can still change, not three settings.
solid answer
~40 sThe job carries a **completeness claim** — a *watermark*: a timestamp asserting nothing older is still expected — and once it passes a group's end that group is eligible to close. A record arriving after that leaves three options. **Drop it**: published numbers never change, and some records are gone; only an explicit late-record count says how many. **Divert it**: write it to a separate repair channel, a second named output, so nothing vanishes and someone can reconcile later, while published numbers still never change. **Restate**: reopen the group and publish a corrected value, which only works if every downstream destination can overwrite a previously published answer rather than only append to it. The choice is a promise about the mutability of published figures — pick it with consumers, then implement it.
code
json · 15 lines{
"repairChannel": "orders-missed-their-hour",
"record": {
"orderId": "A-8123",
"amountMinor": 4990,
"occurrenceTime": "2026-09-19T23:58:11Z"
},
"group": {
"start": "2026-09-19T23:00:00Z",
"end": "2026-09-20T00:00:00Z",
"publishedAt": "2026-09-20T00:03:12Z"
},
"claimWhenItArrived": "2026-09-20T00:06:40Z",
"lateBy": "PT8M29S"
}go deeper
Remember the three names and that they are alternatives: throw the record away, write it somewhere separate, or publish the group's answer again with a new value.
Explain what each option demands. Dropping needs a count to be honest, diverting needs a second output and a reader, restating needs destinations that can overwrite what they already published.
Argue the choice from the consumer side rather than the pipeline side, and say which of the three the existing destinations could actually support today without changing them.
Frame it as a commitment the organisation makes about the mutability of published figures, and recognise that different figures from the same pipeline may deserve different promises.
## Where the choice arises A job grouping by the moment stamped in the payload carries a **completeness claim** alongside the records — a *watermark*: a timestamp asserting that nothing older than it is still expected to arrive. When that claim passes the end of a group, the group becomes eligible to close and its answer can be published. A record stamped inside that group which arrives afterwards is **late**: valid, wanted, and past the point at which the job committed to an answer. At that instant the pipeline is not choosing a behaviour for itself. It is choosing what it has already promised to everyone reading the published number. ## Promise one: drop it The record is discarded. The published figure for that group is final the moment it is published and will never move. - **What the consumer gets**: numbers that never change, which makes every downstream copy, cache and screenshot permanently correct. - **What it costs**: the record itself, and a shortfall nobody can size unless a **late-record counter** — an explicit count of records dropped for arriving too late — is published next to the figure. - **When it fits**: high-volume, low-value-per-record signals where a small fraction changes no decision, and anywhere a consumer genuinely cannot absorb a change. Dropping is a legitimate, declarable choice. What is not legitimate is dropping *by omission*, because that is the default on most runtimes and it gets chosen by nobody. ## Promise two: divert it to a repair channel The record is written to **a separate repair channel** — a second named output holding the records that missed their group, each ideally carrying its own moment, the group it belonged to and the claim at the instant it arrived. - **What the consumer gets**: the same unchanging published figures as dropping. - **What it costs**: a second destination, and a human or a scheduled process that actually reads it. An unread repair channel is a drop with extra storage. - **What it buys over a counter**: the records themselves. A counter tells you a hundred were lost; the channel tells you they were a hundred payments from one region, which is the difference between a statistic and an incident. ## Promise three: reopen and restate The group's answer is republished with a new value — **a corrected restatement**. Typically a **grace period after the group closes** governs how long this remains possible: an interval during which a straggler still updates the answer, after which the record falls back to one of the first two promises. - **What the consumer gets**: the most accurate number available, eventually. - **What it costs**: every downstream destination must be able to overwrite a previously published answer rather than only append to it, and so must every figure derived from it. - **What it demands of readers**: the same query run twice may legitimately return different numbers. ## The three side by side | Promise | What consumers are told | What it costs you | Fits when | |---|---|---|---| | Drop | A published number is final | Records, and accuracy you cannot size without a counter | Per-record value is low; consumers cannot absorb change | | Divert | A published number is final, but nothing is thrown away | A second output and someone who reads it | The records matter individually even if the figure does not move | | Restate | Any published number may still change | Every destination must support overwriting | Accuracy outranks stability, and sinks can express a change | ## What varies between engines Do not assume the mode you know is the class: - A grouping that **emits once and releases the group** has nothing to reopen unless the group's retained state is deliberately kept alive for a grace period. - A grouping that **republishes an updated running result on every input** is effectively restating by default — the interesting question there is not whether to restate but whether the destination noticed. - A **single pass over a finished bounded input** has no running claim during the run at all, so the whole choice moves outside the job, to whoever decides what the next run reads. - Whether a grouping can emit a **second output** for missed records varies; where it cannot, split the stream yourself before the grouping using the same comparison the claim makes. ## Choosing Three questions settle it, and none of them is technical: can the people reading this number tolerate it changing after they have read it; is an individual missed record worth anything on its own; and is there anyone who will actually look at a repair channel. Answer those with the consumers, write the answer down, and the implementation follows.
- Which of the three happens if nobody makes a decision?Dropping, on most runtimes, and usually with no count. The default is the one promise that is never discussed with consumers, which is why it so often turns up as a surprise months later when someone reconciles against a figure computed another way.
- Can a pipeline use more than one of the three at once?Yes, and that is common. Restate within a grace period after the group closes, then divert or drop beyond it. The record's age decides which promise applies, so the contract has to state both the grace period's effect and what happens past it.
- What does reopening a group actually require of the job?It varies by mode. Where the grouping keeps the group's retained state alive it can fold the record in and republish; where the group was released at close there is nothing to reopen. And in a single pass over a finished bounded input there is no running claim to have been passed in the first place.
- Why is a repair channel sometimes better than simply restating?Because restating obliges every destination and every derived figure to handle a value changing. A repair channel keeps published numbers stable while preserving the records, which lets the loss be quantified in business units and reconciled deliberately rather than propagated automatically.
A bank statement printed on the last day of the month, with a transaction that clears the next morning. The bank can ignore it, carry it onto next month's statement as a separate line, or reissue a corrected statement for the month just closed. Three different promises to whoever files the statement — and only the third one obliges them to throw the old copy away.
saying these in an interview costs you the question
- Says accepting the late record is always the right answer.
- Treats diverting and restating as the same thing under different names.
- Presents the choice as a configuration value rather than a promise to consumers.
- Believes dropping is negligence rather than a legitimate, declared choice.
- Assumes reopening a closed group is free because the data is still around.
- Counts a repair channel nobody reads as materially better than dropping.