A team wants to raise a streaming job's out-of-order wait from thirty seconds to five minutes to catch more records — what does that cost and buy?
answer
- the hold applies to every group
- benefit read off the distribution's tail
- each nine costs an order of magnitude
- retained state grows with the wait
- price it against the consumer's decision
basics
~20 sIt adds four and a half minutes of latency to every result the job publishes, not just to affected groups, and holds each group's retained state that much longer. It buys only the share of records the measured distribution places between the two durations.
solid answer
~50 sA wait is a uniform tax. The job holds its completeness claim — a watermark, a timestamp carried alongside the records asserting that nothing older is still expected — a chosen duration behind the newest stamped moment it has seen, and *every* time group is held for that duration before it can be treated as finished, including the groups that were complete in two seconds. So the first cost is four and a half extra minutes on all output, plus proportionally more retained state, because open groups keep their accumulators. The benefit is read straight off the source's measured disorder distribution: the extra records captured are exactly the share falling between thirty seconds and five minutes. On a typical heavy-tailed source that is a fraction of a percent. Ten times the wait for under one percent more completeness is the trade to put on the table.
go deeper
Understand that waiting longer for delayed records is not free: results appear later, and the delay applies to all of them rather than only the ones that were incomplete.
Explain the exchange precisely — a fixed hold behind the newest stamped moment buys the share of the measured distribution below that duration, and the curve flattens quickly.
Demonstrate the whole cost picture in a real job: latency on every result, state held per open group, slower alerting, heavier recovery, and a decision anchored in what the consumer does with the figure.
Own the policy rather than the number: which classes of output may trade latency for completeness, who is permitted to change the dial, and how the choice is re-justified when the source's distribution moves.
## The wait is paid by everything, not by the affected groups A job that waits for out-of-order records is not waiting selectively. It holds its completeness claim — a watermark, a timestamp carried alongside the records asserting that nothing older than it is still expected — a fixed duration behind the newest stamped moment it has seen. Every time group is therefore held for the full duration before it can be treated as finished, including the overwhelming majority that were complete within a second or two. This is the asymmetry that makes the decision hard. The *reason* to lengthen the wait is always a minority phenomenon — the few records arriving well behind the front — while the *cost* lands on all output, all the time. Raising the hold from thirty seconds to five minutes makes every figure the job publishes four and a half minutes older, forever, in exchange for a handful of records per million. ## What it buys, read off the distribution Once the source's disorder is measured as a distribution, the purchase stops being a matter of taste: the share of records captured at a wait of *d* is exactly the share of the distribution at or below *d*. On one measured source the curve might run: | wait | share of records captured | latency added to every result | |---|---|---| | 5 s | 91% | 5 s | | 30 s | 98.5% | 30 s | | 5 min | 99.4% | 5 min | | 1 h | 99.8% | 1 h | | 9 days | ~100% | unusable | The numbers above are one example source, not a universal shape, but the *form* is near-universal: a steep body followed by a long flat shoulder. Each additional nine of completeness costs roughly an order of magnitude of waiting. Going from thirty seconds to five minutes is a tenfold latency increase for under one percentage point of completeness — which may still be the right call for a billing aggregate and is obviously the wrong one for an alerting path. ## Latency is the first cost, not the only one - **Retained state.** Groups that cannot yet be treated as finished keep their accumulators. Memory or local disk held for open groups grows roughly with wait multiplied by arrival rate multiplied by the number of distinct grouping keys in flight. - **Recovery weight.** More open groups means a larger recovery picture to write and to restore, which lengthens restart time. - **Slower detection.** Anything alerted from the aggregate now fires four and a half minutes later, which for an availability or fraud signal can dwarf the value of the extra records. - **Transitive delay.** Every consumer built on this output inherits the delay, and a chain of three such jobs inherits it three times. - **Reduced headroom.** A longer hold raises steady-state resource use, so the job has less slack for a traffic spike. ## How to choose the number The wait is chosen from the distribution *against the decision the output drives*, in roughly this order: 1. **Ask what the consumer does with the figure.** A page that refreshes hourly cannot perceive the difference between a thirty-second and a five-minute hold, so the latency is nearly free there. An on-call alert or a real-time control loop pays the full price. 2. **Price the marginal record.** If the missing fraction of a percent changes no decision anyone makes, it is not worth a tenfold latency increase; if it is revenue that must reconcile, the calculus inverts. 3. **Take the cheapest wait consistent with that.** Sit at the shoulder of the body of the distribution, not out in the tail, and be explicit that you are buying completeness with latency rather than obtaining it. 4. **Say what happens to the remainder.** Choosing a wait always leaves a residue; the choice is only defensible alongside a stated treatment for what arrives afterwards, which is a separate mechanism from the wait itself. ## What varies between engines The latency you actually pay is not always the number you set, so name the mode. On a runtime that executes continuous work as a rapid succession of small finite jobs, the effective hold is your chosen duration rounded up by at least one of those intervals, so a five-minute wait on a one-minute cadence really means up to six. On a record-at-a-time runtime the paid latency tracks the chosen duration closely, because time can advance between any two records. And a single pass over a finished bounded input pays no waiting latency at all: the input is already complete when the read begins, so the same completeness question is answered by *when the run starts*, not by how long it holds — which is precisely why teams migrating a nightly computation to a continuous one are surprised to discover they now have a latency-against-completeness dial they never had before.
- Two jobs read the same source but one feeds an alert and the other a monthly reconciliation. Should they use the same wait?No. The wait belongs to the output's purpose, not to the source. The alert should sit at a short hold and accept a slightly incomplete picture, because detection delay is the dominant cost; the reconciliation can hold far longer, or better, re-derive the figure later, because nothing is decided on it in the meantime.
- Does a longer wait make the job fall behind?Not by itself. Backlog is unprocessed records accumulating because throughput is below the input rate; waiting concerns when a group may be treated as finished, and a job can be entirely caught up while deliberately holding results. The indirect risk is resource pressure: more open groups means more state, which can eventually reduce throughput.
- Your consumer asks for both low latency and high completeness. What do you offer?Two figures rather than one compromise: a short hold for the fast provisional number that drives immediate decisions, and a later, more complete figure derived separately for anything that must reconcile. Trying to satisfy both with one wait produces a number that is simultaneously too slow to act on and still not complete.
Holding a train at the platform for passengers still running down the stairs. Every passenger already seated pays the delay, not just the ones you waited for, and each extra minute of holding catches fewer people than the minute before — until you are delaying a full train for one person who has not even reached the station.
saying these in an interview costs you the question
- Thinks only the groups containing a delayed record are slowed down
- Assumes completeness rises in proportion to the wait
- Treats latency as the only cost and ignores retained state
- Chooses the wait without asking what consumes the output
- Believes a long enough wait eventually reaches full completeness
- Confuses waiting deliberately with the job having fallen behind