skip to content

A team halves how far behind the newest moment seen their pipeline holds its completeness claim - its assertion that nothing older is still coming - to freshen a dashboard. What did they trade?

level: seniorimportance: should knowfreq 55%

answer

  1. one dial, pulling in both directions
  2. everyone pays, a few lose
  3. the curve flattens quickly
  4. not a throughput setting
  5. no finite wait covers everything

basics

~20 s

Latency for correctness, directly. Periods become eligible to close sooner, and more records now arrive after their period was already declared finished. The gain is paid on every result; the loss falls only on the tail of the source's behaviour.

solid answer

~50 s

The distance the claim is held behind the newest observed moment is the single dial on this exchange, and it runs both ways. Holding it further back means every result waits longer and fewer records miss their period; moving it closer means every result arrives sooner and more records miss. Two properties make the decision harder than it looks. First, the exchange is **asymmetric in who pays**: the latency saving lands on every consumer of every group, while the correctness loss lands only on whatever runs behind - often one region, one device class, or one flaky producer, not a random sample. Second, it is **non-linear**: lateness has a long thin tail, so the first part of the wait covers most of it and each further unit buys progressively less. Halving the wait rarely doubles the misses, and doubling it rarely halves them.

go deeper

for a junior

Know the direction of the exchange: waiting less means fresher results and more records missing their period; waiting more means the reverse.

for a middle

Explain that the wait is not a throughput setting, and that the latency is charged to every result while the completeness loss falls only on records running behind.

for a senior

Bring the shape of the distribution into it - steep then flat - and say who actually loses data when the wait is cut, rather than quoting a percentage.

for a principal

Own it as a published promise: what freshness the organisation commits to, what completeness it commits to, and which of the two is allowed to move without a conversation.

## One dial, two directions A completeness claim - the running assertion that nothing older than some timestamp is still coming - is usually held a chosen duration behind the newest moment the job has observed. That duration is the wait, and it is the whole of the exchange: | direction | what improves | what degrades | who notices first | |---|---|---|---| | move the claim closer to the newest moment (aggressive) | every time-based result becomes eligible sooner | more records arrive after their period was declared finished | nobody, unless the misses are counted; the numbers are quietly short | | hold the claim further back (conservative) | fewer records miss their period | every result is delayed by the extra wait, whether or not anything late ever comes | everybody, immediately - the dashboard lags | The asymmetry in visibility is the reason this dial drifts in one direction over a system's life. Freshness complaints arrive by ticket; a two percent shortfall in a daily figure arrives, if at all, months later in a reconciliation. ## Why the exchange is not linear How far behind records actually run is a distribution, not a constant, and in practice it is heavily concentrated near zero with a long thin tail: most records are barely out of order, a few are minutes behind, a very few are hours or days behind because a device was out of coverage. Whatever the precise shape for a given source - measuring it is a separate exercise - two consequences follow directly: - The first part of the wait is worth far more than the rest. Early units of waiting cover the dense part of the distribution; later units cover progressively thinner air. - There is no wait long enough to be safe. A producer offline for a week defeats any bound anyone is willing to pay for, so beyond some point the honest conclusion is that the remaining tail should be handled by whatever the pipeline separately promises about records that miss, not by waiting longer. This is why 'we doubled the wait and the misses barely moved' is the normal outcome rather than a surprise, and why the team in the question may find that halving the wait costs less completeness than anyone feared - and still be wrong to do it, if the records they now lose belong to one identifiable population. ## What the latency actually is The delay this dial buys and sells is the gap between a period ending and that period being treated as finished. It is not throughput: the job processes records exactly as fast as before, and nothing accumulates unprocessed because of it. A job can be perfectly caught up on its input and still be deliberately holding results back, and an aggressive claim will not clear a backlog by one record. Confusing the two produces the classic wrong fix - shortening the wait to make a job that is behind on volume 'catch up'. ## Two dials, not one The wait is how far behind the claim runs. A separate question is whether anything is permitted after the claim has already passed a period's end - whether a straggler may still change an answer that was eligible to close - and that behaviour, along with what becomes of a record that misses entirely, is a different mechanism with its own contract. Keeping them apart matters in an interview, because 'increase the wait' and 'allow a late update' solve overlapping problems at very different costs: the first delays everyone, the second delays nobody and demands more of every downstream consumer. ## What varies between designs - Some runtimes express the wait once, where the record's moment is assigned, and propagate the resulting claim; others allow it per input or per operator, so 'the wait' may not be a single number at all. - Where continuous work runs as a rapid succession of small finite jobs, the job interval is added to whatever wait is configured, so the achievable freshness floor is higher than the wait alone suggests. - Where the source reports its own progress rather than the job inferring it from payloads, the wait can be much smaller for the same completeness, because the assertion rests on the source's promise instead of a guess. ## Answering well Name the exchange - latency on every result against completeness in the tail - then add the two things a junior answer misses: the cost is paid by everyone and the benefit taken from a specific few, and the curve is steep at first and flat later, so the right move is rarely 'round number, doubled'. Finish by saying what you would tell consumers about a figure published under the new setting.

  • Who pays the latency and who pays the correctness?
    Everyone pays the latency: the wait is added to every group's result, whether or not a single straggler ever arrives for it. The correctness cost falls only on whatever runs behind, and that is rarely a random sample - one region, one device class, one producer that buffers. An average-case gain funded by a specific minority's data is worth naming out loud before it ships.
  • Is there a wait long enough to make the claim safe?
    No finite one. A device offline for a week defeats any bound anyone would pay for, and each extra unit of wait is charged to every result while covering a thinner slice of the tail. Past the point where the curve flattens, the remaining records are better handled by whatever the pipeline promises about records that missed their period than by waiting longer for them.
  • Would shortening the wait help a job that is running behind on volume?
    Not at all. The wait governs when a period is treated as finished, not how fast records are processed; a job with unprocessed records piling up has a capacity or throughput problem. Shortening the wait would leave the backlog untouched and quietly reduce completeness at the same time - the worst of both outcomes.

saying these in an interview costs you the question

  • Thinks a shorter wait makes the job process records faster
  • Assumes the lost completeness shows up on its own without being counted
  • Believes doubling the wait roughly doubles completeness
  • Says a long enough wait removes the trade-off entirely
  • Treats the wait as free because most records are on time
  • Assumes the records lost are a random sample of all records