skip to content

A dashboard needs numbers within seconds, but a group is only declared finished minutes later: how do you decide what to publish?

level: principalimportance: should knowfreq 40%

answer

  1. not fast versus correct
  2. publish a defined wrongness
  3. state magnitude, duration, audience
  4. mark what is still moving
  5. two surfaces beat one compromise

basics

~20 s

Decide what wrongness you will publish, not whether to publish any. Emitting before a group is declared finished is defensible when the size and duration of the error are stated, the number is visibly provisional, and each consumer gets the contract it can actually apply.

solid answer

~50 s

This is a product decision dressed as a configuration question, and the honest form of it is: **how wrong may a number be, for how long, in front of whom?** Waiting for the group to be declared finished gives one settled value and a latency floor no extra capacity will lower. Emitting early gives seconds, and pays in corrections a downstream system must absorb — which is only possible if that system can replace a value it already holds. So decide in this order: what each consumer can apply, then the error budget for a provisional number, then the contract. The cleanest outcome is usually two surfaces rather than one compromise — a revisable fast one for the screen, a settled insert-only one for the record — with the provisional surface marked as such so nobody cites it as final.

go deeper

for a junior

Recall that a number published before its group is declared finished can still change, and that a reader has no way to know that unless something says so on the number itself.

for a middle

Explain the components of the settled contract's latency — the completeness wait, the grace period, the runtime's unit of progress — and why adding workers reduces none of them.

for a senior

Show that you choose the contract from what each consumer can apply, and that you would run two surfaces rather than force one interval on consumers with different tolerances.

for a principal

Turn it into a stated error budget: magnitude, duration, audience, visibility, settlement. Publishing a bounded, marked wrongness is a product decision you can defend; publishing an unbounded one is a liability nobody agreed to.

## The trade, stated plainly Nothing over an endless input is ever known to be complete; a job only carries a running assertion that no record older than a stated moment will still arrive, and a group can be declared finished when that assertion passes its end. Everything published before that moment is provisional by construction. Everything published after it waits for it. So the decision is not "fast or correct". It is **which defined wrongness you are prepared to publish, to whom, and for how long** — and a lead who frames it that way gets a real answer from the business, while one who frames it as latency tuning gets "both, please". ## Three publishable positions | Position | First number appears | What is wrong with it | What the consumer must do | |---|---|---|---| | Wait for the group to be declared finished | After the completeness wait, plus any grace period | Nothing, within the claim's own assumption | Insert one row | | Emit early, correct later | As soon as the early firing condition allows | Under-counts by whatever has not arrived yet | Replace a value it already holds | | Restate continuously | Effectively at the runtime's smallest unit of progress | Always current, never settled | Hold a latest-value view and absorb the write rate | None of these is the advanced answer. The advanced answer is knowing which consumers are on which row, and refusing to put two consumers with different tolerances on the same row. ## Writing the wrongness down An error budget for a provisional number states five things, and all five are answerable before any code is written: 1. **Magnitude.** How far below the settled value may an early number sit — as a percentage, and in the units the reader actually acts on. 2. **Duration.** How long may it stay wrong before the corrected value lands. This is the early firing interval plus the completeness wait plus the grace period, not any one of them. 3. **Audience.** Who reads it, and what decision do they take from it. A figure on an internal screen that refreshes and a figure quoted in a monthly statement are not the same risk. 4. **Visibility.** How the reader can tell this number is still moving. A marker on the emission, a column in the destination, a label on the screen — but something, because an unmarked provisional number is indistinguishable from a settled one and will eventually be cited as final. 5. **Settlement.** What event makes it final, and whether the consumer is told. If nothing in the contract distinguishes the last emission, no consumer can ever safely stop waiting. If the business cannot answer (1) and (3), the requirement for sub-second numbers was never real and the conversation should return to the latency the settled contract can offer. ## When the consumer cannot revise at all Some destinations genuinely cannot replace a value: an insert-only record, a downstream export, a system owned by another team on a change schedule you do not control. Three ways out, in order of preference: - **Two surfaces.** A revisable fast surface for the screen and a settled insert-only surface for the record, both fed by the same job under different contracts. The provisional numbers never enter the permanent record, and each consumer gets the contract it can apply. This costs a second write path and the discipline to keep the names distinct. - **Publish only settled values to that consumer** and accept its latency. Honest, and often correct once the error budget conversation has happened. - **Insert provisional rows with their own identity and a marker**, then have the reader filter for settled ones. This works, and it is the option most likely to be misused later by someone who queries the table without the filter. What does not work is sending a revision stream into an insert-only destination and hoping: it accumulates rows, the sums double-count, and nothing errors. ## What varies between runtimes The latency floor is not yours to set alone. Where arrivals are collected for a short span and one finite job runs over the collected set, the earliest possible emission is that span's end, so "within seconds" has a hard floor built into the design. Where records advance one at a time through long-lived operators, the floor is much lower and the cost shows up as write volume instead. Where the model is a finite pass re-run over history, there is no early emission at all — only a more frequent run. A plan that promises sub-second provisional numbers should say which of these the platform is, because on one of them the promise cannot be kept at any price. ## The failure mode to design against The damaging outcome is not a wrong number. It is an unmarked wrong number that somebody screenshots. Once a provisional figure has been pasted into a report, the correction that lands two minutes later reaches nobody, and the next conversation is about whether the pipeline can be trusted rather than about the error budget everyone agreed to. Marking provisional output, and naming the moment it settles, is the cheapest part of this design and the part that is skipped most often.

  • What is the latency floor of the settled contract actually made of?
    Three parts: the wait until the job's assertion that no older record will arrive passes the group's end, any grace period during which the group still accepts records, and the runtime's own unit of progress. None of the three shrinks by adding machines, which is why the answer to a freshness requirement is a contract change, not capacity.
  • Why is running two surfaces usually better than one compromise interval?
    Because the consumers have different capabilities, not just different tastes. One can replace a value and one cannot, so a single middle setting gives the screen numbers that are too slow and the permanent record numbers it cannot apply. Two contracts from one job cost a second write path and no extra correctness argument.
  • What if the business will not state an error budget?
    Then publish settled values only, and say so. Without a stated magnitude and audience, every corrected number becomes an incident after the fact, and the team carries a risk it was never authorised to take. Refusing to publish unbounded provisional numbers is the defensible default.

saying these in an interview costs you the question

  • Frames the choice as fast versus correct
  • Publishes provisional numbers with nothing marking them provisional
  • Assumes more machines will lower the completeness wait
  • Feeds a revision stream to an insert-only record
  • Cannot say what event settles a number
  • Promises sub-second output without checking the runtime's floor