A job carries a timestamp asserting nothing older than it is still expected. Why is that a guess rather than an observed fact?
answer
- the interesting part has not arrived yet
- an observation plus an assumption
- conclusions carrying assumptions are inferences
- wrong in two directions, not one
- one error is silent, the other loud
basics
~20 sNothing in an endless input reports what has not yet been sent. The job infers the claim from the moments it has seen plus a standing assumption about how far behind a record may run, so it can be too early or too late.
solid answer
~50 sCompleteness is a property of the future, and no component holds it. A producer that has sent nothing for an hour may be idle, disconnected, or about to send an hour of buffered records. So the job builds the claim from two ingredients: an **observation** - the largest moment assigned so far, or what the source reports about its own progress - and an **assumption**, the duration it agrees to wait for records running behind. A conclusion carrying an assumption is an inference, and it fails in both directions. Too aggressive, and records genuinely still coming arrive after their period was declared finished; the numbers are simply short and nobody is told. Too conservative, and every result waits for records that never existed. The aggressive error is silent and the conservative error is loud, which is why teams drift toward aggressive.
go deeper
Remember the slogan and why it holds: completeness is asserted, not observed, because nothing in an endless input ever reports what has not been sent yet.
Take it apart into an observation and an assumption, then name both failure directions - declaring a period finished too soon, and holding every result back for records that never existed.
Point out the asymmetry: the aggressive error is silent unless someone counts what missed, so an unexamined pipeline drifts toward it under pressure to look fresh.
Treat the assumption as a published policy, not a setting. Whoever owns the number owes consumers a statement of what was assumed and what may still change afterwards.
## There is nothing to observe An **unbounded input** is a promise of more records, not a set of them. At any instant the job has seen a prefix, and the interesting question is about the part it has not seen. No component anywhere in the system holds the fact 'this producer has nothing older than 09:00 left to send'. Producers are numerous, uncoordinated, sometimes disconnected, and sometimes not yet deployed. A source that has sent nothing for ten minutes is indistinguishable, from inside the job, from a source that is about to deliver ten minutes of buffered records. Completeness, in other words, is a property of the future. It cannot be read; it can only be asserted. **A completeness claim - a watermark - is exactly that assertion: a timestamp the job carries with the records, saying no record with an older moment is still expected.** ## How the guess is actually built The claim has two ingredients, and only one of them is data: - **An observation.** Usually the largest moment assigned to any record seen so far - occurrence time, the moment stamped into the payload by whatever produced the record. Sometimes something better: metadata in which the source reports how far each of its parallel inputs has progressed. - **An assumption.** The **disorder bound**: a chosen duration the job holds the claim behind that observation, standing for 'a record may run this far behind, and no further'. Nobody validated that duration against the future; it is a policy. A conclusion that carries an assumption is an inference. That is the whole of the answer, and everything else follows from it. The weakest variant is the accidental one: a pipeline that never assigned a moment at all is grouping by arrival, which is the same guess made without anyone deciding to make it - and it is the guess that a re-run silently invalidates, because arrival order is a property of the run and not of the data. ## Both directions of error | the claim is | what happens | who notices | |---|---|---| | too aggressive - ahead of reality | records that were genuinely still coming arrive after their period was already declared finished | nobody, unless the misses are deliberately counted; the published numbers are simply short | | too conservative - behind reality | every result is held back waiting for records that never existed | everyone: dashboards lag, downstream consumers wait, someone files a ticket | The asymmetry is the point an interviewer is listening for. One error costs latency and announces itself; the other costs correctness and is invisible by default. Teams under pressure to make a dashboard feel fresh therefore drift toward the silent error, and the drift is not recorded anywhere. ## Why a guess is nevertheless the right design The alternative to guessing is waiting for certainty, and certainty never arrives on an endless input, so that alternative produces no answers at all. A stated, tunable guess converts an unanswerable question - 'is the hour complete?' - into an answerable policy: 'we treat an hour as finished when our claim has passed its end, on the assumption that nothing runs more than N behind'. That policy can be written down, argued about, changed, and measured against reality. An implicit guess cannot. This is also why 'the claim was wrong' is usually not a defect report. A claim that is occasionally wrong is working as designed; a claim that is wrong in a direction and by an amount nobody chose is the actual defect. ## What varies between designs - Where a source reports its own progress per parallel input, and guarantees ordering within each one, the claim can be derived from that report rather than inferred from payloads. It remains an assertion - it now rests on the source's promise instead of on a guessed duration - but it can be far tighter. - A single pass over a finished bounded input makes no guess whatsoever: the end of the input is the boundary, and completeness is observed. - Runtimes differ in where the assumption is expressed. In some it is stated once, where the record's moment is assigned, and propagates; in others each operator recomputes a claim from its inputs. The assumption is the same; the place you would change it is not. ## Saying this in an interview The sentence that scores is 'completeness is asserted, not observed'. Follow it with the two error directions, note that the aggressive one is silent unless someone counts the misses, and finish by saying what the pipeline promises consumers about a number it published while the claim was still a guess.
- If the claim is a guess, what can a downstream consumer actually rely on?A stated policy rather than a guarantee: results are complete up to the claim provided nothing runs further behind than the assumed bound, plus whatever the pipeline separately promises about records that miss their period. A consumer that treats a published figure as final when the pipeline only ever promised 'complete under this assumption' has taken on a risk nobody offered it.
- Can a claim ever be exact rather than inferred?It can get close. Where the source itself reports how far each parallel input has progressed, and guarantees ordering within each, the claim is derived from that report rather than guessed from payloads. It is still an assertion, because it inherits the source's promise: exact only to the extent that promise holds, and worthless if a producer writes records with moments out of order behind the source's back.
Deciding when to serve dinner while guests are still arriving. Nobody telephones to say 'that is everyone'. You infer it from how late people usually are, and you can be wrong twice over: serve now and two more walk in, or hold the food for an hour for a guest who was never coming. The rule you use is useful precisely because waiting for certainty means nobody eats.
saying these in an interview costs you the question
- Treats the claim as a guarantee that nothing older will arrive
- Thinks a long enough wait eventually makes the claim exact
- Believes a claim that was wrong means the pipeline is misconfigured
- Names only the too-early error and never the too-conservative one
- Assumes the job can verify completeness against the source at runtime