skip to content

Two pipelines bucket the same event stream into 15-minute totals with the same step, yet the counts differ — why?

level: seniorimportance: should knowfreq 44%

answer

  1. a step is not yet a grid
  2. where the first edge is pinned
  3. an origin taken from the data moves
  4. only the origin within one step matters

basics

~20 s

Because a step alone does not fix a grid. The first bucket edge is pinned somewhere, and if it is pinned to the earliest stamp in the input, a different input span produces different edges and therefore different groupings.

solid answer

~40 s

A grid is two things: a step, and a position the first edge is pinned to. Designs supply that second thing differently — from the earliest stamp present in the data, from a calendar convention such as midnight, or from a position you declare. The data-pinned form is the trap, because a historical run over a different span, a re-run after a late record arrives, or a first run that starts mid-period all move the earliest stamp and therefore move every edge. Same step, different edges, different records sharing a bucket, and nothing raises. The fix is to pin the origin to a declared constant so the re-spacing is a pure function of span, step and origin — and to reconcile two runs by comparing their edges, not their totals.

go deeper

for a junior

Know that a grid needs a starting point as well as a step: fifteen-minute buckets can legitimately begin at :00 or at :07, and the two group records differently.

for a middle

Explain that shifting the origin by a whole step changes nothing at all, so what matters is its position within one step, and that this is what decides which records share a bucket.

for a senior

Show the operational form: pin the origin to a declared constant so a historical re-run and an incremental run produce identical edges, and reconcile two runs by comparing stamps rather than totals.

for a principal

The judgment is whether bucket edges belong in the pipeline's published contract. Declaring them lets every consumer join on stamps safely; leaving them implicit makes every cross-system comparison guesswork.

## A grid is two numbers, and most people set only one Re-spacing lays a regular sequence of positions over data that carries a time ordering. "Every fifteen minutes" describes the spacing between edges; it does not say where the edges fall. Fifteen-minute buckets beginning at :00, :15, :30, :45 and fifteen-minute buckets beginning at :07, :22, :37, :52 are both correct grids with the same step, and they group the records completely differently. The second number is **where the grid starts**: the position the first bucket edge is pinned to. Only its position *within one step* matters — shifting the origin by a whole step reproduces exactly the same set of edges — so the useful way to think about it is the remainder, not the absolute value. ## Where the origin can come from | source of the origin | what makes it move | safe for repeated runs? | |---|---|---| | the earliest stamp present in the input | any change to which rows the run read | no | | a calendar convention, such as the start of the day or of the week | a change of the convention itself | mostly, if the convention is stated | | an absolute position declared in configuration | nothing but an explicit edit | yes | | the moment the run started | every run | no | The first and last rows are the ones that produce the symptom in the question: two runs, one step, different edges, no error. ## The three ways the earliest stamp moves underneath you 1. **A historical re-run.** The historical run reads a wider span than the incremental run, so its earliest stamp is older and its grid is anchored elsewhere. The two outputs then cannot be compared bucket for bucket, and appending one to the other produces a sequence whose edges change partway through. 2. **A late record.** A record arrives after the first run and before the re-run, earlier than anything the first run saw. The re-run's grid shifts, and buckets that were never supposed to change do. 3. **A partial first period.** The very first run of a new feed starts mid-hour, so the grid is pinned to an arbitrary minute and stays pinned there for every subsequent run that inherits the same anchoring rule. ## What it corrupts, and what it does not - **Per-bucket totals change**, because different records share a bucket. - **The grand total does not change.** Both runs still partition the same records, so summing everything agrees and reconciling on the total proves nothing. - **The row count barely moves** — at most by one at each end — so a shape check will not catch it either. - **Two feeds re-spaced separately cannot be put onto one grid** unless both used the same origin, which makes the anchor a cross-pipeline concern rather than a local one. - **A comparison against yesterday's output is off by a fraction of a bucket**, which looks like small genuine movement rather than like a defect. ## Working rules - Declare the origin as a constant and pass it in. Never let the rows a run happened to read determine the grid. - Treat the re-spacing as a pure function of the span, the step and the origin. If two of the three are fixed, the output must be reproducible. - Assert the first edge of the result equals the declared origin plus a whole number of steps. That single assertion catches every case above. - Reconcile two runs by comparing stamps: take the difference between their earliest stamps modulo the step. A non-zero remainder proves different edges, whatever the totals say. - For a weekly step, remember the anchor also chooses a weekday. Two weekly grids with the same step and different anchors are offset by whole days and will never line up. - Publish the origin alongside the data if anyone downstream will join on your stamps. ## The sentence that makes it stick A step tells you how wide the buckets are. An origin tells you where they are. Only the second one is usually left to whatever the data happened to contain, and it is the one that makes two correct runs disagree.

  • Shifting a grid's origin by exactly one step changes nothing. Why?
    Because the grid repeats with a period equal to the step. Moving the first edge a whole step earlier or later regenerates the same set of edges, so the same records share the same buckets and only the extent at the ends can differ. What matters is the origin's position within a single step, not its absolute value.
  • How would you detect that two runs used different bucket edges when both results look plausible?
    Compare stamps, not values. Take the earliest stamp of each result and difference them modulo the step; a non-zero remainder proves the edges differ. Totals will not reveal it, because both runs partition the same records and the grand total agrees either way.
  • Why does this bite harder when each entity's data is re-spaced in its own run?
    Because an origin taken from the data is then taken from that entity's earliest stamp, so every entity gets its own grid. The outputs cannot be set side by side on a shared sequence of positions, and any later comparison between entities is comparing differently-aligned buckets.

saying these in an interview costs you the question

  • Believing the step alone determines the buckets
  • Letting the earliest stamp in the batch anchor the grid
  • Reconciling two runs by totals instead of bucket edges
  • Assuming a re-run over a wider span reproduces the same buckets
  • Expecting a mismatch of this kind to raise an error