skip to content

A reading stamped exactly at 09:00 lands on a bucket edge of an hourly grid — which bucket takes it, and what stamps that row?

level: middleimportance: must knowfreq 58%

answer

  1. two knobs, not one
  2. an edge value joins one bucket only
  3. closure decides membership, naming decides the stamp
  4. same stamp, different records

basics

~20 s

Two independent defaults decide it: which edge closes each bucket, so an edge value joins the bucket before it or the one after, and which edge names the row, so the same fold may be stamped 09:00 or 10:00.

solid answer

~40 s

A bucket runs between two edges, and exactly one of those edges belongs to it — that is the first choice, and it decides which bucket a record landing exactly on an edge joins. The second, quite separate, is which edge the resulting row is stamped with; that moves no records at all, only labels. Both are usually left at a default, and the defaults are not the same everywhere or even across steps within one tool. The consequence worth saying out loud is that two combinations can both produce a row stamped `09:00` while those rows hold different records. So a stamp is not evidence of which records were folded, and reconciling two systems means checking the edge conventions before comparing any numbers.

go deeper

for a junior

Know that a bucket includes one of its two edges and excludes the other, and that a reading landing exactly on an edge therefore belongs to one bucket rather than both.

for a middle

Explain the two choices separately: one decides which records share a bucket, the other decides only what the row is called, and they combine into four legitimate outcomes.

for a senior

Show that you diagnose by stamps before numbers — a uniform one-step offset is a naming difference, while a few disagreeing buckets point at closure and at records sitting on edges.

for a principal

The angle to own is that these conventions belong in the published contract. Once a consumer joins on your stamps, an edge convention is an interface, and changing it is a breaking change.

## One call, two independent decisions When stamped records are re-spaced onto a regular grid — a regular sequence of positions laid over the data, with everything falling between two successive edges folded into one value — the edges have to partition the records exactly. Two separate conventions do that work, and both usually arrive as defaults nobody chose. 1. **Which edge closes the bucket.** A bucket lies between two edges, and exactly one of them belongs to it. If the earlier edge belongs, the bucket covers that edge and everything up to but not including the later one, so a reading stamped exactly at 09:00 joins the bucket running 09:00 to 10:00. If the later edge belongs, the bucket starts just after the earlier edge and runs up to and including the later one, so the same reading joins the bucket running 08:00 to 09:00. 2. **Which edge names the resulting row.** Once the fold has happened the row needs a stamp, and the two natural candidates are the bucket's earlier edge and its later edge. This choice decides only the label; it does not move a single record. The two are orthogonal, giving four combinations, all defensible. Tools set defaults for both, those defaults are not the same across tools, and within one tool they are not always the same for every step. ## The four combinations, for one reading | edge that closes | edge that names | bucket the 09:00 reading joins | that row's stamp | |---|---|---|---| | earlier | earlier | 09:00 to 10:00 | 09:00 | | earlier | later | 09:00 to 10:00 | 10:00 | | later | earlier | 08:00 to 09:00 | 08:00 | | later | later | 08:00 to 09:00 | 09:00 | The first and last rows are the trap. Both produce a row stamped 09:00, and those two rows hold different records. **A stamp is not evidence of which records a row folded.** ## Why a bucket takes only one of its two edges Buckets are half-open — one edge in, one edge out — so that every record lands in exactly one of them. If both edges were inclusive, a record falling exactly on an edge would be folded into two buckets at once, and the sum of the buckets would exceed the sum of the input. If neither were, an edge record would vanish. Half-open intervals make the partition exact, and the only remaining question is which end is the open one. This matters far more on machine-generated feeds than on human-entered data, because instruments report on the second, on the minute, and at midnight. On such a feed a large share of records sit exactly on an edge, so the closure choice is not an edge case at all. ## Telling the two apart when two systems disagree - **Every value offset by one position, with the values matching one for one** — that is the naming choice. The groups of records are identical; only the stamps moved. - **Most buckets agreeing and a few differing** — that is the closure choice, and the ones that differ are the buckets that had a record sitting exactly on an edge. - **The grand total agreeing while a per-stamp comparison disagrees** — true of either, which is why a grand total is no evidence at all here. - **A figure crossing a reporting line at a month end** — a single record stamped at midnight on the first can be folded into the previous period or the next one, and only the closure convention says which. ## Working rules - State both choices explicitly in the code and in the specification, even when the default is the one you wanted; a default you relied on silently is a default someone can change. - Before comparing two systems' numbers, compare their stamps: if the earliest stamps differ by exactly one step, you are looking at a naming difference and the numbers are probably fine. - Where records can land exactly on an edge — whole minutes, whole hours, midnight — write down which bucket owns them, because that is the part a reader cannot infer from the output. - Do not assume a tool applies the same convention to a monthly step as to an hourly one. Verify per step, not per tool. - When a downstream consumer joins on the stamp, the naming choice becomes part of your contract with them, not an internal detail.

  • Two systems re-space the same feed hourly and every total in one is stamped an hour later than the other. Which of the two choices explains it?
    Which edge names the row. A different naming choice moves every stamp by one step and leaves the groups of records untouched, so the values match one for one once you shift the stamps. A closure difference instead moves individual edge records between buckets, so per-bucket totals change while the grand total stays the same.
  • Why are buckets half-open rather than closed at both ends?
    So that each record lands in exactly one bucket. With both edges inclusive, a record sitting exactly on an edge would be folded twice and the buckets would sum to more than the input. Half-open intervals make the partition exact; all that remains is to say which end is open.
  • On which kind of feed does the closure choice matter most?
    Machine-generated ones. Instruments report on whole seconds, whole minutes and midnight, so a large fraction of records sit exactly on a grid edge and move as a block when the convention changes. On irregular human-entered data an exact edge hit is rare, and the same choice is nearly invisible.

Marks on a tape measure. A cut at exactly the 3 cm mark belongs to one of the two segments either side, and somebody has to say which. Quite separately, you can call that segment "the 3 cm segment" after where it starts or "the 4 cm segment" after where it ends. Two conventions, both arbitrary, and neither one is recoverable from the measurement itself.

saying these in an interview costs you the question

  • Thinking the row's stamp tells you which records it holds
  • Assuming a record on an edge is counted in both buckets
  • Treating closure and naming as a single setting
  • Assuming every tool closes and names the same edge
  • Reading a one-step shift as a data problem rather than a labelling choice