A stakeholder asks for an alert on errors in the last five minutes. What must be pinned down before it can be computed?
answer
- a duration is not a boundary
- three readings, different numbers
- grid, rolling, or continuously trailing
- also key, clock, predicate
- burst straddling a grid boundary
basics
~20 sThe phrase fixes only a length. Still undefined: whether the five minutes advances continuously or resets on a grid, how often it is recomputed, what groups the records together, which moment it is measured against, and what counts as an error.
solid answer
~50 sThe phrase states one number and leaves the shape open. *Last five minutes* could mean back-to-back five-minute spans, where a fresh count starts on the grid every five minutes; or a five-minute span restarted every step — every minute, say — so the figure rolls; or no interval at all, just a running count per key with a five-minute expiry on what it counts. These give different numbers for the same input, and a spike straddling a boundary is invisible to the first and obvious to the second. Four other things are undefined: the **refresh step** if the shape is the rolling one, the **grouping key** (per service, per host, or one global figure), which **moment** the five minutes is measured against — the one the pipeline assigned from the record, or the one the worker observed — and the predicate for *error*. Settle those five and the grouping is stated; the threshold and the alert contract are separate conversations.
go deeper
Notice that the phrase gives a length and nothing else. Ask whether the five minutes restarts on a grid or rolls, and ask what the count is grouped by before writing anything.
Lay out the readings and show they differ on the same input: a burst straddling a grid boundary is split between two groups, while a rolling shape always has a group containing all of it.
State the cost of the reading you recommend — a refresh step multiplies output rows and puts the same burst in front of the threshold many times — and hand back the decisions that are not yours.
Push on what the alert is for. If the point is to notice a failing service quickly, the honest design may be a running value with an expiry and no interval-shaped output at all, which changes what every consumer downstream is built to read.
## Why the phrase is not yet a computation Every aggregate over an endless input needs a boundary, and a phrase like *the last five minutes* names a duration without naming a boundary. Three different shapes all honour those words: | reading | shape | behaviour | |---|---|---| | a fresh five-minute count, on the grid | back-to-back equal spans | one result every five minutes; the count restarts from zero at each boundary | | a rolling five minutes | five-minute span restarted every step | one result per step; consecutive results share most of their records | | always the trailing five minutes | no interval; a running value with a five-minute expiry | one continuously restated value per key | They are not interchangeable. On the same input the first will report a quiet interval where the second reports a breach, because a burst split across a grid boundary is divided between two groups and neither total crosses the threshold, while a rolling span always has a group that contains the whole burst. ## The five things to settle 1. **The shape.** Which of the three readings above. This is the question to ask first, because the others depend on it. 2. **The refresh step**, if the shape is the rolling one. Five minutes refreshed every minute is five groups covering each record and five times the output rows; refreshed every ten seconds it is thirty of each. The step is bought, not free. 3. **The grouping key.** One global count, one per service, one per host, one per customer. This decides how many results exist at once and, just as importantly, what the alert means: a global figure hides one host failing completely, and a per-host figure alerts on every deployment. 4. **The moment the boundary reads.** The one the pipeline assigned from the record, describing when the thing happened, or the one the worker observed. Which to choose, and what it does to correctness, is a separate subject — but the requirement is not complete until it says which, and a pipeline that never assigned a moment is silently grouping by arrival. 5. **The predicate.** What an error is: a status class, a level in a log line, a named exception category. Interviewers rarely dwell here, but a grouping over an unspecified filter is not a specification. ## The things that are adjacent and should be named, not solved A good answer names these and hands them back rather than silently picking: - **When the result is handed downstream**, and whether a group may be handed down more than once as a corrected value, is a separate decision that can be attached to any of the three shapes. - **Whether a record that turns up after its group was declared finished is accepted**, and what happens to a number already published, is likewise separate. - **What holding the groups open costs** is a real constraint but a different subject: what belongs in the shape conversation is only that a shorter refresh step means proportionally more open groups. - **The threshold and the alert policy** are not part of the grouping. A rolling shape re-evaluates the same underlying burst against many overlapping groups, so whoever owns alerting has to expect repeats. ## Worked translation The answer an interviewer wants sounds like a proposal with the choices exposed: - *I would read it as a five-minute span restarted every minute, per service, measured on the moment carried on the record, over log lines at error level or worse.* - *That is five overlapping groups covering every record, so five rows per service per five minutes, and a burst is caught wherever it falls rather than only when it aligns with a grid.* - *If the cost of those rows is not wanted, the alternative is back-to-back five-minute spans, and the cost of that is a burst split across a boundary going unreported.* - *If they want a figure that is never stale, the third reading — a running value with a five-minute expiry on what it counts — is closest to what people mean by the last five minutes, and it produces no interval-shaped rows at all.* ## Why this is asked The question tests whether a candidate can turn business language into a definite grouping without inventing the missing parts silently. Choosing the grid shape because it is the cheapest to compute, and not saying so, is the failure mode: the number is then wrong in a specific, predictable way that nobody downstream knows about. Naming the shape, the step, the key, the clock and the predicate — and saying which of them you guessed — is the whole of the skill here.
- Why would back-to-back five-minute spans miss a burst that a rolling shape catches?Because a burst can straddle a boundary. Ninety seconds of errors split as sixty seconds in one group and thirty in the next leaves neither total above a threshold that the combined ninety would have cleared. A rolling shape restarts a span every step, so some span contains the whole burst regardless of where it falls.
- What does choosing a one-second refresh step for a five-minute figure buy and cost?It buys detection delay bounded by a second instead of a minute. It costs three hundred groups covering every record, so three hundred times the result rows for the same input, three hundred open groups per key at any moment, and a threshold that sees the same burst three hundred times. The step should be chosen against the reaction time anyone actually has.
- Does the grouping key change the answer or only the number of rows?Both, and the first matters more. A global count of errors can sit comfortably below a threshold while one host is failing every request, and a per-host count will alert on every routine restart. The key decides what the alert is actually watching, so it is a requirement question, not a partitioning detail.
saying these in an interview costs you the question
- Picks the grid shape silently because it is cheapest to compute
- Treats five minutes as fully specifying the grouping
- Never asks what the count is grouped by
- Assumes the boundary is measured on arrival without saying so
- Promises a rolling figure and delivers a count that resets