On an irregularly-sampled feed, why does a moving window of seven cover different records when seven means days rather than records?
answer
- seven of what, exactly
- records fixed, or time fixed
- irregular spacing separates the two
- count-defined reaches back further after a gap
basics
~20 sA count-defined window always covers exactly seven records and therefore a variable span of time. A span-defined window always covers exactly seven days and therefore a variable number of records. On irregular data the two select different sets.
solid answer
~50 sThe word "seven" is doing two different jobs. A **count-defined window** covers the current record and the six before it in the ordering, so it always has seven records and covers however much time those seven happened to take. A **span-defined window** covers every record whose stamp falls within seven days of the current position, so it always covers seven days and has however many records happened to arrive. They coincide only when the spacing is exactly one unit with no holes and no duplicate stamps. After a six-hour outage a count-defined window of seven is still dutifully averaging seven readings, most of them hours old, and calls the result "recent"; a span-defined window of thirty minutes at the same position has one record and is enormously noisier than its neighbours. In most designs a bare number means records, and asking for a span requires a duration instead — but that is a convention, so read the length you actually wrote.
go deeper
Learn to ask "seven of what?" before anything else. One reading fixes the number of records and lets the time vary; the other fixes the time and lets the number of records vary.
Explain which reading a bare number usually gets, and what changes for each form when the feed goes quiet or bursts. Be able to say why the count-defined answer after an outage is misleading rather than merely stale.
Show the diagnosis. Given a metric that misbehaved after an incident, work out from the covered record counts and covered spans which form was in use, and say what the metric's own description implied it should have been.
The angle is definition ownership. A metric phrased as "the trailing thirty days" has a contractual meaning; argue for the span form as the default in the shared layer so that arrival-rate changes cannot silently redefine what a number reported upward means.
## The same word, two selections "A window of seven" is ambiguous, and the ambiguity is not academic: the two readings select genuinely different records from the same data. - A **count-defined window** covers a fixed number of *records* — the current one and the six before it in the ordering. It therefore covers a **variable span of time**, however long those seven records happened to take. - A **span-defined window** covers a fixed amount of *time* — every record whose stamp falls within seven days of the current position. It therefore covers a **variable number of records**, however many happened to arrive. They coincide only when the observations are spaced exactly one unit apart, with no holes and no duplicate stamps. Real feeds rarely oblige. In most designs a bare number means records, and a span has to be asked for with a duration rather than an integer. That is a convention rather than a law, and it is one of the places where reading the length you actually wrote matters more than remembering what you meant. ## Worked: what an outage does Take a sensor that normally reports every five minutes. | situation | count-defined window of 7 records | span-defined window of 30 minutes | |---|---|---| | normal, five-minute spacing | 7 records, covering 30 minutes | 7 records, covering 30 minutes | | just after a six-hour outage | 7 records, covering 6 hours 5 minutes | 1 record, the new one | | a burst of ten readings a minute | 7 records, covering 42 seconds | about 300 records | The count-defined window's first answer after the outage is labelled a recent average and is computed almost entirely from readings six hours old. Nothing marks it, and the column's type, length and range all look fine. The span-defined window is honest about the gap — it covers what actually arrived — but its answer at that position is computed from a single record and is therefore far more variable than the answers around it. A threshold tuned on the dense stretch will fire on it. Honest is not the same as safe. ## Holes are not quite the same as irregular spacing A hole is a position where something was expected and nothing was observed; an irregular feed may have no notion of expected positions at all. Both make a count-defined window's reach in time vary, but they differ in what you can do about them. Laying a regular sequence of positions over the data first makes the two window forms coincide again, at the cost of deciding what value sits at each position nothing was observed at — a decision with consequences of its own, and not one to make casually just to tidy a window. There is a second interaction worth naming: a span-defined window can cover records whose value is absent. The window covers seven days' worth of records but fewer than that many real observations, which is what the minimum observation count is measured against. ## What a span-defined window needs in order to exist A count-defined window needs only an ordering. A span-defined window needs the tool to know **which values the span is measured against** — which column carries the ordering. Designs differ in how that is supplied. In some, the time operations dispatch on the type of the row labels, so the timestamps must first be promoted to what rows are named and ordered by; in others the ordering column is named as an argument and the table stays a plain set of typed columns throughout. Some designs offer no span-defined form at all, and the requirement then has to be assembled by hand. Two further preconditions apply to the span form specifically: the ordering must be monotonic, and duplicate stamps need a defined disposition, because a span's edge can land in the middle of a run of identical stamps. ## Which to choose - Choose **span-defined** when the quantity is physical or contractual — the average over the last hour, the total in the trailing thirty days. A count of records is not what that sentence means, and it will silently redefine the metric whenever the arrival rate changes. - Choose **count-defined** when the records are the unit — the last twenty trades, the last fifty builds — or when the arrival rate is genuinely fixed and you want a predictable amount of arithmetic per position. - Either way, when the spacing is irregular, emit the **number of records each answer covered** as a column beside the answer. It costs one column and it is the difference between a number you can defend and one you cannot. ## The check 1. Pick the sparsest stretch in the data and the densest one. 2. For one position in each, write down the stamp of the earliest record the window covered, the stamp of the position itself, and the count of records between them. 3. If the elapsed time differs wildly between the two stretches, you have a count-defined window. If the record count differs wildly, you have a span-defined one. 4. Either is fine. What must be true is that the one you have is the one the sentence describing the metric claims.
- Which of the two forms is affected by duplicate stamps at the same position?Both, differently. A count-defined window counts duplicates as separate records, so a run of identical stamps eats the window's capacity and shortens its reach in time. A span-defined window either takes all of them or splits a run at its edge, depending on whether the edge is inclusive — which is a defaulted choice worth reading rather than assuming.
- If the feed is dense and regular, does the distinction still matter?It matters the moment that stops being true, which is usually an incident rather than a design change. A count-defined window on a feed that has always been regular is correct right up until the first outage, and the answer it produces then is the one most likely to be looked at. Writing the span form costs nothing extra and removes the failure mode.
saying these in an interview costs you the question
- Says a window of seven always covers a week.
- Assumes a count-defined window is stable in time when observations are irregular.
- Believes the two forms only diverge when records are missing.
- Expects every design to offer a span-defined window at all.
- Reads a span-defined answer without checking how many records it covered.