For a moving window of thirty records, what decides whether the first twenty-nine positions produce a value at all?
answer
- length is not availability
- how many real observations before answering
- a second parameter beside the length
- it also fires on interior holes
basics
~20 sA minimum observation count — how many real observations a window must cover before it answers. At the full length the leading positions hold no value; below it they answer early, from fewer observations than the length implies.
solid answer
~50 sThe length says how far back the window may reach; it does not say how many records must actually be there. A separate **minimum observation count** decides that: the number of real observations a window must cover before it produces a value rather than an absent one. Set it equal to the length and the output begins at position thirty; set it to one and every position answers, the first from a single record. The setting is a decision about comparability, not about tidiness, because nothing in the output distinguishes a value computed from two observations from one computed from thirty — same column, same type, same appearance. It also fires away from the head: absent values inside the window, and quiet stretches under a span-defined window, both reduce the covered observation count below the threshold in the middle of the data. Defaults differ across designs, so state it rather than inherit it.
go deeper
Know that the window's length and the number of observations it must actually have are two different settings, and that the first positions of a sequence do not have enough history behind them.
Explain both directions of the trade: a strict threshold gives comparable values and an empty head, a lenient one gives a full column whose early values rest on almost nothing and look identical to the rest.
Demonstrate that you check the interior too. Positions with no value in the middle of the data mean the threshold is firing on holes or quiet stretches, and that is a finding about the feed, not a cosmetic issue.
The angle is what the pipeline publishes. Argue for carrying the covered-observation count as a first-class column so consumers can make their own reliability decisions, instead of the pipeline making one irreversibly on their behalf.
## What the threshold counts A moving window of fixed length declares how far back it may reach. It does **not** declare how many records must actually be there. Those are two parameters, and the second one — the **minimum observation count** — is how many real observations a window must cover before it produces a value instead of an absent one. Early in a sequence there simply are not thirty records behind the current position. The window covers what exists: one record at the first position, two at the second, and so on. The threshold decides what happens in that region. - Set the threshold equal to the window's length, and the first twenty-nine positions hold no value; the output begins at position thirty. - Set it to one, and every position from the first holds a value — computed from one record, then two, then three. - Set it somewhere between, say ten, and the output begins at position ten, computed from ten observations, then eleven, until the window is full. ## It is not only the leading edge Two other situations reach the same threshold, and both are easy to forget because they happen where nobody is looking. 1. **Absent values inside the window.** A window covering thirty records, some of which hold no value, is covering fewer than thirty *observations*. Whether that position answers depends on the same threshold, in the middle of the data rather than at the head. 2. **A window whose length is a span of time.** On an irregular feed the number of records inside a fixed span changes at every position. A quiet stretch can push the covered count below the threshold, so the output has holes wherever the feed went quiet — which is often exactly where someone will later look for an explanation. ## The defaults differ, so read them Designs in this space do not agree here. Some produce a value only once the window is full. Some produce a value from the very first position. Some expose the threshold as a parameter of its own and some fold it into the operation with no way to change it. A pipeline ported from one to another can silently gain or lose a block of values at the head of every derived column, with nothing raised anywhere. ## The real trade is comparability | threshold | leading edge of the output | what the early values mean | |---|---|---| | equal to the length | a block holding no value | every published value rests on the same number of observations | | one | full from the first position | the first values rest on one, then two, then three observations | | somewhere between | a shorter empty block | a declared floor on how few observations may stand behind a value | The point of the strict setting is not neatness. It is that **nothing in the output distinguishes a value computed from two observations from one computed from thirty**: same column, same type, same appearance on a chart, same treatment by whatever reads it next. A mean of two observations is far more variable than a mean of thirty, and a spread or a quantile over two observations carries almost no information at all — yet published in the same column it will be compared with the rest as though it were the same kind of number. ## What the strict setting costs A block of positions holding no value at the head of a column is not free either: - a later step that drops rows with any absent value removes those rows, and the data is now shorter than the input by an amount nobody wrote down; - if several derived columns use different lengths, the longest one silently decides how much of the head is lost for all of them; - a chart or an aggregate over the whole column now covers a shorter span than its title claims. So the choice is between publishing values you cannot compare and losing rows you did not intend to lose. Both are defensible; neither is a default worth inheriting silently. ## What to do 1. State the threshold explicitly, in the same place you state the length, so a reader sees both decisions together. 2. If you accept early values, carry the number of observations each answer used as a column beside it. A consumer can then filter on it instead of guessing, and a chart can grey out the region rather than pretending. 3. Decide once, at the boundary of the pipeline, whether the incomplete head is trimmed or kept — and if trimmed, record how many rows went and why. 4. Check the interior as well as the head. Count how many positions hold no value in the middle of the data; a non-zero count there tells you the threshold is also firing on holes or on quiet stretches, which is usually news.
- Why is setting the threshold to one not simply the friendliest choice?Because it publishes values of wildly different reliability in one column with nothing to tell them apart. The first answer rests on a single observation and the thousandth on thirty, yet both are plain numbers of the same type. Anything reading that column downstream — a threshold, a chart, an aggregate — treats them identically, and the least trustworthy values sit at the head where people look first.
- If the head of a column holds no values, what should happen to those rows?Decide it deliberately and once. Trimming them is defensible, but it shortens the data by an amount that must be recorded, and if several columns use different window lengths the longest one silently decides the cut for all of them. Keeping them is also defensible, provided every downstream consumer is prepared for positions with no value rather than discovering them by surprise.
saying these in an interview costs you the question
- Treats the empty leading positions as a bug rather than a declared choice.
- Sets the threshold to one so nothing is empty, then compares the values freely.
- Assumes every design leaves the leading positions empty by default.
- Forgets that absent records inside the window also count against the threshold.
- Drops the empty leading rows without noticing the data got shorter.