A feed that reports only when its value changes is re-spaced onto a one-minute grid — what happens to statistics computed afterwards?
answer
- a finer grid invents positions
- rows now count positions
- time-weighted, not report-weighted
- plateaus shrink the spread
- mark what was actually observed
basics
~10 sEvery grid position nothing was measured at now carries a manufactured value, so the rows count grid positions rather than observations. Averages become time-weighted, variability shrinks, and nothing about it raises an error.
solid answer
~40 sMoving to a finer grid cannot create measurements; it creates positions, and something has to stand at each one. For a value that holds until it next changes, the usual choice is carrying the last observed value forward — taking the most recent value at or before the position. That is defensible for the value itself, but it changes what the rows mean. A statistic over the grid weights each value by how long it stood; the same statistic over the raw reports weights each report equally, and the two answer different questions. Spread shrinks, because plateaus repeat one number many times. Positions before the first report have nothing to carry. And an uncapped carry presents a three-day-old reading as a current one. Record which rows were observed, so a consumer can still tell.
go deeper
Know that moving to a finer grid cannot create measurements: every position between two reports holds something the operation produced, not something anyone measured.
Explain the weighting change — a statistic over the grid weights each value by how long it stood, while the same statistic over the reports weights each report equally.
Demonstrate the guards: a column marking observed against manufactured rows, a cap on how far a value may be carried, and never letting a later observation fill an earlier position.
The standing decision is what the team publishes: the raw reports, the regular grid, or the reports with the grid derived on read. Publishing only the grid discards provenance nobody can reconstruct afterwards.
## A finer grid creates positions, not data Re-spacing onto a grid finer than the data's own arrivals produces one position per step regardless of how many records exist. Four reports in a day against a one-minute grid gives 1,440 positions and four observations. Whatever appears at the other 1,436 positions was produced by the operation, not by the instrument. ## What can stand at a manufactured position - **The column's absent marker** — the position exists and is honest about having no value. - **The most recent value at or before it**, carried forward. For a value that holds until it next changes, this is usually the physically correct reading — it looks *backward* in time for its source, which is what makes it safe. - **A later value, filled backward into an earlier position.** This looks *forward* in time for its source, so the row carries something that did not exist at its own stamp. - **An interpolated value**, computed from the observations either side — which also reads a later observation into an earlier row. - **Nothing at all.** In designs where re-spacing only produces keys that records created, a finer grid yields no new rows whatever; you have to construct the positions explicitly first and then match the data onto them. Because the last of those exists, "I re-spaced it finer and the row count did not change" is a real outcome and not a mistake on your part. ## What each statistic becomes once the grid is filled | statistic | over the raw reports | over the filled grid | |---|---|---| | count | number of reports | number of grid positions — fixed by the step, independent of the data | | mean | each report weighted equally | each value weighted by how long it stood | | standard deviation | spread across the reports | smaller, because plateaus repeat one value many times | | correlation with another feed | over positions where both reported | distorted; the effective number of independent observations is far below the row count | | maximum and minimum | over the reports | the same values, but possibly stamped at a position nothing was measured at | The weighting row is the one that matters most. Neither weighting is wrong. "The average reading across the reports" and "the average level over the day" are different questions, and re-spacing silently moves you from the first to the second. A candidate who can say which question the business asked is the one who passes. ## Three hazards worth naming 1. **The leading edge.** Nothing precedes the first report, so leading positions have nothing to carry into them. They stay absent, unless someone fills them from a later value — at which point a row stamped before the first report carries information that did not exist at that stamp, and anything reading that row is seeing the future. 2. **An uncapped carry.** A value carried indefinitely turns a reading from three days ago into an apparently current one. A limit on how far a value may travel from its observation keeps the hole visible instead of papering over it, and makes a stale feed look stale. 3. **Lost provenance.** Once filled, an observed row and a manufactured row are identical in the output. Nothing in the data records which was which unless you add a column that does. ## Working rules - Carry a marker column recording whether each row was observed or manufactured, and propagate it. It costs one narrow column and it is the only way anyone can audit the result later. - Decide the weighting deliberately and write down which question you answered: per report, or per unit of time. - Put a cap on how far a value may be carried, chosen from how long the quantity is physically plausible, not from what makes the chart look continuous. - Never let a later observation fill an earlier position in anything a model will read. The direction in time is the whole distinction; the verbs for the two operations sound almost identical and are opposites. - State the grid in the output's contract, so a consumer knows the row count is a property of the grid rather than a measure of how much was observed. - Check what your own tool does before assuming: some take the filling as part of the same call, some produce positions and leave them absent, and some produce no extra positions at all. ## The line to remember A finer grid never adds information. It adds rows, and every statistic that divides by a row count, or treats rows as independent, quietly changes meaning the moment those rows exist.
- A report gives the average value of a step-valued feed. Why might the grid version and the raw version disagree sharply?Because they weight differently. Across the raw reports, a value that stood for one second counts as much as one that stood for six hours. Across a regular grid, each value counts once per position it covers, so long plateaus dominate. Both are defensible, and the disagreement is a difference of definition rather than a defect.
- Why is filling a leading position from a later report different in kind from carrying an earlier value forward?Because it moves information backwards in time. A row stamped before the first report would carry a value that did not exist at that stamp, so anything reading that row sees something unavailable then. Carrying forward only ever reuses what was already known at the position being filled.
- What does a cap on how far a value may be carried actually protect?It protects the distinction between a quiet feed and a dead one. Without a cap, a feed that stopped reporting looks identical to one whose value has not changed. With a cap chosen from how long the quantity stays plausible, the gap reappears as absent values and a downstream check can see it.
saying these in an interview costs you the question
- Treating a carried value as an observation
- Reading the row count as a count of measurements
- Assuming an average over the grid equals the average of the reports
- Filling an earlier position from a later report without noticing the direction
- Carrying a value forward with no limit on its age
- Believing more rows means more information