Daily sensor readings are re-spaced onto a monthly grid, and one month holds no readings. What can appear at that position?
answer
- the position nobody measured
- absent row, filled row, or no row
- grid built first, or keys produced
- count expected positions, compare row count
basics
~20 sOne of three things, depending on the design: a row carrying an absent value, a row the same call already filled, or no row at all. Confirm which by comparing the returned row count with the months the span should contain.
solid answer
~50 sRe-spacing to a regular grid means laying a regular sequence of positions over the data and deciding what value sits at each one. For a position no record fell into, designs differ in a way that is easy to miss. Some build the full grid first and then place records into it, so every month exists and an unobserved one appears with an absent value. Some take a fill rule as part of the same call, so the hole is never materialised and the month appears already carrying a value. Others implement the re-spacing as a grouping on a truncated stamp, so a month that received nothing is a key that was never produced and the row is simply missing. The resolving check is the same in all three: work out how many positions the span should contain, and compare that with the row count you got back.
go deeper
Know that a coarser grid can contain a position no record fell into, and that what shows up there is not always a row holding an absent value.
Be able to describe all three outputs and say which one a grouping-shaped implementation produces, and why a full-grid implementation cannot produce the same thing.
Show the habit rather than the fact: compute the expected number of grid positions, compare it with the row count, and probe a span you know is empty before trusting anything downstream.
The standing call is whether pipelines may emit sparse grids at all. A declared complete grid costs a little storage and buys every consumer the right to assume one row per position.
## What re-spacing does **Re-spacing to a regular grid** — laying a regular sequence of positions over data that carries a time ordering, and deciding what value sits at each one — is what the market calls *resampling*: the data is re-expressed at positions you chose rather than at the positions it happened to arrive on. It needs two things. The first is **the target spacing**, the declared step of the grid; when that step is a calendar unit such as a month, it is a rule rather than a fixed number of seconds. The second, whenever the grid is coarser than the data, is **an aggregate** that folds whatever records fall inside each bucket into one value. Nothing in either of those inputs says what should happen at a position no record fell into. That decision is made for you, and it is made differently by different designs. ## The three things that can appear at an unobserved position | how the operation is built | what the unobserved month looks like | how you notice | |---|---|---| | the full grid is constructed first, then records are placed into it | the row exists and its value is the column's absent marker | the row count matches the span, and the hole shows in the values | | the fill rule is taken as part of the same call | the row exists and already carries a value | the row count matches the span and nothing looks wrong at all | | the operation is a grouping on a truncated stamp | there is no row | the row count is short by exactly the unobserved positions | The third is the one that surprises people, and it is not a defect. Splitting records by a derived key, applying a fold to each piece and recombining — the shape re-spacing has underneath in several designs — can only produce keys that some record actually created. A month that received nothing never becomes a key, so it never becomes a row. The first design instead materialises every position the step implies between the span's ends, and only then discovers that one of them received nothing. ## Why the difference is not cosmetic - A **row count** over the result means two different things: the number of positions in the span, or the number of positions that received at least one record. - Two sequences re-spaced separately and then set side by side only line up position for position if both carry every position. One sparse result against one complete result no longer share a grid. - A reader scanning the rows will not notice a month that is absent from them altogether, whereas an empty-valued row is conspicuous. - What an aggregate then does with an absent value is a separate question, but the two designs hand it different input, so numbers can already diverge before anything downstream has made a choice. - Any step that assumes a fixed number of rows per year — twelve monthly positions, fifty-two weekly ones — works on the complete shape and breaks quietly on the sparse one. ## The check that settles it, in any design 1. From the first and last positions you intend to cover and the declared step, compute how many grid positions the span should contain. 2. Compare that number with the row count that came back. Equal means the grid is complete; short means positions were never produced, and the difference is exactly how many. 3. Point the operation at a span you know is empty and read what comes back. One run tells you which of the three behaviours you have, and it is worth recording next to the pipeline. That check costs two lines and it is the portable one. It does not depend on knowing which design you are on, and it keeps working when someone swaps the implementation underneath it. ## What the operation needed in order to find the ordering One more thing varies, and it changes how the question is even phrased. In some designs the time operations dispatch on the type of the row labels, so the timestamp has to be **promoted to the row labels** — made the thing rows are named and ordered by, rather than one more field alongside the others — before re-spacing is possible at all. In others there is no row-label concept, and you name the time column as an argument instead. Both reach the same grid, and neither changes the answer above. That is exactly why the row-count check is the thing to carry between tools, rather than the shape of the call. ## The short version to say out loud A coarser grid can contain positions that received nothing. Whether such a position arrives as an absent value, as an already-filled value, or not at all is a property of the tool, not of the data. Measure it once, write it down, and assert the row count from then on.
- A re-spaced monthly result has eleven rows for a full calendar year. What are the two explanations, and how do you separate them?Either the span genuinely ends early, or one month received no records and the implementation produced no row for it. Separate them by reading the first and last stamps in the result: if they bracket the whole year, an interior position was never produced, and the stamp that is absent from the sequence names it.
- Why is a row with an absent value more useful than no row at all, when neither carries data?Because it asserts that the position exists. A consumer can count rows, set two re-spaced sequences side by side position for position, and see the hole. A missing row is indistinguishable from a shorter span, and it silently changes how many positions everything downstream believes it has.
saying these in an interview costs you the question
- Assuming every month always appears in the output
- Treating a missing row and an absent value as the same thing
- Reading a short result as a short span rather than dropped positions
- Assuming nothing in the output was manufactured by the call itself
- Checking only the first and last rows instead of the row count