skip to content

Time-Indexed Data

Operating along a time ordering rather than across records: re-spacing it, sliding a window over it, and matching two feeds sampled on different clocks. The failures are silent and directional.

on this pageshow

explore

questions

23

Daily sensor readings are re-spaced onto a monthly grid, and one month holds no readings. What can appear at that position?

level: juniorimportance: must knowfreq 66%

answer

  1. the position nobody measured
  2. absent row, filled row, or no row
  3. grid built first, or keys produced
  4. count expected positions, compare row count

basics

~20 s

One of three things, depending on the design: a row carrying an absent value, a row the same call already filled, or no row at all. Confirm which by comparing the returned row count with the months the span should contain.

solid answer

~50 s

Re-spacing to a regular grid means laying a regular sequence of positions over the data and deciding what value sits at each one. For a position no record fell into, designs differ in a way that is easy to miss. Some build the full grid first and then place records into it, so every month exists and an unobserved one appears with an absent value. Some take a fill rule as part of the same call, so the hole is never materialised and the month appears already carrying a value. Others implement the re-spacing as a grouping on a truncated stamp, so a month that received nothing is a key that was never produced and the row is simply missing. The resolving check is the same in all three: work out how many positions the span should contain, and compare that with the row count you got back.

go deeper

for a junior

Know that a coarser grid can contain a position no record fell into, and that what shows up there is not always a row holding an absent value.

for a middle

Be able to describe all three outputs and say which one a grouping-shaped implementation produces, and why a full-grid implementation cannot produce the same thing.

for a senior

Show the habit rather than the fact: compute the expected number of grid positions, compare it with the row count, and probe a span you know is empty before trusting anything downstream.

for a principal

The standing call is whether pipelines may emit sparse grids at all. A declared complete grid costs a little storage and buys every consumer the right to assume one row per position.

## What re-spacing does **Re-spacing to a regular grid** — laying a regular sequence of positions over data that carries a time ordering, and deciding what value sits at each one — is what the market calls *resampling*: the data is re-expressed at positions you chose rather than at the positions it happened to arrive on. It needs two things. The first is **the target spacing**, the declared step of the grid; when that step is a calendar unit such as a month, it is a rule rather than a fixed number of seconds. The second, whenever the grid is coarser than the data, is **an aggregate** that folds whatever records fall inside each bucket into one value. Nothing in either of those inputs says what should happen at a position no record fell into. That decision is made for you, and it is made differently by different designs. ## The three things that can appear at an unobserved position | how the operation is built | what the unobserved month looks like | how you notice | |---|---|---| | the full grid is constructed first, then records are placed into it | the row exists and its value is the column's absent marker | the row count matches the span, and the hole shows in the values | | the fill rule is taken as part of the same call | the row exists and already carries a value | the row count matches the span and nothing looks wrong at all | | the operation is a grouping on a truncated stamp | there is no row | the row count is short by exactly the unobserved positions | The third is the one that surprises people, and it is not a defect. Splitting records by a derived key, applying a fold to each piece and recombining — the shape re-spacing has underneath in several designs — can only produce keys that some record actually created. A month that received nothing never becomes a key, so it never becomes a row. The first design instead materialises every position the step implies between the span's ends, and only then discovers that one of them received nothing. ## Why the difference is not cosmetic - A **row count** over the result means two different things: the number of positions in the span, or the number of positions that received at least one record. - Two sequences re-spaced separately and then set side by side only line up position for position if both carry every position. One sparse result against one complete result no longer share a grid. - A reader scanning the rows will not notice a month that is absent from them altogether, whereas an empty-valued row is conspicuous. - What an aggregate then does with an absent value is a separate question, but the two designs hand it different input, so numbers can already diverge before anything downstream has made a choice. - Any step that assumes a fixed number of rows per year — twelve monthly positions, fifty-two weekly ones — works on the complete shape and breaks quietly on the sparse one. ## The check that settles it, in any design 1. From the first and last positions you intend to cover and the declared step, compute how many grid positions the span should contain. 2. Compare that number with the row count that came back. Equal means the grid is complete; short means positions were never produced, and the difference is exactly how many. 3. Point the operation at a span you know is empty and read what comes back. One run tells you which of the three behaviours you have, and it is worth recording next to the pipeline. That check costs two lines and it is the portable one. It does not depend on knowing which design you are on, and it keeps working when someone swaps the implementation underneath it. ## What the operation needed in order to find the ordering One more thing varies, and it changes how the question is even phrased. In some designs the time operations dispatch on the type of the row labels, so the timestamp has to be **promoted to the row labels** — made the thing rows are named and ordered by, rather than one more field alongside the others — before re-spacing is possible at all. In others there is no row-label concept, and you name the time column as an argument instead. Both reach the same grid, and neither changes the answer above. That is exactly why the row-count check is the thing to carry between tools, rather than the shape of the call. ## The short version to say out loud A coarser grid can contain positions that received nothing. Whether such a position arrives as an absent value, as an already-filled value, or not at all is a property of the tool, not of the data. Measure it once, write it down, and assert the row count from then on.

  • A re-spaced monthly result has eleven rows for a full calendar year. What are the two explanations, and how do you separate them?
    Either the span genuinely ends early, or one month received no records and the implementation produced no row for it. Separate them by reading the first and last stamps in the result: if they bracket the whole year, an interior position was never produced, and the stamp that is absent from the sequence names it.
  • Why is a row with an absent value more useful than no row at all, when neither carries data?
    Because it asserts that the position exists. A consumer can count rows, set two re-spaced sequences side by side position for position, and see the hole. A missing row is indistinguishable from a shorter span, and it silently changes how many positions everything downstream believes it has.

saying these in an interview costs you the question

  • Assuming every month always appears in the output
  • Treating a missing row and an absent value as the same thing
  • Reading a short result as a short span rather than dropped positions
  • Assuming nothing in the output was manufactured by the call itself
  • Checking only the first and last rows instead of the row count
open as a page

What does making the timestamp the row labels of a table buy, and what does it cost?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Promoting the timestamp makes it what rows are named and ordered by, so a range can be asked for by naming two moments and time operations find the ordering unasked. Only one of the record's times can hold that slot.

open as a page

A column of timestamps arrives with no zone attached. What does each value actually state, and what can you not do with it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A stamp with no zone is a clock reading, not a point on the world timeline: it says what a clock showed, not when. Until the recording zone is attached, it cannot be safely compared with, or converted for, another source.

open as a page

A quote feed and a price feed never share a timestamp, so how is each quote matched to the price in effect then?

level: juniorimportance: must knowfreq 48%

basics

~20 s

Match each quote to the most recent price record whose stamp is at or before the quote's stamp - an inexact ordered match rather than an equality match. The driving side sets the row count: at most one matched record each.

open as a page

Over an ordered sequence, how does a moving window of fixed length differ from a window anchored at the start?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A moving window of fixed length keeps its edges a fixed distance apart, so old records drop out as new ones enter. A window anchored at the start never moves its left edge, so each answer covers everything so far.

open as a page

A reading stamped exactly at 09:00 lands on a bucket edge of an hourly grid — which bucket takes it, and what stamps that row?

level: middleimportance: must knowfreq 58%

basics

~20 s

Two independent defaults decide it: which edge closes each bucket, so an edge value joins the bucket before it or the one after, and which edge names the row, so the same fold may be stamped 09:00 or 10:00.

open as a page

A stamp with no zone is subtracted from a zone-carrying instant. What can a tool do, and why might the answer depend on the machine?

level: middleimportance: must knowfreq 62%

basics

~20 s

Three behaviours are live across designs: refuse the mixed operation, adopt the process's own configured zone for the zoneless side, or assume UTC for it. Two of the three return a plausible number that is wrong by whole hours.

open as a page

A feature column was smoothed with a window centred on each row; why could its value at that row's stamp not have been computed then?

level: seniorimportance: must knowfreq 57%

basics

~20 s

A centred window is stamped in the middle of the records it covers, so roughly half of them come from after that stamp. The value therefore encodes observations that had not happened at the moment the row describes.

open as a page

Monthly totals from a re-spaced daily series show February well below January — is that a real decline?

level: middleimportance: should knowfreq 51%

basics

~20 s

A calendar month is 28 to 31 days, so a per-bucket sum scales with the bucket's width and February sits below its neighbours even at a flat rate. Divide each total by its own bucket's width before reading a decline.

open as a page

A table's row labels are timestamps that are not in ascending order — what does that quietly break?

level: middleimportance: should knowfreq 52%

basics

~20 s

Asking for a range by two moments, and anything that walks the ordering, assume the labels are non-decreasing. Out of order you may get a refusal, a slow full comparison, or a plausible wrong subset. Verify the order yourself.

open as a page

Two readings share one timestamp in a table labelled by time — what changes when you ask for that label?

level: middleimportance: should knowfreq 44%

basics

~20 s

Asking for a repeated label returns every row carrying it, so the result is a sub-table rather than a single row. The shape now depends on the data rather than on the code, and in several designs nothing raises.

open as a page

Adding one day to a stamp in a zone that shifts its clocks: which two answers can that produce, and when do they differ?

level: middleimportance: should knowfreq 48%

basics

~20 s

Calendar arithmetic targets the same clock time on the next calendar day; duration arithmetic adds exactly twenty-four hours. Across a clock shift they land on different moments, because the calendar day is twenty-three or twenty-five hours long.

open as a page

A match to the latest earlier reading always succeeds while any earlier reading exists, so what is a staleness tolerance for?

level: middleimportance: should knowfreq 38%

basics

~20 s

A staleness tolerance caps how old a matched value may be. Beyond that age the match is abandoned and an absent value is produced instead, so a reading from hours ago is never silently presented as the value in effect now.

open as a page

For a moving window of thirty records, what decides whether the first twenty-nine positions produce a value at all?

level: middleimportance: should knowfreq 48%

basics

~20 s

A minimum observation count — how many real observations a window must cover before it answers. At the full length the leading positions hold no value; below it they answer early, from fewer observations than the length implies.

open as a page

On an irregularly-sampled feed, why does a moving window of seven cover different records when seven means days rather than records?

level: middleimportance: should knowfreq 60%

basics

~20 s

A count-defined window always covers exactly seven records and therefore a variable span of time. A span-defined window always covers exactly seven days and therefore a variable number of records. On irregular data the two select different sets.

open as a page

A feed that reports only when its value changes is re-spaced onto a one-minute grid — what happens to statistics computed afterwards?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Every grid position nothing was measured at now carries a manufactured value, so the rows count grid positions rather than observations. Averages become time-weighted, variability shrinks, and nothing about it raises an error.

open as a page

Two pipelines bucket the same event stream into 15-minute totals with the same step, yet the counts differ — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because a step alone does not fix a grid. The first bucket edge is pinned somewhere, and if it is pinned to the earliest stamp in the input, a different input span produces different edges and therefore different groupings.

open as a page

A record carries when the event happened, when it reached you and when it became valid — which belongs in the row labels?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The moment your questions are asked against, which for analysis is normally when the event happened, because that is the ordering the world had. Label by arrival instead and the same source re-run produces different numbers. Keep all three as columns.

open as a page

A column of local wall-clock readings crosses a clock shift. Which two local times cause trouble when a zone is attached, and why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

At a clock shift one local hour occurs twice and one never occurs. A repeated reading names two possible instants; a skipped one names none. Attaching a zone must therefore raise, choose a side, or move the value.

open as a page

A table stores each row's UTC offset rather than the zone it came from. What can you no longer answer from it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An offset is one answer from a zone's dated rulebook. Storing it pins the moment but throws the rule away, so the captured moment and its clock reading survive while local time at any other moment cannot be worked out.

open as a page

A feature table matched each label row to the nearest sensor reading by stamp, and the model scored far better offline than live - what did the match do?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Nearest takes whichever reading is closer, earlier or later. Every row where the later reading was closer carries a value that did not exist at the label's stamp, so the table contains information from the future. Only the latest-at-or-before direction cannot.

open as a page

A moving average over one table holding many devices' readings looks wrong at each device's first rows; what happened?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The window walked straight across the boundary between devices, so each device's first answers mix in the tail of the device before it. The window has to be restricted so it never reaches past an entity boundary.

open as a page

An inexact ordered match over a table of many sensors returned a full, plausible result with no error - which two preconditions went unchecked?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

That the match stayed inside one sensor, and that both inputs were in stamp order. Neither is inferred: without an entity restriction one sensor's reading attaches to another's record, and an out-of-order input is walked as if ordered, producing a plausible table.

open as a page