A table's row labels are timestamps that are not in ascending order — what does that quietly break?
answer
- a precondition nobody verifies
- boundary search needs sorted labels
- a plausible subset, not an error
- appending is how order is lost
- assert non-decreasing at the boundary
basics
~20 sAsking for a range by two moments, and anything that walks the ordering, assume the labels are non-decreasing. Out of order you may get a refusal, a slow full comparison, or a plausible wrong subset. Verify the order yourself.
solid answer
~50 sTime-labelled rows come with an unwritten precondition: the labels run in non-decreasing order. A range between two moments can then be answered by locating two boundary positions and taking the block between them, which is both cheap and exact. On unordered labels a boundary search finds *a* position rather than *the* position, and the block between two such positions is an arbitrary subset. Designs respond differently — some verify and refuse, some fall back to comparing every row, and some walk the input as if ordered and hand back a plausible table. The same assumption spreads to anything that steps along the ordering: a span walking the rows, a grid laid over them, a match against a second feed. Confirm the labels are non-decreasing where the data enters, once, rather than hoping the operation checked.
go deeper
Know that operations over time-labelled rows expect the labels to run in non-decreasing order, and that the rows do not arrive that way just because they carry stamps.
Explain why the order is what makes a range cheap — two boundary positions and the block between them — and name the three things a tool may do when the order does not hold.
Show where in a pipeline the order is established and what re-establishes it after an append, and explain how you would recognise a result that was silently short.
Decide whether the order is a producer's promise or your own assertion, and make that a standing rule: a claim you do not verify is a dependency on someone else's retry behaviour.
## Why order is a precondition here at all When a table's rows are named by timestamps and those timestamps run in **non-decreasing order** — each label at least as large as the one before it — a request for "everything between these two moments" has a cheap exact answer. The tool locates the position of the first boundary and the position of the second and returns the contiguous block between them. No row outside the block has to be looked at. That shortcut is correct *only* because of the order. On labels that are not sorted, a boundary search still terminates and still returns a position; it just is not the position you wanted. The block between two such positions is a contiguous run of rows that happens to sit there, and it can include moments outside the requested range and exclude moments inside it. **Nothing about the result looks wrong**: it has rows, the rows have plausible stamps, and the count is plausible too. ## The three behaviours you can actually get | Behaviour | What you see | The trap | |---|---|---| | Verify and refuse | An error naming the ordering | None — this is the friendly case, and you cannot count on it | | Fall back to a full comparison | The right rows, more slowly | You never learn the order was wrong; it bites later somewhere that does not fall back | | Proceed on the assumption | A subset that looks fine | The answer is silently wrong, and the row count will not tell you | The honest summary is that **whether the check happens is a property of the design, not of the operation's meaning**. Treat the precondition as yours to establish. ## Where the assumption spreads Range selection is the visible case. The same assumption is underneath anything that steps along the ordering: - **A span of fixed length walking the rows** covers whichever rows are physically adjacent. If the rows are out of order, "the last seven" is seven rows that are not the seven most recent moments. - **Laying a regular grid over the data** assigns records to positions by their stamps but reports them in the order it walked them; an unsorted input can produce positions out of sequence or double-visited. - **Matching two feeds on their stamps** walks both inputs once on the assumption that both advance in time. Out of order, the pointer on one side moves past a record it should have matched, and the result is silently short. - **Reading "the last value"** means the last *row*, which is the latest *moment* only when the order holds. - **Reading each row's value from the previous position** gives the previous row, which is the previous moment only under the same condition. ## Establishing it yourself 1. **Ask the question directly.** Many tools expose a way to ask whether a set of labels is non-decreasing, computed once and remembered, which makes the check nearly free to repeat. 2. **Where they do not, compare the column with a copy of itself read one position earlier**, and require every difference to be at least zero. That is one whole-column expression, not a loop, and it can also point at the first violating position. 3. **Decide which order you actually need.** Non-decreasing allows neighbours with equal stamps; strictly increasing does not. Operations differ in which one they want, and a data set with two readings at one instant satisfies the first and fails the second. 4. **Do the check where the data enters**, once, rather than defensively before every operation. A single assertion at the boundary is worth more than five scattered ones downstream. ## Order is a precondition, not an operation Putting rows into order is its own piece of work with its own decisions, and it is not what this is about. The point here is narrower and easier to miss: several time operations behave as though the order already holds, and most of them will not tell you when it does not. Two habits cover it — establish the order at the boundary, and re-establish it after any step that appends. **Appending is the usual way a sorted table stops being sorted**: a late record added at the end carries an earlier moment than its neighbour, and from that row on, every shortcut above is resting on a condition that no longer holds. ## What this looks like in a review The question to ask of a time pipeline is not "is the data sorted" but "**where is that established, and what re-establishes it after the append**". A pipeline that answers "the source emits in order" is relying on a producer's behaviour it does not control, which is a different kind of claim from a check.
- How do you check the order without writing a per-row loop?Compare the label column with a copy of itself read one position earlier and require every difference to be at least zero. That is a single whole-column expression, it costs one pass, and the position of the first negative difference tells you where the order first breaks. Many tools also expose a ready-made way to ask the same question.
- Two records carry exactly the same moment. Does that break monotonic order?Not necessarily. Non-decreasing order allows equal neighbours; strictly increasing order does not. Which one an operation needs varies, so a data set with simultaneous readings can satisfy one requirement and fail the other. Find out which your step wants before you conclude the data is fine.
- The source says its records are emitted in time order. Is that enough?It is a claim about a producer you do not control, not a check. Retries, parallel senders and a restart can all break it, and a single out-of-order append is enough to invalidate every shortcut that follows. Keep the claim, add the assertion, and let the assertion be the thing you rely on.
saying these in an interview costs you the question
- Believes an out-of-order range selection always raises rather than returning something
- Treats "the rows came from a stamped log" as a guarantee of order
- Checks only the first and last labels and calls the order verified
- Thinks only range selection cares, and windows and grids do not
- Confuses the order rows are stored in with the order a report displayed them