skip to content

After rows are removed, a saved offset still resolves while a saved row label no longer does — why, and which failure is safer?

level: seniorimportance: should knowfreq 46%

answer

  1. one scheme is recomputed, one is not
  2. positions describe the current layout
  3. the stale offset still resolves
  4. prefer the failure that announces itself
  5. read back an identifying value

basics

~20 s

Removing rows renumbers positions contiguously while leaving the survivors' labels alone, so an offset resolves against the new layout and names a different record. The label failure is safer: it cannot hand back a plausible wrong row.

solid answer

~50 s

A removal compacts the rectangle. Positions are a description of the current layout, so they are recomputed from the start and every offset past the removal point now names a different record — and still resolves, because the number is still inside the row count. The surviving rows keep the labels they had, so the set of labels is simply non-contiguous, and a label belonging to a removed row has nothing to match. What happens then is design-specific: some raise, some hand back a row of absent values, some return an empty result. The point is that none of those quietly gives you a different record, which is exactly what the offset does. The safer failure is the one that cannot be mistaken for success, so references that must survive a step boundary are labels or identifier values, never positions.

go deeper

for a junior

Recall the asymmetry: removing rows renumbers positions but leaves the survivors' labels untouched, so a position saved beforehand no longer points at the same record.

for a middle

Explain why the stale offset still resolves — the number is inside the new row count, so the request succeeds — and why a set of labels with gaps in it is perfectly normal rather than a sign of damage.

for a senior

Demonstrate the operational judgment: choose the scheme whose failure announces itself, keep positions out of anything that crosses a step boundary, and prove the selection with a read-back of an identifying value.

for a principal

Set the rule and defend it by the failure class it removes: identifiers cross boundaries, positions are derived at the point of use, row counts are asserted across restrictions, and labels are never renumbered for tidiness.

## What a removal does to each scheme Restricting a table — keeping some rows and discarding the rest — has an asymmetric effect on the two ways of naming a row, and the asymmetry is the whole answer. - **Positions are recomputed.** An offset is a description of where a row sits in the current layout, and the layout just changed. The survivors are numbered contiguously from the start again, so a row that was at position 900 may now be at position 612. Any offset captured before the removal now names whichever row landed there. - **Labels are left alone.** A label belongs to the row, so the survivors keep theirs. The set of labels is now non-contiguous, with gaps where the removed rows were, and that is fine — a set of labels has no obligation to be dense. That is the mechanical difference. What makes it an interview question is the second-order consequence. ## Loud against silent | the saved reference | what happens after the removal | how you find out | |---|---|---| | an offset inside the new row count | resolves, returns a different record | you do not; the output is plausible | | an offset past the new row count | fails, or is clamped, depending on the design | immediately, if it fails | | a label that survived | resolves, returns the same record | nothing to find out | | a label whose row was removed | does not match | at the lookup, in whatever way the design reports it | The first row of that table is the dangerous one, and it is dangerous for a structural reason rather than a careless one: **an offset that has gone stale is indistinguishable from an offset that is correct.** The request succeeds. The row has the right columns and the right types. Its values are in range. Every assertion anyone thought to write about shape still passes. The only thing wrong with it is that it is somebody else's record, and no downstream step has any way to notice. A failed label lookup has the opposite property. It cannot be mistaken for success, so it stops the run at the step that is actually wrong, with the identifier that is actually missing in hand. That is a better outcome even though it is the one that looks worse in the log, and saying so plainly is what the question is fishing for. ## "It raises" is not universal — say which behaviour you mean This is where an answer that only knows one tool falls over. The designs in this family do not agree on what an unmatched label does: - some **raise**, treating a missing label as a programming error; - some hand back a **row of absent values**, treating the lookup as an alignment against a requested set of labels; - some return an **empty result**, treating the lookup as a match that found nothing; - and where a design offers no row labels at all, the question does not arise, because the only address is the offset and there is nothing durable to save. The common property across all of them — and the one you can rely on — is that **none of them returns a different record**. That is the claim worth making, rather than asserting a particular error behaviour that is true of the tool you happen to know. One more variation to keep straight: some designs will renumber the labels for you on request after a restriction, producing fresh consecutive labels. That is a legitimate thing to want, but it destroys the property this whole question rests on. After a renumbering, a saved label is no safer than a saved offset, because the label no longer identifies the row — it describes a position under another name. ## The discipline this implies 1. **Nothing crosses a step boundary as a position.** If a reference has to survive into the next step, into a cache, into a report or into tomorrow's run, it is a label or an identifier value. Offsets are derived at the point of use and discarded. 2. **Derive and consume in one expression.** "Order by this, then take the first ten" is safe because nothing can intervene. "Order by this" in one step and "take rows 0 to 9" in another is a promise that no one will ever insert a step between them. 3. **Assert the row count across a restriction.** You usually know roughly how many rows should survive. An assertion there catches the restriction that removed far more than intended, which is the event that silently invalidates every offset downstream. 4. **Prove which record you got.** When a specific record matters, read back an identifying value from the row you selected and compare it against what you asked for. It is the only portable check, and it costs one line. 5. **Do not renumber labels out of tidiness.** Consecutive labels look neater and are strictly less informative; the gaps were telling you something. ## What separates a senior answer A middle answer explains the renumbering correctly and stops. A senior answer goes on to the operational point: the failure mode you want is the one that announces itself, so you deliberately choose the addressing scheme whose failure is loud, and you spend a line on a read-back where the record identity actually matters. Then it says what varies, because "it raises" is one design's answer and not the class's.

  • How would you prove, in one line, that a selection returned the record you intended?
    Read an identifying value back out of the selected row and compare it with what you asked for. It does not depend on which design you are on, it catches both a stale offset and a label that matched the wrong number of rows, and it fails at the selection rather than three steps later in a total nobody recomputes by hand.
  • Why is renumbering labels after a restriction a bad habit, even though it looks tidier?
    Because it converts identities into positions under another name. Once the survivors are relabelled consecutively, a label saved earlier no longer identifies the record it was taken from, and the one property that made labels safe across a step boundary is gone. The gaps in a non-contiguous set of labels are information, not untidiness.
  • A design offers no row labels at all. What plays the label's role?
    A value you designate: a column of identifiers carried alongside the numbers, or a parallel sequence matched by position. References across step boundaries are written in terms of that value, and the offset is re-derived from it at the point of use. The discipline is the same; it is just not supplied by the tool.

saying these in an interview costs you the question

  • Expects a removal to leave gaps in the positions of the survivors.
  • Thinks a stale offset will fail rather than return a different row.
  • Asserts every design raises on a label that cannot be matched.
  • Treats the loud failure as the worse outcome because it stops the run.
  • Stores a position between steps and re-uses it on the next run.
  • Renumbers labels after a restriction for tidiness, losing the identity.