When a table is reordered, which of a row label and a row offset still names the same record, and why?
answer
- two names for the same row
- identity against coordinate
- one survives reordering, the other relabelling
- the row carries its label with it
- a stale offset still resolves
basics
~20 sThe row label still names the same record; the offset does not. A label is an identity the row carries with it, while an offset only describes where the row happens to sit right now.
solid answer
~40 sTwo addressing schemes reach the same rectangle. Addressing by label names a row by what it is called, so the reference survives a reordering and breaks only if the labels themselves are rewritten. Addressing by offset names a row by how far down it currently sits, so it survives a relabelling and breaks the moment the rows move — a reorder, a removal, an insert. Neither scheme is more correct; they answer different questions, and the mistake is not picking the wrong one but not noticing there was a choice. Designs differ in how you ask: some expose two separately named addressing surfaces, some overload a single subscript and infer the scheme, and an unlabelled rectangle of numbers has no labels at all, so an offset is the only name a row there has.
go deeper
Recall the one-line distinction and be able to say it without hedging: a label names what the row is, an offset names where it currently sits. Knowing which of the two survives a reordering is the whole of the expected answer.
Explain the mechanics: why a stale offset resolves successfully instead of failing, what a removal does to the numbering of the survivors, and what changes when a design offers one overloaded subscript instead of two named surfaces.
Show the production judgment. Name the steps that invalidate an offset between derivation and use, say why the silent wrong record is worse than a reference that cannot resolve, and describe how you would prove which record you actually got.
Frame it as a contract rather than a preference: where offsets are permitted in a codebase, what must never cross a step boundary as a position, and what class of silent failure the rule buys you against the readability it costs.
## Two names for one row A rectangle of data with named columns can carry one further piece of structure: a **row label**, a value identifying the row itself independently of where the row sits. Where that structure exists, every row can be reached two ways, and the two are not two spellings of one idea. - **Addressing by label** names a row by what it is called — "the record labelled `4471`". The label is carried by the row, like a serial number stamped on a part. - **Addressing by offset**, equivalently by position, names a row by how far down it currently sits — "the fourth row from the start, in whatever order the table is in right now". A label is an **identity**. An offset is a **coordinate**. Every behavioural difference between the two schemes follows from that one sentence. ## What each reference survives | what happens to the table | reference by row label | reference by offset | |---|---|---| | the rows are put in a different order | names the same record | names a different record, silently | | some rows are removed | names the same record, or fails to resolve if that record went | still resolves; names a different record | | a row is inserted above it | names the same record | names a different record | | the row labels are rewritten | may no longer resolve | completely unaffected | Read the table by column and the slogan falls out: **a label survives a reordering, an offset survives a relabelling.** Neither survives both. ## Why the offset failure is the dangerous one When a label reference fails, it usually fails audibly: the record is gone and there is nothing to hand back. Designs differ in how they say so — some raise, some hand back a row of absent values, some return an empty result — but none of them quietly gives you a different record. A stale offset does exactly that. As long as the number is inside the current row count the request succeeds and returns a row: a real row, of the right shape, with plausible values. Nothing downstream can tell it from the intended one. Three things routinely make an offset stale between the moment it is derived and the moment it is used: 1. A reordering step is added upstream, perhaps by someone else, perhaps for an unrelated report. 2. Rows are restricted away, so the survivors are renumbered from the start. 3. The source grows or shrinks between runs, so an offset chosen from last week's output lands somewhere else this week. ## The designs do not agree on how you ask This is the part that does not travel between ecosystems, so state it rather than assume it. - Some labelled designs expose **two separately named addressing surfaces**, one per scheme. The surface you call states which scheme you meant and nothing is inferred. - Some expose **one overloaded subscript** and infer the scheme from what is inside it. That is fine until the row labels are themselves integers, at which point the same expression can address a different row depending on data nobody looked at. - **An unlabelled rectangle of numbers has no labels at all**, so an offset is the only name a row has. Code written against it cannot say "the record called 4471" and must carry an identifier in a column of its own. - Some grammars discourage row labels entirely and expect you to restrict rows by a condition over a column you nominate as the identifier, so the label-against-offset question is settled by convention rather than by the tool. ## When an offset is the right answer "Always address by label" is too strong, and an interviewer will push on it. An offset is exactly right when the question really is about position in an order you just imposed yourself: - the first n rows after an ordering step performed in the same expression; - a sample, a chunk boundary, or the head of a file used to eyeball the data; - pairing two sequences that are matched by position and have no labels to match on. The common thread is that the offset is derived and consumed with no step in between that could reorder or restrict. **An offset is a short-lived name.** ## The habit to demonstrate Say which question the reference is asking. If the answer is "that specific record", the reference must be a label — or a value in a column you treat as the identifier — and it must be carried across step boundaries as such. If the answer is "wherever the tenth row happens to be right now", an offset is correct, and it should be computed at the point of use: never stored, never handed to another step, and never written into a report as though it named something durable.
- What happens to each scheme when rows are removed from the middle of the table?The surviving rows keep the labels they had, so the set of labels is simply non-contiguous and every surviving label still names its own record. Positions are recomputed from the start, so every offset after the removal point now names a different row than it did — and it still resolves, which is why the breakage is silent.
- A rectangle of numbers carries no row labels. How do you write a durable reference to one of its rows?You cannot do it with an address at all, because the only address is the offset and the offset is not durable. You carry an identifier as a value — a column of keys alongside the numbers, or a separate sequence of identifiers matched by position — and re-derive the offset from that identifier at the point of use.
A row label is the catalogue number printed inside a book; an offset is "third from the left on that shelf". Reshelve the section and the catalogue number still finds the book, while the position now finds someone else's.
saying these in an interview costs you the question
- Says the two schemes are interchangeable, just different syntax for one thing.
- Believes an offset reference is stable because the underlying data did not change.
- Thinks a label reference stops working once the table has been sorted.
- Assumes every rectangle of values carries row labels, so offsets are never needed.
- Counts the first row as row one in every host without checking the convention.