skip to content

A result is carried through five filtering and reordering steps and must be reattached to the records it came from - what must travel with the values?

level: seniorimportance: should knowfreq 46%

answer

  1. position is not identity
  2. what survives a reorder
  3. silent, full-length, wrong rows
  4. the carrier differs by design
  5. carry a field when identity is not first-class

basics

~20 s

An identity per record - either a row-label structure the design carries for you or an ordinary field you keep and match on. Position is not identity: the first filter or sort breaks correspondence between two holdings, and nothing reports it.

solid answer

~50 s

Between two holdings that carry no identity, the only correspondence is position: same length, same order. Every filter, sort, grouping or de-duplication along the way breaks that, and the breakage is silent - you get a full-length result whose values are attached to the wrong records. What survives those steps is an identity attached to each row, and designs differ in where it lives: some materialise a row-label structure beside the values so it is carried for you; some tabular designs have no row-identity concept at all, so identity has to be an ordinary field you carry deliberately and match on; some keep row names but drop them across most operations. The working rule is to know which of those you are in, and to carry an explicit identifying field wherever identity is not first-class. Nothing checks that identity for uniqueness either.

go deeper

for a junior

Take away one rule: rows that line up today stop lining up as soon as something filters or sorts them, so keep whatever identifies a record with the values rather than relying on order.

for a middle

Explain the mechanism: position is a property of the current arrangement, identity is a property of the record, and only the second survives a step that changes which rows are present or in what order.

for a senior

Demonstrate that you know the failure is silent and full-length, that agreeing lengths are not evidence, and that you have a habit - carry an identifying field where the design does not carry identity for you.

for a principal

The call worth owning is a standing convention: what identifies a record in this codebase, who is responsible for carrying it, and whether the team accepts positional correspondence anywhere at all.

## Position is a correspondence, not an identity When two holdings have no per-row identity, the only thing relating row 7 of one to row 7 of the other is that both are seventh. That is a genuine correspondence and it is perfectly usable - as long as nothing has disturbed either side. It is not an identity, because it is a property of the holding's current arrangement rather than of the record. The distinction only matters once a pipeline has steps in it, which is to say always. ## What breaks it, and how quietly A correspondence by position is broken by any step that changes which rows are present or what order they are in: - **filtering** - one side now has fewer rows, and every position after the first dropped row refers to a different record; - **sorting** - both sides may still be the same length, which is the dangerous case, because the lengths agreeing looks like evidence; - **grouping** - the result has one row per group, and positions now index groups rather than records; - **de-duplication** - rows vanish somewhere in the middle, and nothing marks where; - **reading the two holdings from different sources**, which may simply never have agreed on order in the first place. In none of these does anything raise an error. If the two sides are the same length, the operation succeeds and produces a full-length result with values attached to the wrong records. If they are different lengths, you get a failure - which is the lucky outcome, because it is loud. ## Where identity lives: three designs, three obligations 1. **The design materialises a row-label structure beside the values.** Identity is first-class and carried by default: filter the rows and the survivors keep their labels, so a later step can bring two holdings together by identity. Your obligation is to know when an operation replaces or resets those labels, and to remember that nothing checks them for uniqueness. 2. **The design has no row-identity concept at all.** Rows are addressed by position, full stop. Identity has to be an ordinary field that you select through every step and match on explicitly at the end. Your obligation is that the field survives: a step that keeps only the measures will drop it, and it will do so without comment. 3. **The design keeps row names but discourages them** - restricted, and dropped by most operations. The obligation is the most awkward of the three, because the mechanism exists and cannot be relied on, so in practice you treat it like the second case. None of the three is wrong. What is wrong is writing code that assumes the first while running on the second. ## The two halves of what labels buy, seen here together | what you need at the end | what has to have travelled | what fails if it did not | |---|---|---| | attach a value to the right record | an identity per row, carried or explicit | full-length result, wrong rows, no error | | say what each number means | field names on the result | a rectangle nobody can check by reading | | prove the two sides still describe the same records | identity present on both sides | no evidence exists in the values either way | That last row is the one people miss. Without an identity on both sides there is nothing in the data that could confirm or deny the correspondence - not the length, not the values. The absence of evidence is the whole hazard. ## The interpretability half The same pipeline has a second failure that is slower and just as expensive. A result carried positionally moves the meaning of every offset out of the data and into whoever wrote the code. Five steps later nobody can check the result by reading it; they can only re-derive it. Names keep the mapping attached to the values, where a reviewer and the next step both see the same thing. Neither names nor identity enforce anything, but both keep meaning and correspondence from drifting away from the numbers they belong to. ## The working rule - Decide at the start of the pipeline what identifies a record, and say it out loud in the code. - If the design does not carry identity for you, carry it as an ordinary field and select it through every step. - Never treat equal length as evidence of correspondence; it is the failure mode's best disguise. - Remember that whichever carrier you use, nothing checks it for uniqueness - reattaching by identity assumes something the holding does not guarantee.

  • If nothing raises an error, what would tell you the correspondence had broken?
    Only something outside the values. Agreeing lengths prove nothing: a step that reordered rows leaves the same count, and so does one that replaced them. Where neither side carries an identity there is no evidence in the data either way, which is exactly why correspondence by position is a production hazard rather than a style preference.
  • Is carrying an identifying field yourself worse than a design that carries identity for you?
    It is more explicit and more work. A carried field has to be selected through every step and is dropped by any step that keeps only the measures; a first-class identity travels by default but is unenforced and can be replaced by an operation. Neither removes the need to state what a record is.
  • Why is interpretability part of the same argument rather than a separate one?
    Because both failures come from meaning living outside the data. A positional result moves what each offset means into the code and into the person who wrote it, so after five steps nobody can check by reading - only by re-deriving. Names and identity both keep meaning attached to the values.

saying these in an interview costs you the question

  • Assumes two results of the same length line up row for row
  • Says every table carries an identity per row automatically
  • Expects a mismatch to raise an error rather than produce wrong rows
  • Relies on the original read order surviving a sort or a filter
  • Treats a row label as guaranteed unique because it is called a key
  • Thinks field names alone are enough to reattach rows to records