skip to content

Your team's tables move between a design that carries row identity and one that does not. What standing rule do you set for where identity lives?

level: principalimportance: nice to knowfreq 40%

answer

  1. a standing commitment, not a per-script call
  2. column at the boundaries
  3. labelling as a local optimisation
  4. portability against per-lookup speed
  5. write requirements a reviewer can apply

basics

~20 s

A defensible default: identity lives in an ordinary column, and is promoted into the row labels only inside a step that performs many lookups by name, then demoted before the data is handed on. Write the rule around what you require, not what one tool offers.

solid answer

~50 s

The decision is a standing commitment, not a per-script choice, because it decides how portable your transforms are. The rule I would defend is that identity is an ordinary column everywhere data crosses a boundary between steps, tools or people, and that a labelling is a local optimisation taken inside a step that performs many retrievals by name or many span selections, then given back. The argument for it: an identifier column exists in every design, including the ones with no row-identity concept, so a transform written that way survives a change of tool; it is visible in every output a human reads; and nothing about it depends on an unenforced property. The argument against it is real too — where one design is in play and lookups dominate, a labelling carried throughout is faster and reads better. State the rule so a reviewer can apply it without benchmarking.

go deeper

for a junior

Notice that identity can live in a field or in the holding's row labels, and that only the field exists in every design you might meet.

for a middle

Be able to compare the two placements on cost and on retrieval, and to say what each one assumes about the data being unique.

for a senior

Argue the choice from measured access patterns and from the blast radius of an unenforced assumption spreading across pipeline stages.

for a principal

Own the standing rule: write it as requirements a reviewer can apply, name the exception explicitly, and make sure it survives a change of tool.

## What is actually being decided Every table your team passes around answers one question, usually by accident: **where does the entity's identity live?** There are two honest answers. - **In the row labels** — the per-row keys a holding stores beside the values, in designs that have the feature. Named retrieval is direct, the identity is not one of the fields, and a design can build a lookup structure over it. - **In an ordinary column** — a field like any other. Every design has fields, so this works everywhere; retrieval means testing the column, which is a pass over the values. Because the choice repeats hundreds of times across a codebase and is expensive to change later, it is a standing rule rather than a per-script judgment, and that is exactly why an interviewer asks it of a lead. ## The default I would set, and why > **Identity is a column at every boundary. A labelling is a local optimisation inside one step, and it is given back before the data leaves.** The reasons, in the order I would argue them: 1. **Portability.** Some tabular designs carry no row-identity concept at all, and others have row names that most operations quietly discard. A transform written against a column is expressible in all of them; a transform written against labels is expressible in some. 2. **Visibility.** A column appears in everything a human looks at. Identity that lives outside the fields is easy to lose track of in review and easy to drop in a step that was not thinking about it. 3. **No dependence on an unenforced property.** Labels are never validated. A labelling carried across a whole pipeline means every stage inherits an assumption that nothing anywhere checks, and the failure — a retrieval returning several rows where one was expected — surfaces far from where the duplicate entered. 4. **Predictable cost.** Testing a column is a pass, on every run, in every order. Label-based access is faster when its preconditions hold and degrades quietly when they stop holding, which in a scheduled job trades a known cost for a conditional one. ## The honest case against it A rule you cannot argue against is a rule you have not thought about. | when the opposite rule wins | why | |---|---| | one ecosystem, no plans to move | portability buys nothing, and the labelling reads better | | many retrievals by name over one holding | the full-length allocation amortises and the per-lookup saving is real | | repeated span selections over ordered identifiers | bounded search over ordered labels is a genuinely different complexity class from a pass | | interactive exploration | naming a row is more ergonomic than testing a field, and the cost of a mistake is a rerun | And the promote-then-demote discipline has its own price: the conversions are code, they can be forgotten, and a step that promotes and never demotes leaves the next stage holding an assumption it did not ask for. ## Write the rule in terms of requirements, not tools The principal-grade part of this answer is not which rule you pick — it is that the rule survives a sentence being true in one design and false in another. Phrase it as requirements the team can check in review: - *Identity must be readable in any output a human inspects.* - *A transform must be expressible in a holding with no row identity.* - *No stage may depend on an identity being unique unless that stage itself establishes it.* - *An optimisation that depends on the data arriving in a particular order must be visible at the point it is taken.* A rule written that way keeps working when the team adopts a different tool; a rule written as "always set the labelling first" is a statement about one design wearing the clothes of a policy. ## What makes it a rule rather than a preference Three properties: it is decidable by a reviewer without measurement, it names the exception explicitly so the exception does not have to be smuggled, and it is cheap to follow in the common case. If a rule needs a benchmark to apply, it will be ignored; if it forbids the fast path outright, it will be broken in the one place it mattered and nobody will say so. ## How to present it in an interview Give the rule in one sentence, give the two strongest arguments for it, then give the case where you would overturn it and say what evidence would make you. Finish on the thing that generalises: a commitment about identity is a commitment about which sentences your codebase is allowed to assume, and the sentences in this area are true of one design and false of another more often than anywhere else in this subject.

  • What evidence would make you overturn the default and carry a labelling throughout?
    Measured access patterns: a pipeline dominated by retrievals by name or by span selections over ordered identifiers, running on one design with no near-term plan to move. I would also want the duplicate risk understood at the point data enters, because carrying a labelling spreads an unenforced assumption across every later stage.
  • What is the cost of the promote-then-demote discipline itself?
    Conversions are code that can be skipped or forgotten, and the promotion allocates. A step that promotes and does not demote hands the next stage a holding shaped differently from the contract, which is the failure the rule was written to prevent. The discipline is only worth it if the boundary is a real one — a module, a file handed on, a team edge.
  • Why is 'always set the labelling first' a weak rule even in a single-tool team?
    It states a mechanism instead of a requirement, so it gives a reviewer nothing to judge and gives the team no way to see when it stopped applying. Rules phrased as requirements — identity must be visible in output, no stage may assume unchecked uniqueness — still mean something when the tool changes.
  • Does this decision affect how much you can trust a row count?
    Yes, indirectly. Identity carried as an unenforced label invites stages to assume one row per entity without saying so, and the first duplicate makes several counts quietly wrong at once. Identity as a field does not remove duplicates, but it keeps the multiplicity in view where a reader can see it.

saying these in an interview costs you the question

  • Picks a rule without naming the case where it should be overturned
  • Assumes every design in play carries row identity
  • Treats named retrieval as universally faster regardless of access pattern
  • Believes a labelling makes the identifier unique
  • States the rule as one tool's mechanism rather than as a requirement