skip to content

A table with integer row labels is subscripted with 3 — what decides whether that names a label or an offset?

level: middleimportance: should knowfreq 58%

answer

  1. the integer describes nothing by itself
  2. something outside it picks the scheme
  3. labels win when the labels are integers
  4. the meaning depends on data, not on the line

basics

~20 s

The design does. Where two separately named addressing surfaces exist, the surface you called decides and nothing is guessed. Where a single subscript is overloaded, an integer resolves as a label once the row labels are integers, and as an offset otherwise.

solid answer

~50 s

The integer itself carries no information about the scheme, so something else has to resolve it. Under two separately named addressing surfaces there is no ambiguity: the surface names the scheme, and `3` means a label on one and an offset on the other. Under a single overloaded subscript over a labelled table, the design inspects the row labels: if they are integers, `3` is read as the label 3; if they are not, it falls back to reading it as a position. That makes the meaning of the expression depend on data the author may never have looked at — and if the labels were rewritten to integers upstream, the same line starts addressing a different row without changing. A third wrinkle: in some subscripts a bare integer addresses a column rather than a row, so the axis is inferred too.

go deeper

for a junior

Recall that an integer inside a row reference does not say which scheme is meant, and that a table can have integer row labels which the reference may match instead of counting to a position.

for a middle

Explain the resolution rule: two named surfaces leave nothing to infer, while one overloaded subscript prefers the label reading when the labels are integers and falls back to a position otherwise.

for a senior

Describe how the bug arrives without the line changing — a relabelling or a removal upstream flips the reading — and why an inline fixture with consecutive integer labels cannot detect it.

for a principal

Argue the general rule: an overloaded address asks the data what the author meant, so the codebase standard is that the scheme is explicit at the call site and bare integers never cross a step boundary.

## An expression with two readings `3` inside a row reference is not self-describing. It could be the row *called* 3 or the row *at position* 3, and on a table whose row labels are integers both readings resolve to a real row — usually a different one. The question is what resolves the ambiguity, and the honest answer is that it depends on the design, which is exactly why it is asked. ## What resolves the integer | design | what a bare integer in a row reference means | |---|---| | no row labels at all | an offset — there is nothing else it could be | | two separately named addressing surfaces | whichever scheme the surface you called names; nothing is inferred | | one overloaded subscript, non-integer row labels | an offset, by fallback | | one overloaded subscript, integer row labels | a **label**, because the label reading is preferred when it is available | | a subscript whose first axis is the columns | a column reference, not a row reference at all | Two things are worth pulling out of that table. First, the overloaded subscript does not raise on the ambiguity — it resolves it silently, and a resolution you did not intend produces a row rather than an error. Second, the rule is **data-dependent**: the same source line means one thing on a table whose labels are text and another on a table whose labels are integers. ## Why that makes it a genuinely nasty class of bug The expression never changes; the data underneath it does. Three realistic routes into it: 1. **A relabelling upstream.** A step that renumbers the row labels, or one that rebuilds the table from scratch so the labels come out as consecutive integers, converts every overloaded integer subscript downstream from an offset reference into a label reference in one commit that never touched them. 2. **A restriction upstream.** Rows are removed. If the surviving labels are kept, they are integers with gaps — so the integer subscript still resolves as a label, and the label reading and the offset reading now diverge by however many rows were dropped above it. 3. **A test fixture that does not match production.** A small fixture built inline often ends up with consecutive integer labels starting at zero, where the two readings coincide exactly. Every test passes. Production data, whose labels are real identifiers, takes the other branch. In all three the symptom is the same and it is the worst symptom available: a plausible row, of the right shape, from the wrong record. ## Removing the ambiguity The fix is not cleverness, it is refusing to let the expression be ambiguous in the first place. 1. **Use the explicit surface** where the design offers one. If the scheme is in the name of the thing you called, there is no inference and no data dependency. This is the whole reason those surfaces exist. 2. **Where only one overloaded subscript exists, do not put a bare integer in it.** Address by a label you can see is a label — text identifiers, timestamps — or derive the position and pass it through the offset-only path the design provides. 3. **Do not let a reference cross a step boundary as a bare integer.** Pass the identifier; let the consuming step resolve it. 4. **Make the fixture look like production.** If real labels are identifiers, the test table's labels must be identifiers too, or the branch that runs in production is never exercised. ## What the interviewer is listening for A weak answer picks one reading and states it as the rule: "an integer in a subscript is a position". That is true of unlabelled rectangles and of a dedicated offset surface and false of an overloaded subscript over an integer-labelled table, which is precisely the case the question is about. A strong answer names what resolves the integer, says that the resolution depends on the labels rather than on the expression, and points out that the safe designs are the ones that never have to guess. There is a wider principle underneath, and it is worth saying out loud: **an overloaded address is a design that asks the data what you meant.** Whenever a single surface serves two schemes, some input decides for you, and the decision is invisible at the call site. The remedy is always the same shape — make the scheme explicit in what you wrote, so no data can change the meaning of the line.

  • Why do tests so often miss this, even with good coverage?
    Because an inline fixture usually ends up with consecutive integer labels starting at zero, where the label reading and the offset reading name the same row. Both branches agree, so the test cannot distinguish them. Production labels are real identifiers, take the other branch, and the difference appears for the first time in output nobody diffs.
  • A design offers only one overloaded subscript. How do you write an unambiguous row reference?
    Keep bare integers out of it. Address by a label whose type cannot be mistaken for a position — an identifier string, a timestamp — or route positional work through whatever offset-only path the design provides. If neither is available, resolve the identifier to a position immediately before use, in the same expression, so nothing can intervene.

saying these in an interview costs you the question

  • States flatly that an integer in a subscript is always a position.
  • Assumes the design raises rather than silently choosing a reading.
  • Thinks integer row labels are rare enough to ignore in practice.
  • Believes a passing test proves which reading production takes.
  • Calls it a syntax question rather than a data-dependent resolution.