skip to content

Two columns of 1,000 values each, taken from different tables, are added together — what decides which value pairs with which?

level: juniorimportance: must knowfreq 58%

answer

  1. the line does not choose
  2. property of the tool, not the code
  3. labels first, position only without them
  4. an unmatched label leaves a hole

basics

~10 s

The tool decides, not the expression. Designs that carry a per-row identifier pair the two operands by that identifier first; designs that carry none pair strictly by position and refuse operands of different lengths.

solid answer

~50 s

Nothing in the line you wrote chooses the pairing rule — the tool holding the data does. Some designs attach **row labels** (a per-row identifier carried beside the columns) and pair the two operands by label before computing anything, so the value at position 3 on one side may be added to the value at position 900 on the other, and a label present on only one side comes back as a cell holding the absent-value marker, the tool's representation of a value that is not there. Other designs — a bare rectangle of numbers addressed by position only — have no labels at all, pair strictly by position, and refuse two operands of different lengths. The same expression therefore means two different things. Before trusting the result, say what makes a row on one side correspond to a row on the other, and make the operation pair on that.

go deeper

for a junior

Recall that two columns can be paired in two different ways — by a per-row identifier or by position — and that which one happens depends on what the data is held in, not on how you wrote the line.

for a middle

Explain what a per-row identifier does to an arithmetic operation: the pairing happens before the arithmetic, a label on one side only leaves a hole, and identifiers that look like counting numbers stop matching positions the moment one side is filtered.

for a senior

Show that you would state the intended correspondence before writing the expression, force it explicitly, and check the output row count against a number you predicted rather than reading the first few rows on screen.

for a principal

Weigh a codebase where every combination inherits an implicit pairing default against one where the correspondence is declared at every seam, and be able to say which of the two failures you would rather your team debug under time pressure.

## The question underneath the expression Two columns are combined with an ordinary arithmetic expression. Before a single addition happens, the tool has to decide **which value on one side is paired with which value on the other**. That decision is made by the library, not by the line you wrote, and the two rules in common use produce different results from identical source text. Four terms, defined once: - **Row labels** — the per-row identifier a tool carries beside the columns. Present in some designs, absent entirely in others. - **Labelled table** — a rectangle of named columns that, in some tools, also carries one of those identifiers for each row. - **Bare numeric array** — a rectangle of numbers addressed by position only, carrying no labels at all. - **Label pairing** — the tool silently lining two operands up by their row labels before it computes anything. Its opposite is **positional pairing**: the first value with the first, the second with the second. ## The two rules, side by side | What the operands carry | How they are paired | What a mismatch does | |---|---|---| | Row labels on both sides | By label, before any arithmetic runs | A label on one side only yields a cell holding the **absent-value marker** — the tool's representation of a value that is not there | | No label concept at all | Strictly by position | Two operands of different lengths are refused at the call | | Labels that happen to be consecutive counting numbers | Still by label — the numbers are identities, not positions | Two sides counted from different starting points overlap only partly | The third row is the one that fools people. A label set reading 0, 1, 2, 3 looks exactly like a set of positions, and for as long as both sides were built the same way it behaves like one. Filter one side, and the rows that survive keep the identifiers they already had rather than closing up; from then on the two sides agree by label only where the filter happened to keep a row. ## Why this is a property of the tool - The expression is identical in both worlds. Nothing in it says *pair by label* or *pair by position*. - Whether the operands carry labels at all is a design decision taken by whoever wrote the library, and the designs in wide use genuinely differ on it. - In one process you can hold both kinds at once: a labelled table beside a bare rectangle of numbers. The same arithmetic between two of the first behaves differently from the same arithmetic between two of the second. - Nothing raises. Label pairing that finds few common labels returns a full-length result of absent values; positional pairing over equal-length operands that do not actually correspond returns plausible wrong numbers. ## The three failure shapes 1. **Silent reordering.** Both sides carry the same labels in different orders. Positional pairing would be wrong here and label pairing is right — which is why someone who assumed position is surprised by a *correct* answer and learns the wrong lesson from it. 2. **Holes.** The two label sets overlap only partly, so rows whose label exists on one side only come back carrying the absent-value marker. 3. **Growth.** A label occurring more than once on both sides is paired with each of its partners in turn, so the result is longer than either input. ## Making the correspondence something you wrote down The durable habit is to stop relying on the default and state what the correspondence is: 1. Say out loud what makes row 400 on one side the same thing as some row on the other. If the honest answer is *they came out of the same ordered step over the same rows*, the correspondence is position. If it is *they are the same account*, the correspondence is a value. 2. When it is position, **drop to positions** deliberately — discard the labels so the operation pairs strictly by position — and assert that the two lengths are equal in the same breath. 3. When it is a value, put that value in a column on both sides and match the two tables on it. The correspondence is then written in the code and survives anyone filtering or reordering either side later. 4. Either way, compare the output row count against a number you predicted before you look at any value. ## What varies, and how to say it There is no single true sentence about pairing for this whole family of tools. Some attach an identifier to every row and treat it as the identity of that row, so any operation between two operands pairs by label first and the unmatched labels become holes. Others have no label concept, pair strictly by position, and treat a length mismatch as an error. A candidate who says flatly *it pairs by label* is describing one design as though it were the class. The answer that travels across ecosystems is: it depends what the operands carry, here is how each behaves, and here is how I would make it explicit rather than inherit a default.

  • If both operands carry the same labels in the same order, does label pairing give the same answer as positional pairing?
    Usually yes: identical label sets in identical order pair the same way under either rule. Two caveats matter. The equality is a property of the data at run time rather than of the code, so a filter or a reorder upstream breaks it without touching your line. And if a label occurs more than once on both sides, label pairing multiplies the rows even when the two sides look identical.
  • How do you make the correspondence explicit instead of relying on the default?
    Two routes. If position genuinely is the correspondence, discard the labels so the operation pairs strictly by position, and assert the two lengths are equal in the same step. If the correspondence is a value both sides carry, put it in a column on each side and match the two tables on it — that keeps working after either side is filtered or reordered.

Two decks of numbered cards. Pairing by label is dealing so that card 7 always meets card 7 wherever it sits in each deck, and a number present in only one deck comes back with no partner. Pairing by position is dealing off both tops at once, which is right only if someone stacked the two decks in the same order and can say who.

saying these in an interview costs you the question

  • Says element-wise arithmetic always pairs the two operands by position.
  • Says every tabular design pairs on row labels before computing.
  • Assumes equal lengths guarantee that the two sides correspond.
  • Expects a wrong pairing to raise rather than return plausible values.
  • Treats row labels as display decoration with no effect on results.