In a tool that pairs operands by row label, one side carries label A twice and the other three times — what comes back?
answer
- not a lookup, a pairing
- count occurrences on each side
- the product, not the larger count
- two times three is six rows
basics
~20 sSix rows come back for that label. Every occurrence on one side is paired with every occurrence on the other, so the output row count for a label is the product of its two occurrence counts, not the larger of them.
solid answer
~50 sLabel pairing — the tool lining two operands up by their **row labels**, the per-row identifier it carries beside the columns — is not a lookup. It pairs every matching occurrence with every other, so two occurrences of A on one side and three on the other give 2 x 3 = 6 output rows for A, and the whole result is that product summed over every label. The row arithmetic is the same wherever label pairing happens; what differs between tools is whether you are told, which ranges from complete silence, through a warning, to a refused call. The usual source of an unintended repeat is a step that put two pieces one under the other: in a design that carries row labels each piece brings its own along, so the same label is now present twice for reasons that have nothing to do with the data.
code
pseudocode · 11 linesfirst = [ (A, 10), (A, 20), (B, 5) ] # each row is (row label, value)
second = [ (A, 1), (A, 2), (A, 3), (B, 7) ]
for label in labels present on both sides:
for l in rows of first with that label: # A: 2 occurrences
for r in rows of second with that label: # A: 3 occurrences
emit (label, l.value + r.value)
# label A -> 2 * 3 = 6 output rows
# label B -> 1 * 1 = 1 output row
# total -> 7 rows, from inputs of 3 rows and 4 rowsgo deeper
Recall that a pairing on row labels can return more rows than either side started with, and that this happens when the same label is present more than once on both sides.
Explain the arithmetic precisely: for one label the output row count is its occurrence count on one side multiplied by its occurrence count on the other, and the total is that product summed over every label.
Predict the output row count before running the step, compare it and the distinct-label count afterwards, and trace a repeat back to the step that created it rather than trimming rows at the end.
Decide whether pipelines under your care may rely on implicit label pairing at all, given that the same inflation is silent in one tool, a warning in another and a refused call in a third.
## Pairing is not lookup The mental model that produces the wrong answer here is *lookup*: for each row on one side, go and find its partner on the other. That model has a single partner built into it, so it predicts a result no longer than the side you started from. What actually happens is a **pairing over occurrences**. **Row labels** are the per-row identifier a tool carries beside the columns — present in some designs, absent entirely in others. **Label pairing** is the tool lining two operands up on that identifier before it computes anything. When a label is present more than once on a side, it has more than one occurrence to offer, and the pairing emits every combination. ## The arithmetic For one label value, the number of output rows is the number of its occurrences on one side **multiplied by** the number on the other. The total output is that product summed over every label. | Occurrences on one side | Occurrences on the other | Output rows for that label | |---|---|---| | 1 | 1 | 1 — the case everyone has in mind | | 1 | 4 | 4 | | 2 | 3 | 6 | | 3 | 0 | 3 rows carrying the absent-value marker, in designs that keep a label present on one side only | The **absent-value marker** is the tool's representation of a value that is not there. Note that the last row of the table is a different defect from the first three: holes rather than growth. They frequently arrive in the same result, and describing one does not describe the other. Two consequences worth stating explicitly: - The result can be longer than either input, which no single-partner model predicts. - The number of **distinct labels** does not change while the **output row count** grows. That pair of numbers is the whole diagnostic; either one alone looks plausible. ## Where the repeat came from An unintended repeated label almost never comes from the data. It comes from a pipeline step: - **Putting two pieces one under the other.** In a design that carries row labels, each piece brings its own identifiers with it, so the result carries both sets and every identifier that both pieces used is now present twice. In a design that has positions only, there is nothing to duplicate — the result is simply longer, and the question does not arise. - **Adopting a non-unique attribute as the identifier.** A per-day identifier over data with several records per day repeats by construction. - **Re-running a step that produced part of the data twice**, so identical identifiers arrive from two runs. ## Why nothing necessarily tells you The row arithmetic is universal wherever label pairing happens. The **diagnostics are not**. Some designs return the inflated result in silence; some emit a warning; some accept a declaration of the relationship you expect and fail the call when the data violates it. Any answer that promises you will be warned is describing one design as though it were the class, and a candidate who relies on the warning will not see the defect in a tool that does not emit one. ## Catching it 1. **Predict the output row count before you run the step.** If you cannot state the number you expect, you are not in a position to notice a wrong one. 2. **Compare the output row count with the input counts afterwards**, alongside the count of distinct labels. Growth with an unchanged distinct-label count is a pairing product. 3. **Count occurrences per label on each side** before the operation when the answer matters, and deal with the repeat where it was created rather than trimming rows off the end afterwards. 4. **Remove the ambiguity instead of surviving it.** If position is the real correspondence, discard the labels and pair strictly by position after asserting equal lengths. If the correspondence is a value both sides hold, put it in a column and match the two tables on it, where the behaviour is at least something you chose. ## The thing people say that is wrong *It takes the longer of the two counts.* It does not; three occurrences against two is six rows, not three, and the gap between the two predictions widens fast. Two labels at 3 x 3 contribute eighteen rows from twelve input rows. Repeat that across a real dataset and a step that was supposed to preserve size returns several times what went in, with every value in it individually correct — which is precisely why it survives review.
- A label occurs twice on one side and not at all on the other. What comes back for it?Two rows in designs that keep a label present on one side only, each carrying the absent-value marker where the other operand should have been. Nothing multiplies, because there is nothing to multiply against. This is the other failure shape — holes rather than growth — and the two often appear in the same result, so finding one is no reason to stop looking.
- The output row count grew but the number of distinct labels did not. What does that tell you?That the growth is a pairing product rather than new data. No label was added, so nothing arrived; the row count rose because at least one label carries several occurrences on each side and every combination was emitted. Report those two numbers together as a habit — either one on its own looks entirely plausible.
- Does discarding the labels before the operation fix this?It removes the multiplication, because positional pairing emits one row per position and cannot inflate. It does not fix the underlying problem: the repeat is usually the trace of an upstream step that produced the same rows twice, and pairing those rows by position simply lines up the wrong ones silently. Fix the step that created the repeat.
saying these in an interview costs you the question
- Says the result takes the length of the longer operand.
- Thinks repeated labels are collapsed before the pairing runs.
- Believes a tool always warns before returning an inflated result.
- Assumes row labels cannot repeat, so the case cannot arise.
- Reads the first rows on screen instead of the output row count.