A transform subtracting a shorter operand from a longer one is moved to a different tool, still runs, but returns different numbers. What could the second tool have done?
answer
- the expression did not change
- three rules, three silent answers
- labels or bare positions decides everything
- two of the three never raise
- assert positions and absent count
basics
~20 sIt resolved the size conflict by a different rule. The same written expression can return a rectangle of every pairing, a result filled by repeating the shorter operand, or a result over both sets of row labels that is mostly absent - and only one of those three designs refuses outright.
solid answer
~50 sThe expression did not change; the rule underneath it did. Three resolutions exist in this family. A rule that lines axis lengths up from the last axis backwards and stretches any length-one axis can turn two short operands into a rectangle of every pairing. A rule that repeats the shorter operand over the longer gives a full-length result, warning rather than stopping when the lengths do not divide evenly. And operands that carry row labels are not stretched at all - they are matched by those labels, so the result spans both label sets and carries the absent-value marker wherever a label appeared on one side only. Two of the three produce output without an error, which is why "the port ran" is not an acceptance test. Before moving such a transform, establish whether each operand carries row labels or is a bare positional buffer, and assert the result's number of positions on both sides of the move.
go deeper
Recall that the same subtraction can mean different things in different tools when the two operands are different sizes, and that running successfully is not proof of a correct result.
Name all three resolutions and what each returns: a rectangle of pairings, a repeated fill, or a result over both label sets that is mostly absent. Say which of them raises and which do not.
Diagnose from the fingerprint - inflated counts, a nonsense tail, or a jump in absent values - and name the property that decides the rule: whether each operand carries row labels or is a bare positional buffer. Then give the assertions you would leave behind.
The angle is what a team commits to when it runs the same transform on two tools: which invariants are asserted at every step boundary, and whether a dual-run comparison with a stated tolerance is worth its cost for the duration of a migration.
## Why the same expression means three things The written form - subtract one operand from another - is identical everywhere. What differs is the rule each design applies when the two operands do not hold the same number of values. There are three resolutions in play across this family, and they disagree not about speed or ergonomics but about **what number comes back**. | the rule the design applies | what it returns for a size conflict | does it stop? | |---|---|---| | Line the axis lengths up from the last axis backwards; stretch any axis of length one | The larger length on every stretched axis - which, when both sides have a length-one axis, is a rectangle of every pairing | Only when no length-one axis is available | | Repeat the shorter operand over the longer | A full-length result; where the longer length is not a whole multiple of the shorter, usually a partial fill | Rarely - typically a warning | | Do not stretch; match the operands by their row labels | A result spanning both sets of labels, with the absent-value marker at every position present on one side only | No | The **absent-value marker** is whatever a tool stores where there is no value - for some designs a borrowed floating-point sentinel, for others a separate validity bit beside a value of the narrow width. ## What each failure looks like from the outside Each resolution has a distinct fingerprint, and recognising it is most of the diagnosis: - **Too many rows, all of them plausible numbers.** A stretch of two length-one axes multiplied the result. Downstream aggregates come out inflated by a whole factor, and no value is individually wrong. - **The right number of rows, and the tail is nonsense.** The shorter operand was laid down repeatedly. The first stretch of the result is often correct, which is exactly why a preview of the head reassures. - **The right number of rows, and most of them absent.** Row labels were matched rather than positions. Sums come back smaller than expected, or absent entirely, depending on how the next step treats absence. ## The property to establish before the move The single fact that decides which rule applies is **whether each operand carries row labels or is a bare positional buffer**. An operand carrying row labels knows an identity for each of its values and can be matched on that identity. A bare positional buffer carries no labels at all and can only be matched by position, which is why a size conflict there is resolved by stretching or refusing rather than by matching. So the pre-move question is not "does the second tool support this operator" - all of them do. It is: do both operands still carry labels after the port, and if they do, are those labels the same labels? A transform that worked by position for years can start matching by label the first time an operand arrives from a step that attached labels to it, without a single line of the expression changing. ## Why "it ran" is a weak acceptance test Two of the three resolutions do not raise. One of them only warns, and a warning in a scheduled run goes to a stream nobody reads. So a successful port tells you the operator exists, nothing more. The three acceptance signals that actually mean something are: 1. **The result's number of positions** matches the row count the step is contracted to produce. 2. **The count of absent values in the result** is what it was before the move - a jump means matching by label replaced matching by position. 3. **A total or a checksum over the result** agrees with the old implementation to a stated tolerance rather than exactly, because a reorganised accumulation over floating-point values differs in its last digits. ## Making it hard to regress - Assert the result's number of positions immediately after the expression, not at the end of the pipeline, so the assertion names the step that broke. - Assert the absent-value count too, where the operands may carry labels; it is the only cheap signal that distinguishes label matching from positional matching. - Keep the size expectation in the code rather than in a comment or a document. The rule that the code runs under is a property of the runtime, and the assertion is the only part of your intent that travels with the transform. - Run both implementations over the same input during the move and compare the outputs, rather than reasoning about which rule ought to apply. The deeper point for a senior candidate: a size mismatch is a defect in **your operands**, not in the tool. Every one of the three rules is a defensible design. What is not defensible is writing an expression whose correctness depends on which of them happens to be underneath it.
- What single property of the two operands do you establish before porting such an expression?Whether each one carries row labels or is a bare positional buffer. Label-carrying operands are matched on identity and yield absence where labels do not meet; bare buffers can only be matched by position, so a conflict is stretched or refused instead. That one fact selects which of the three rules the expression will run under.
- The ported result has the expected row count but far more absent values. What does that point to?Matching by row label rather than by position. The positions are there because the result spans both label sets, and the absences mark every label that appeared on only one side. It means the two operands' labels disagree - which a positional match had been quietly papering over before the move.
- Why compare the old and new outputs with a tolerance rather than for equality?Because a differently organised accumulation over floating-point values gives a slightly different total. A pass that adds in blocks and a pass that adds strictly in order disagree in the last digits, so an exact comparison fails on a correct port. Over exact representations, such as integers in range, equality is the right test.
saying these in an interview costs you the question
- Assumes the same expression means the same thing in every tool
- Expects a size mismatch to surface as an exception everywhere
- Says a mostly-absent result proves the input data was bad
- Accepts a port because it ran without an error
- Thinks matching by label and by position differ only in speed