You widen a table so two measures sit side by side, then hand those two columns to a numeric routine as a plain rectangle of numbers. What must you establish at that boundary?
answer
- the object changed, not just the arrangement
- label matching against position matching
- the ordering stops being checked for you
- invented cells and data-derived column positions
basics
~20 sWhat lines the two operands up. A labelled table matches them on their row labels; a plain rectangle of numbers matches strictly by position, so crossing that boundary stops the ordering being checked for you and a wrong answer arrives with nothing raised.
solid answer
~50 sThe boundary is not just a change of arrangement, it is a change of object, and the two objects resolve a mismatch differently. A **labelled table** carries row labels and aligns its operands on them, so two columns taken from differently ordered sources still meet the right rows. A **plain rectangle of numbers** — a block of one type addressed by position, with no labels to align on — matches first with first. Hand it two columns that came from differently ordered sources and it produces a wrong answer with no marker anywhere. So establish three things before you cross: that the rows are in a known, identical order, or that you align while you still can; what happened to the cells the widening had to create because a subject never had that measure, since a rectangle has no notion of them; and whether the column positions you are about to index are stable, given that a widening takes its headers from the data.
go deeper
Know that there are two kinds of object here: one carries row labels and one is addressed purely by position. Which one you are in decides what happens when two operands disagree about order.
Explain both matching rules and what each does with a row present on one side only — holes in one world, silently mispaired numbers in the other.
Treat the crossing as the risky line. Take both operands out of the same rearranged table, select by name, and decide the disposition of the cells the widening invented before any routine averages over them.
Make it a rule that a layout or alignment claim never crosses between object models without the mechanism being named, and put the crossing in one reviewed place rather than wherever a step happened to need it.
## Two objects, two ways of lining operands up The widening put the two measures side by side, which was the point. The step after it changes something else entirely: it changes the kind of object the numbers are in. - **A labelled table** — a rectangle whose columns are named and each carries one type, with row labels attached — resolves a mismatch by **label**. Two operands drawn from differently ordered sources are matched row label to row label, and where one side has a label the other does not, the result carries the absent-value marker there. The mismatch shows up as holes. - **A plain rectangle of numbers** — a block of one single type addressed by position, with no column names or row labels to align on — resolves a mismatch by **position**. First with first, second with second. There is nothing to notice a disagreement with. | | Labelled table | Plain rectangle of numbers | |---|---|---| | Operands matched by | row label | position | | Differently ordered sources | matched correctly | combined in the wrong pairs | | A label present on one side only | a hole in the result | no concept of it | | Evidence when it goes wrong | visible holes | none | That asymmetry is the whole hazard of this boundary. While you are in the labelled world, ordering is somebody else's problem. The moment you leave it, the ordering you were implicitly relying on stops being checked, and the failure is a plausible number rather than an error. ## Three things to establish before you cross 1. **What lines the operands up on the far side.** If the far side matches by position, then the row order is now load-bearing. Either align the operands while they are still labelled and cross with a single object, or take both columns out of the *same* rearranged table in one step so they cannot have come from different orderings. 2. **What happened to the cells the widening had to create.** A widening fills the complete grid of row keys against headers, and where a subject never had that measure there is an **invented cell** — a cell that exists only because the arrangement demanded it. A labelled table will carry an absent-value marker there; a numeric block generally has no separate notion of absence, so those cells arrive as whatever the crossing turned them into. Decide explicitly whether those rows belong in the computation at all, before the routine averages over them. Worth noting in passing: where a tool marks absence by borrowing a value from the numeric range itself, gaining an invented cell can move a whole-number column to a different representation, whereas a tool that carries a separate validity bit keeps the width it had. Know which trigger you are on; the representation change itself is a separate subject. 3. **Whether the column positions are stable.** A widening takes its headers from the data, so the set and order of columns on the far side is a function of the input rather than of your code. Indexing the third column by position across runs is a bet that next month's input contains the same distinct values in the same order. Select by name while you still have names, and cross with exactly the columns you meant. ## Why this is a boundary question rather than an arithmetic one The numeric routine is not doing anything wrong. It is doing precisely what an object addressed by position does. The mistake happens one line earlier, at the crossing, where a candidate assumed a property of the object they left applies to the object they entered. That is the general shape of this leaf's failure: a layout claim carried from one world into another without checking the mechanism it depended on. The same care applies in reverse. Coming back out of a numeric block into a labelled table, the results have positions and no labels; attaching them to the wrong rows is the same error travelling the other direction, and it is just as silent. ## How to make the crossing safe - Do the widening and the column selection **in one place**, so both operands provably come from the same rearranged table and the same row ordering. - Select the columns **by name** at the last moment you still have names. - Decide the disposition of invented cells **before** crossing, not by seeing what the numbers look like afterwards. - If the two operands genuinely come from different sources, align them **while both are still labelled**; do not fix it by sorting both and hoping. ## What an interviewer is listening for The strong answer names both alignment rules and says which object has which — it does not assert one of them as the way things work. Candidates who have only ever used one of the two designs tend to state that design's rule flatly, and that is exactly the habit this question is built to find. The second half of a strong answer is the invented cells and the data-derived column set, because both are consequences of the widening that only bite once the labels are gone.
- The two operands came from two differently ordered sources. Where do you fix that?While both are still labelled, so the match is made on row labels and any label present on only one side shows up as a hole you can see. Sorting both and crossing on faith replaces a checked match with an assumption. Once you are in a position-matched object there is nothing left that could disagree with you.
- Why is indexing a widened result by column position risky?Because a widening takes its headers from the data, so the set and order of columns is a function of the input rather than of your code. A distinct value that appears for the first time next month shifts every position after it. Select by name while names exist, and cross the boundary with exactly the columns you meant.
- The routine returned numbers that look plausible but are wrong. How would you confirm the alignment is the cause?Compute the same result for a handful of subjects in the labelled world, where the match is made on row labels, and compare. If the labelled answer differs from the positional one, the operands were paired by position against an ordering you did not guarantee. The disagreement is the evidence; the plausible numbers never will be.
saying these in an interview costs you the question
- Assumes the two operands are lined up by row label whatever object they sit in
- Treats the crossing as a change of arrangement rather than a change of object
- Sorts both sides and calls the ordering guaranteed
- Indexes the widened result by column position across runs
- Ignores the cells the widening created for subjects that never had that measure
- Expects a mismatch to raise something rather than produce a plausible number