Two pieces carry the same five column names in a different order and are stacked - what can land where, and how would you notice?
answer
- which rule: name, refusal, or position
- sets have no order
- equal widths mean no error and no holes
- right shape, wrong contents
- compare each column across the piece boundary
basics
~20 sIt depends entirely on the reconciliation rule. Reconciled by name, column order is irrelevant and the result is correct. Reconciled by position, the second piece's values land under the wrong headings with no holes and no error. Establish which rule the operation follows before trusting it.
solid answer
~50 sA stack has to decide how the pieces' columns correspond, and the rules differ. **By name** is the common case and makes column order irrelevant: the result is right. **By refusal**, the call fails until you make the pieces agree, which is the cheapest outcome. **By position**, the names are never consulted, so with two equal-width pieces every value from the second piece lands one heading over - and because the widths match, there are no holes and nothing is raised. The give-away is not in the shape of the result, it is in the content: look at each column separately for the rows that came from each piece, and you will see the values change character at the boundary. The defence is to compare each piece's ordered sequence of column names against an expected one before combining.
go deeper
Know that a stack has to decide how the pieces' columns correspond, and that not every tool does it by heading. If one might pair them by slot order, the same names in a different order stop being harmless.
Name the three rules - by name, by refusal, by position - and say what each does with a re-ordered piece. Explain why the positional case raises nothing: equal widths mean a well-defined operation, no holes, and the predicted row count.
Diagnose by content rather than shape: mark which piece each row came from, summarise each column per piece, and look for values changing character at the boundary. Then make it impossible by selecting the columns explicitly on every piece before combining.
The angle is where correctness is anchored. Depending on the combine operation's reconciliation rule puts a load-bearing assumption in a library; selecting an explicit column sequence on every piece moves it into your own code, where it survives a new producer or a change of tooling.
## Three rules, and only one of them ignores the order When two pieces are stacked one under the other, the operation has to decide which column of the first piece corresponds to which column of the second. Tools in this family answer that in three different ways: - **By name.** Columns are paired by their headings. Column order is then irrelevant, a name present on one piece only becomes a partly filled column, and this scenario is harmless. - **By refusal.** The call fails unless the two pieces present the same columns, sometimes in the same order. You find out at the seam, which is the outcome you would design for. - **By position.** The first column of one piece pairs with the first column of the other, whatever they are called. The headings of the result come from the first piece, and the second piece's values are placed underneath them in slot order. The question is therefore not "what happens when column orders differ" but "which of these three rules does my combine step follow", and it is a fair interview answer to say so before saying anything else. ## Why positional reconciliation is the dangerous one here Under the positional rule, with two pieces of equal width and identical names in a different order, every condition that would normally raise something is satisfied: | Signal you would normally rely on | What it shows here | |---|---| | An error or warning | none: the widths agree, so the operation is well defined | | The combined row count | exactly the sum you predicted | | The combined column count | exactly the width of each piece | | Holes from the absent-value marker | none, because no name was missing | | The column headings | all present, all spelled correctly | Everything about the shape is right. Only the contents are wrong, and they are wrong in a way that looks like data rather than like a bug: a column of amounts that contains dates for part of its rows, a region column that holds product names below a certain row, a count that is suddenly enormous. Contrast that with the name-mismatch failure, which announces itself with extra, partly filled columns. Positional reconciliation produces a result that is the right shape and quietly wrong, which is the worse of the two. ## How you would notice 1. **Summarise each column separately for the rows that came from each piece.** Keep a marker of which piece each row came from - added deliberately before the stack - and compare the two halves of every column. Distinct counts, ranges, and the share of empty values will all move abruptly at the boundary when values landed under the wrong heading. 2. **Check whether a column now holds two kinds of thing.** Values that were homogeneous within each piece and are heterogeneous in the result point directly at a slot-order problem. 3. **Compare the pieces' ordered column-name sequences before combining**, not just their sets. A set comparison passes here, which is exactly why this failure survives the check most people write. 4. **Look at a handful of rows from each piece side by side.** Reading two real rows is faster than any summary and usually settles it in seconds. The key insight for an interviewer is step three: the natural defence against the *name-mismatch* failure - comparing column names as sets - is blind to this one. Sets have no order. Guarding both failures means comparing the ordered sequence. ## Making it impossible rather than detectable - **Select the columns you want, in the order you want, on every piece, immediately before combining.** This one habit makes the pieces agree by construction, and it is cheap because the selection is written once and applied to each piece. - **Compare each piece's ordered column names against a single expected sequence** and fail on a difference nobody sanctioned. This converts a wrong number three steps downstream into a failure at the seam. - **Do not rely on "our tool reconciles by name"** as a complete answer. It is the right first question, but it is an argument about the tool rather than about the pipeline, and it stops being true the moment a piece is produced by a different path, or the operation is replaced by one from a different library during a migration. - **Keep a column recording which piece each row came from.** It costs one column and turns every question of the form "did the combine do this?" into something you can answer by grouping on it. The general shape of the lesson is the one this whole subject keeps teaching: combining operations are forgiving by design, the forgiveness is what makes them useful, and the price is that the result of a wrong combine is a plausible table rather than an error. The candidate who states the expected shape *and* the expected content before running the operation is the one who catches it.
- Why does comparing the two pieces' column names as sets miss this failure?Because a set has no order. Both pieces carry the same five names, so the set comparison passes, the column counts agree, and nothing in the shape of the result is wrong. Catching a re-ordered piece requires comparing the ordered sequence of names, or making the order irrelevant by selecting the columns explicitly on every piece first.
- If your tool reconciles a stack by name, is a re-ordered piece still worth guarding against?Yes, but for a different reason. The correctness now depends on a property of the tool rather than of your pipeline, and that property can change: a piece produced by another code path, or a migration to a different library, silently moves you onto another rule. Selecting the columns explicitly before combining costs a line and removes the dependency.
- What single cheap addition makes this class of failure easy to investigate afterwards?A column recording which piece each row came from, added before the stack rather than inferred after it. With it, every column can be summarised per piece in one step, and a value distribution that changes character at the boundary is immediately visible. Without it you are guessing where one piece ended.
saying these in an interview costs you the question
- Assumes column order never matters when adding records.
- Expects every stack to consult column names before combining.
- Thinks two pieces of equal width cannot be combined wrongly.
- Compares column names as sets and calls the pieces verified.
- Believes holes would appear if something had gone wrong.
- Inspects only the combined table, never each piece's columns.