skip to content

A per-group numbering returned different rows this month with no code change — what should you check about its ordering?

level: seniorimportance: should knowfreq 50%

answer

  1. an invisible input to the numbering
  2. declared at the call, or inherited
  3. which step last arranged the rows
  4. a row count will not catch it
  5. assert that row 1 is the right row

basics

~20 s

Check whether the numbering states its own ordering or inherits the row order it happens to find. An inherited ordering is owned by whatever step ran before it, so an unrelated upstream change silently renumbers every group.

solid answer

~50 s

There are two designs, and the failure only exists in one of them. If the numbering takes the ordering as an argument, it is self-contained and an upstream step that rearranged rows cannot touch the numbers. If the numbering instead numbers rows in whatever order it finds them, the ordering was established somewhere earlier, and *anything* that changed that order — a different source file, a rewritten earlier step, a step that now emits rows in a different arrangement — renumbers every group without touching a line of numbering code. So: read the call and ask what it is ordered by. If the answer is not in the call, find the last step that established an order and ask whether that step is deterministic and whether anything now sits between it and the numbering. Then close it: state the ordering at the numbering if the surface allows, otherwise establish it immediately before with nothing in between, and assert the result.

go deeper

for a junior

Take away one habit: when you see a per-group numbering, find out what it is ordered by before you trust it. If the call does not say, the answer lives somewhere earlier.

for a middle

Explain the two designs and why only the inherited one has this failure. Name the upstream events that can rearrange rows without any change to the numbering code.

for a senior

Run the diagnosis in order — read the call, trace the last arranging step, reproduce with last month's input, diff row identities not counts — and close it with a stated ordering plus an assertion.

for a principal

The call you own is whether pipelines are allowed to rely on inherited row order at all, and what standing assertion accompanies every per-group numbering whose output someone acts on.

## The two designs Numbering rows inside a group — giving each row of a group a position under an ordering, as part of a grouped operation that keys rows, computes per key and reassembles — gets its ordering from one of two places, and which one your tool uses decides whether this failure can happen at all. 1. **Declared at the call.** The ordering is an argument to the numbering itself. The numbering establishes the order it needs, uses it, and leaves the table's own arrangement alone. No earlier step can change the numbers. 2. **Inherited from the rows.** The numbering assigns positions in the order the rows are currently sitting in. The ordering is therefore a property of the pipeline up to that point, not of the numbering, and it is held by whichever step last arranged the rows. Both designs are real and both ship. The common error is to assume design 1 while working in design 2 — writing the numbering, seeing correct output because an earlier step happened to leave the rows the right way round, and shipping a result whose correctness rests on a coincidence nobody wrote down. ## Why the answer changed with no code change In design 2 the numbering has an invisible input. Any of these changes it without appearing in a diff of the numbering: - the source data arrives in a different arrangement this month; - an earlier step was rewritten and no longer emits rows the way it did; - a step was inserted between the ordering step and the numbering; - the earlier step's arrangement was never guaranteed in the first place and only looked stable. The result is a numbering that is internally consistent, passes every type check, has the right row count, and answers a different question than it did last month. When the numbering feeds a cut — keeping the top few rows of each group — the visible symptom is that *different rows survived*, which is exactly the report that reaches you. ## How to run the diagnosis 1. **Read the call and answer one question: ordered by what?** If the ordering is written at the numbering, stop — this is not your failure, and the cause is in the data or in the grouping key. 2. **If the ordering is not in the call, find the step that last arranged the rows.** Trace forward from it to the numbering and list everything in between. Every one of those steps is a suspect. 3. **Ask whether that arranging step is deterministic for equal rows.** If two rows of a group are equal under whatever it ordered by, their relative arrangement — and therefore their numbers — was never pinned down. 4. **Reproduce deliberately.** Rerun last month's input through today's pipeline. If the numbers differ, the ordering is inherited and something upstream moved; if they match, the input changed. 5. **Compare the two results as rows, not as counts.** The row count of a fixed-size cut is unchanged by a renumbering, which is why this failure survives a count check. Diff the identities. ## Closing it properly | Fix | What it buys | When it applies | |---|---|---| | state the ordering at the numbering | the numbering stops having an invisible input | the surface accepts an ordering argument | | establish the order immediately before, nothing in between | narrows the blast radius to one adjacent step | the surface consumes the order it finds | | make the ordering separate every pair of rows in a group | removes the equal-rows case, so arrangement stops mattering | always worth doing when the surviving row is acted on | | assert the invariant after the numbering | turns a silent renumbering into a loud failure | any pipeline whose output someone acts on | The assertion is the part teams skip and the part that pays. A cheap one: for a sample of groups, check that the row numbered 1 really is the row that wins under the ordering you believe in. That fails loudly on a renumbering, where a row count does not. ## The claim to avoid making It is tempting to say *numbering rows inside a group requires sorting the table first*. That is true only of the inherited design; in the declared design no prior arrangement is needed and the table's own order is untouched. The portable statement is the one worth carrying into an interview: **either the ordering is an argument to the numbering or it is the arrangement the rows are already in — and in the second case a distant step owns your answer.**

  • Why does a row-count check fail to catch this?
    Because a renumbering rearranges which rows are kept, not how many. A fixed-size cut returns the same count from the same groups whichever rows won. Only a comparison of row identities, or an assertion that the top row matches the ordering, exposes it.
  • If the surface has no ordering argument, what is the best you can do?
    Establish the order in the step immediately before the numbering with nothing between them, make it specific enough to separate equal rows, and assert the result. That does not remove the dependency but shrinks it from the whole pipeline to one adjacent, visible step.
  • Is this the same problem as an ordering that is simply wrong?
    No. A wrong ordering is a bug you can read in the code and reproduce forever. This one is a correct-looking call whose input arrives from outside it, so it reproduces only on the pipeline that produced the arrangement — which is why it surfaces as a result that changed by itself.

saying these in an interview costs you the question

  • Assumes every numbering sorts the rows itself before numbering.
  • Treats an ordering established anywhere earlier in the script as safe.
  • Says a result cannot change while the code is unchanged.
  • Checks only the row count, which a renumbering leaves intact.
  • Blames the data before asking what the numbering ordered by.
  • Assumes two tools take their ordering the same way.