Orders are grouped by customer and each row numbered inside its group by date; what does a five-order customer receive?
answer
- per row, not per group
- restarts at every group
- 1 up to the size of the group
- an ordering you chose, not the numbering's
basics
~20 sFive numbers, one per row, running 1 to 5 and restarting at every customer. A position inside a group is a per-row value, not a single value for the group, and the chosen ordering decides which row is 1.
solid answer
~40 sThe customer's five rows each get their own number, 1 through 5, and the numbering restarts from 1 at the next customer — so a table of ten thousand orders across two thousand customers holds no number higher than the largest group's size. The result has exactly as many rows as the input; nothing is folded away. Which row holds 1 is decided entirely by the ordering you chose for the numbering: earliest date first puts the oldest order at 1, latest first puts the newest there. The numbering contributes no ordering of its own — it consumes one. That makes three separate things worth keeping apart: the grouping key (one value per group), the size of the group (one value per group), and the position inside the group (one value per row).
go deeper
Recall the shape: one number per row, starting again at 1 in each group, and the ordering you picked decides which row is 1. Say the result has the same row count as the input.
Explain that the numbering consumes an ordering rather than creating one, and that the ordering can be an argument to the call or inherited from an earlier step depending on the tool.
Show that you check the result: maximum value against the largest group's size, and one group spot-checked against the ordering you believe was used. Silent mis-numbering is not visible in a row count.
The angle worth owning is the standing rule: whether per-group numbering in your pipelines must always declare its ordering explicitly, so no distant step can change a downstream answer.
## What the operation hands back A **grouped operation** is one pass that gives every row a key, runs a computation once per key, and reassembles the answers. Numbering rows inside a group is the variety of that pass whose answer is **one value per input row**. The customer with five orders receives five numbers, and the table that comes back has as many rows as it started with — nothing has been folded away into a summary. Inside each group the numbers run from **1 up to the size of that group**, and they **restart at every group**. Two customers with five orders each both carry the numbers 1 to 5. There is no number 6 anywhere in that result unless some customer has six or more orders, and no number is unique across the table: `1` appears once per group. ## What the ordering decides A position only means something relative to an ordering, and the ordering is yours to pick — earliest date first, largest amount first, whatever the question needs. The numbering itself supplies none of it. - Which row of the group carries **1** is entirely a consequence of the ordering you chose for the numbering. - Reverse that ordering and every group's numbers are mirrored: the row that held 1 now holds a number equal to the size of its group. - Changing only the ordering does **not** change the number of rows in the result, and does not change which rows belong to which group. It changes only which row carries which number. - The one case the ordering does not settle by itself is rows that are equal under it. What those rows receive is a separate decision with real consequences, and it is the part interviewers push on. ## Four things that are easy to confuse | Thing | How many values | What it depends on | |---|---|---| | the grouping key | one per group | the key expression alone | | the size of the group | one per group | how many rows share that key | | a position inside the group | one per row | the ordering chosen for the numbering | | a position across the whole table | one per row | an ordering over every row, ignoring the grouping key | The last row of that table is the one juniors reach for by mistake. A position across the whole table is a single ascending run from 1 to the row count; a position inside a group is many short runs, each starting again at 1. ## Where designs differ This is a mechanism several tools implement, and they do not implement it the same way. Two forks matter even at this level: 1. **Where the ordering comes from.** Some surfaces take the ordering as an argument to the numbering itself, so the call is self-contained. Others number the rows in whatever order they are currently sitting in, so the ordering has to be established by a step before it. Both designs are real, and the question *what is this numbered by* is answered either in the call or somewhere upstream. 2. **What order the result arrives in.** Some designs hand the numbers back aligned onto the rows where they already were, leaving the table's own row order untouched; others hand back the rows arranged by the ordering used for the numbering. Read the result rather than assuming either. A third difference is what equal rows receive, which differs by tool and by option and decides which rows survive later. ## Reading a result you did not write When a column of small integers turns up beside your rows, three checks tell you what it is: - **Does it restart?** Scan down: if the values go 1, 2, 3, 1, 2, 1, 2, 3, 4 it is per-group; if they run 1 to the row count without repeating, it is a position over the whole table. - **Does its maximum match a group's size?** The largest value anywhere should equal the number of rows in the largest group, never the row count of the table. - **What is it ordered by?** Pick one group and check that its number 1 really is the row you would expect under the ordering you think was used. If it is not, the ordering is not the one you think. ## Why an interviewer asks it It separates a candidate who can predict a result from one who can only produce it. The prediction here is cheap to state and easy to check: same row count as the input, values from 1 to the size of each group, restarting at every group, with the row holding 1 chosen by an ordering somebody picked on purpose. A candidate who says *one number per customer* has the shape of the operation wrong, and everything built on it — keeping the top few of each group, carrying a total down a group's rows — will be wrong the same way.
- How many rows does the result have once the numbering is added?Exactly as many as the input. Numbering is a per-row answer: every row keeps its place and gains a value. That is different from a pass that collapses each group to a single row, where the result's row count is the number of groups instead.
- If the groups are wildly different sizes, what does that do to the numbering?Nothing structural. Each group numbers independently from 1 to its own size, so a group of two ends at 2 while a group of fifty thousand ends at fifty thousand. The largest number anywhere in the result is simply the size of the largest group.
saying these in an interview costs you the question
- Says each group receives a single number.
- Expects the numbering to continue across group boundaries.
- Thinks the numbers are positions in the whole table.
- Believes the numbering imposes an ordering of its own.
- Assumes changing the ordering changes the result's row count.