Swapping the order of two grouping keys changes what about the collapsed result, and what does it leave untouched?
answer
- same pairs, different arrangement
- outer part against inner part
- reading order, not values
- swap for readability, never for correctness
basics
~20 sKey order changes the layout - which key is the outer part of the identity, and so how the result is arranged and reads - but not a single aggregated value, because the set of pairs is identical either way.
solid answer
~40 sThe identity of a group is the pair of key values, and a pair is the same pair however you write it down: the set of groups and every number computed from them are identical whichever key is named first. What changes is presentation. In a design that composes the keys into one row label made of two parts, the first key named becomes the outer part, so the result nests one way rather than the other; where the split leaves an ordering behind, that ordering runs by the outer key first. In a design returning the keys as ordinary columns, the column order changes and little else. So key order cannot explain two reports disagreeing on a number, though it often explains why one of them is unreadable.
go deeper
Recall that naming the two grouping keys in the other order gives the same groups and the same numbers. What moves is how the result is arranged and how it reads.
Explain the mechanism: the pair is unordered as a membership test but ordered as a layout, so the first key named becomes the outer part and drives any ordering the result carries.
Use it deliberately. Choose the order that makes the result readable and cheap for the next step to consume, and never let a review blame a numbers discrepancy on key order, which cannot produce one.
Fix an arrangement convention for figures several teams read. The same numbers published in two arrangements look like two analyses, and people will reconcile them long before anyone notices they already agree.
## The identity is a pair, and a pair is the same pair either way Two records belong to the same group when they agree on every grouping key. That test does not consult the order in which the keys were named: "region A with product X" and "product X with region A" describe the same set of records. So the set of groups is identical under either naming, the records folded into each group are identical, and every number computed from them is identical. Say that part first in an interview, because it is the part that decides correctness. Whatever else swapping two grouping keys does, it cannot change a value, cannot add or remove a group, and cannot change the result's row count. ## What key order does change The identity is unordered as a membership test but ordered as a layout. The first key named is the **outer** part of the identity and the second the **inner**: | aspect | keyed region then product | keyed product then region | |---|---|---| | the set of identities | the same pairs | the same pairs | | every aggregated value | identical | identical | | the row count | identical | identical | | the outer part of a composed label | region | product | | column order, where keys come back as columns | region first | product first | | any ordering the result does carry | region varies slowest | product varies slowest | | what the result reads as | for each region, its products | for each product, its regions | The last line is why anyone cares. The two results contain the same facts and answer different questions comfortably. A reader scanning for "how does this product do across regions" finds the second arrangement obvious and the first exhausting, even though a machine reads them identically. The same is true of a later step: one arrangement may let it walk contiguous blocks of one key while the other makes it hop. ## An ordering is a by-product, not a guarantee It is tempting to say the result comes back sorted by the outer key. That depends on the split mechanism, and both kinds exist in this family: - A split that **orders the key values as part of forming the groups** leaves an ordering behind, and with two keys that ordering naturally runs outer key first, inner key within it. - A split that **buckets keys by a hash** produces no ordering at all. Keys come back roughly in the order they were first seen, and swapping the two keys changes which identities are adjacent without producing anything you could call sorted. So the honest formulation is: key order decides which key is outer, and the outer key drives any ordering the result happens to carry. If the ordering matters to a reader or to a later step, ask for it explicitly rather than inheriting whatever the split happened to leave behind. That habit also survives a change of tool, which an inherited ordering does not. ## Using key order deliberately - **Put the key the reader scans by first.** A report read "per region" wants region outer. This is the single most useful application of the rule and it costs nothing. - **Put the coarser key first when the result is long.** A few large outer blocks read better than hundreds of tiny ones, and they compress better in whatever is written out. - **Match the order the next step expects.** If a downstream consumer walks the result expecting one key to vary slowest, honour that rather than making it re-sort. - **Do not use key order to fix a number.** If two results disagree on a value, the cause is elsewhere: a different filter, a different source, a different definition of a derived key. Key order is incapable of producing a discrepancy. ## The mistake in both directions Two opposite misconceptions show up, and each is half right. The first is that key order changes the answer - it does not, and a candidate who believes it will waste a debugging session on the wrong hypothesis. The second is an overcorrection: that key order is meaningless because the groups are the same. It is not meaningless. It fixes which key is the outer part of the identity, which decides the arrangement, the column order, and whatever ordering the result carries. Two reports that are pair-for-pair identical can be genuinely different to read, and that difference is the one thing key order owns.
- If key order cannot change the numbers, why does the result sometimes look completely different?Because the arrangement changed, not the content. With a composed label the outer key blocks the rows and the inner key varies fastest, so the same values appear in a different sequence, and where the split leaves an ordering behind that ordering now runs by a different key first. Read the row identities rather than the row positions and the two results match pair for pair.
- Is the result guaranteed to come back ordered by the outer key?No. Some split mechanisms order the key values while forming the groups and some do not; a hash-bucket split produces no ordering at all and returns keys roughly in the order they were first seen. Key order decides which key is outer, and the outer key drives any ordering that does exist. If the ordering matters, ask for it explicitly rather than inheriting it from the split.
Filing the same invoices into folders by region and then by month, or by month and then by region, produces the same set of bottom folders holding exactly the same invoices. Only the cabinet's arrangement differs, and with it how easy it is to answer a given question by walking the drawers.
saying these in an interview costs you the question
- Says swapping the grouping keys changes the aggregated numbers
- Says key order has no effect at all on the result
- Expects the row count to change when the keys are reordered
- Assumes every result comes back ordered by the outer key
- Cannot say which key becomes the outer part of the identity