skip to content

A grouped total returns its keys in no obvious order — what does that say about the split?

level: middleimportance: should knowfreq 50%

answer

  1. the order was never asked for
  2. how the groups were formed shows through
  3. sorted runs against hashed buckets
  4. by-product, not guarantee

basics

~20 s

That the split formed groups by hashing key values into buckets rather than by ordering them. Any ordering in a grouped result is a by-product of the split mechanism, not a guarantee — order the result explicitly if you need it.

solid answer

~50 s

Tie the ordering to the mechanism. A **sorted split** orders the key column so equal keys form a contiguous run, and an ordering by key falls out of that as a by-product. A **hashed split** drops each key value into a bucket by its hash and leaves no ordering at all; the keys typically come back in the order they were first seen, because that is the order the bookkeeping recorded them in. Some tools also sort the keys after the fact as a default. So an unordered result usually means a hashed split with no post-sort. The rule to state out loud is that ordering here is a by-product, not a contract: it can change with a tool version, a different key type, or a switch of split strategy, and it will change silently. If a consumer depends on key order, order the result as an explicit step of its own.

go deeper

for a junior

Know that a grouped result is not promised in key order, and that if a report needs the keys in order you must ask for that order somewhere.

for a middle

Explain the fork: a sorted split leaves an ordering behind as a by-product, a hashed split leaves none and usually returns keys in first-appearance order.

for a senior

Recognise the silent failure — a consumer built on an undocumented order that moves when a version, a key type or an input order changes — and make the requirement explicit in code.

for a principal

Set the team rule that no output contract may rest on an ordering nobody asked for, since that is a dependency no review can see and no test reliably catches.

## The question behind the question An interviewer asking this is checking one thing: do you know that a grouped result's order is produced by *how the groups were formed*, rather than being part of what a grouped operation promises? A candidate who answers "grouped results come back sorted" and a candidate who answers "grouped results come back in arbitrary order" are making the same mistake in opposite directions. ## Two ways to form groups | split mechanism | how groups are found | ordering left behind | typical cost shape | | --- | --- | --- | --- | | sorted split | the key column is ordered so equal keys sit in one contiguous run | an ordering by key falls out for free | pays an ordering cost up front, then sequential reads | | hashed split | each key value is dropped into a bucket by its hash | none at all | no ordering cost, scattered reads | Both produce exactly the same grouping — the same rows in the same groups — and differ in cost and in by-products. A third case sits on top of both: a tool may sort the keys *after* forming the groups, as a default or on request, in which case the result is ordered whichever mechanism ran underneath. ## First-appearance order is not the same as no order When a hashed split returns keys "unordered", what usually comes back is **first-appearance order**: the order in which each distinct key value was first encountered while sweeping the rows. That is not random, and that is exactly the trap. - It looks stable, because rerunning the same job over the same file reproduces it. - It has a plausible story behind it, so people rationalise it as intentional. - It is nevertheless an artefact of the bookkeeping, and nothing documents it as a contract. ## Why a by-product is dangerous in a way that a guarantee is not The failure mode is silent. Nothing raises, nothing warns, and the change is not in your code: 1. Something upstream re-orders the input, so first-appearance order changes and a report's rows move. 2. The key column's type changes, and the tool switches from a hashed to a sorted strategy or the reverse. 3. A tool version changes the default, or the planner picks a different strategy for the same statement at a different data size. 4. A downstream consumer — a chart, a file diff, a fixed test — was quietly built on the old order, and now disagrees with no error anywhere. The cost is rarely a wrong number. It is a diff that will not settle, a chart whose categories reshuffle between runs, or a test that fails on a machine where the data arrived differently. ## The rule State it as two clauses: - **Ordering in a grouped result is a by-product of the split mechanism.** Say which mechanism would produce which by-product if asked; do not state a flat outcome as though every tool behaved the same way. - **If order matters, make it explicit.** Order the result as a step of its own after the grouping, so the requirement lives in the code rather than in an implementation detail you did not choose. ## Two things this is not - It is **not** about whether equal keys kept their original relative order through the split. That is a property of the sort underneath, and a different subject. - It is **not** an argument for sorting the input before grouping. A sorted input does not oblige the split to preserve that order into the result, and the ordering cost is paid twice if the split orders the keys anyway. ## What good sounds like in the room "Whether the keys come back in order depends on how the split formed them. A sorted split leaves an ordering behind; a hashed split does not and will usually give you first-appearance order. Either way it is a by-product, so I would not build a consumer on it — if the report needs key order, I would order the result explicitly and let the reader see that requirement."

  • If a hashed split leaves no ordering, what order do the keys actually appear in?
    Usually the order in which each key value was first seen, because that is the order the bookkeeping recorded them in. It is an artefact of the implementation rather than a contract, and a change of tool, version or key type can change it with no warning.
  • Why does it matter that the ordering is a by-product rather than a guarantee?
    Because a by-product disappears silently. Code and downstream consumers that grew to depend on it keep working until something upstream changes the split strategy, and then output changes with nothing raising an error to point at the cause.
  • Does sorting the input table before grouping give you an ordered result?
    Not reliably. The split is free to hash the keys regardless of the input's order, and a sorted split will order the keys on its own terms anyway. You have paid an ordering cost and still have no guarantee about the result.

saying these in an interview costs you the question

  • Says a grouped result always comes back sorted by key.
  • Says a grouped result is always in arbitrary order, whatever the mechanism.
  • Treats first-appearance order as a documented guarantee.
  • Sorts the input before grouping and calls the result ordered.
  • Assumes the same tool keeps the same order after a version or key-type change.