skip to content

An ordering key is absent on a quarter of the rows — where do those rows land in the result?

level: middleimportance: should knowfreq 54%

answer

  1. not a default worth memorising
  2. first, last, optional, or undecided
  3. the direction can move them
  4. order on an absent-or-not key first

basics

~20 s

Wherever the design puts them, which is not something to assume. Some place them last, some first, some expose an option that says which end, and some leave it to the comparison. State the placement you want, or arrange it with a derived key.

solid answer

~40 s

Placement of rows whose ordering key is absent is a per-design choice, and designs in this space genuinely disagree. Some document a fixed end, commonly last. Some offer an absent-value placement option — the knob that says which end those rows go to — and a few resolve it relative to the direction, so flipping from ascending to descending also moves the absent rows. Some make no decision and leave placement to the underlying comparison, in which case the rows need not even end up together. The portable move is not to remember whose default is what: set the placement option where one exists, or order first on a derived key that says only whether the real key is absent, which puts those rows at an end you chose under any design.

go deeper

for a junior

Know that a row whose ordering key has no value still appears in the output and has to go somewhere, and that where it goes is a property of the tool rather than something you can read off the data.

for a middle

Explain that placement is a fixed documented end in some designs, a settable option in others, and a side effect of the comparison in the rest, and that a direction change can move those rows in the designs that tie the two together.

for a senior

Show the habit of making placement explicit rather than inherited, and connect it to the downstream step that reads one end of the result, because that is where an inherited default turns into a wrong answer with nothing raised.

for a principal

Decide once for the codebase whether rows with an absent ordering key are placed at a stated end or removed before ordering, and write the rule down. It costs one derived key; the absence of a rule costs a silent change of meaning on a tool upgrade.

## A row with no key still needs a place Every row in the input appears somewhere in the output of an ordering step — the operation that puts rows into an order you specify, taking one or more keys each with its own direction. A row whose ordering key has no value is no exception: it has to be somewhere. The question this leaf owns is *where*, and whether the tool lets you say. This is narrower than it sounds, and worth keeping narrow. What an absent value does inside arithmetic, and what a comparison against one yields, are separate subjects. Here the only thing at stake is the position in the ordered result. ## Four behaviours that all exist | Design behaviour | What you get | What it costs you | |---|---|---| | A documented fixed end | The absent rows reliably at one end, usually last | Nothing, until the pipeline runs on a different tool | | A settable placement option | Whichever end you asked for | You must actually ask; the unset default is still a per-design choice | | Placement resolved against the direction | One end ascending, the other end descending | Flipping the direction quietly moves them | | No decision at all | Whatever the comparison left behind | The rows need not be adjacent to one another | The last row is the one people find surprising. If the underlying comparison does not treat an absent value as an extreme, the rows carrying one are not being ordered relative to anything, and they can appear scattered through the result rather than gathered at an end. ## Why this becomes a correctness bug rather than a cosmetic one An ordering step is almost never the last step. What follows it usually assumes something about the ends. - A listing that shows the head of the ordered result will show rows with no key at all, and will show fewer real rows than the count suggests. - A step that takes rows from one end to represent "the largest" silently changes meaning when the absent rows are at that end. - A comparison between two runs of the same report starts failing when a tool upgrade changes the default, with no code change to point at. - A visual scan of the output looks fine: the rows are present, the counts add up, and nothing raised. That combination — no error, plausible output, a cause outside your own diff — is what makes placement worth deciding rather than inheriting. ## Making placement yours 1. **Decide whether those rows belong in this result at all.** Often the honest answer is that a row with no key is not a candidate for a ranked listing, and removing it before the ordering step is clearer than placing it. 2. **If they stay, set the placement explicitly where the design offers an option**, and write the setting even when it matches the default, so a reader knows it was a decision. 3. **If the design offers nothing, order on a derived key first.** Compute a value that is one where the real key is absent and zero where it is present — or the reverse — and make that the first ordering key, with the real key second. The derived key is an ordinary comparison that every design agrees about, so the absent rows land at the end you chose regardless of tool. 4. **Re-check after any change of direction.** Where placement is resolved against the direction, switching a listing from ascending to descending moves the absent rows too, and the derived-key approach is immune to that because the derived key has its own direction. ## Two things not to say in an interview The first is "they go to the end" stated flatly. It is a real behaviour of real tools and it is exactly the claim that fails on the next tool along. Say instead that placement is a documented fixed end in some designs, a settable option in others, and an artefact of the comparison in the rest, and that the safe habit is to specify it. The second is that absent rows are simply dropped. Some designs do exclude them from the comparison and append them afterwards, which is a placement decision rather than a removal — the rows are still in the output, and the count still includes them. Removal is something you can choose to do, and doing it deliberately, before the ordering step, is usually better than hoping the step did it for you.

  • How do you make absent-key placement identical under two tools that disagree about it?
    Order on a derived key first: a value that is one where the key is absent and zero where it is present, then the real key second. The derived key involves no absent values, so every design compares it the same way, and the rows land at the end you picked. It also survives a change of direction on the real key.
  • Why does a step that takes rows from one end of an ordered result make placement a correctness issue?
    If the absent rows sit at that end, the step returns rows with no key, and fewer real rows than the caller asked for. The output is well formed and the count is right, so nothing flags it. Placing the absent rows at the far end, or removing them before ordering, makes the step's input honest.

saying these in an interview costs you the question

  • States flatly that rows with an absent key always come last.
  • Assumes the placement stays put when the direction is flipped.
  • Believes absent rows are dropped from an ordered result entirely.
  • Relies on a placement observed once and never written down.
  • Thinks absent rows must at least end up next to each other.