How do you cut an input so that the dominant key and the rare record shapes both survive into the development sample?
answer
- assembled, not drawn
- head, middle, tail of shapes
- search for the rare case
- cap the heavy key, record the factor
- manifest says what is missing
basics
~20 sMeasure key frequencies once, then build the cut in three deliberate parts: the heaviest keys kept whole or at a recorded fraction, a stable draw over the ordinary middle by whole key, and one searched-for example of every rare shape. Chance supplies none of these.
solid answer
~50 sTreat the cut as something you assemble, not something you draw. First run one grouped count over the whole input to learn how records are spread across keys. Then keep three strata for different reasons: the **head** - the keys carrying the most records - because an even cut is a different problem from the real one; the **middle**, drawn by whole key with a deterministic rule, to supply ordinary volume; and the **tail of shapes**, found by querying for each awkward case you know of - a null in the join column, a negative or zero amount, an unusual encoding, a key with exactly one record, a duplicate. If the head key is too large to keep whole, keep a stated fraction of its rows and record the factor, so nobody reads its total as real. Rebuild the cut as the input drifts.
go deeper
Know that a useful cut is put together on purpose - the biggest key and the odd records are chosen, not left to a random draw that will usually skip both.
Explain the three strata and why each exists: the head preserves the uneven shape, the middle supplies ordinary volume by whole key, and the tail of shapes has to be found by query.
Show the operational half: a rebuildable recipe, a manifest recording reduction factors and shapes searched for but not found, and an explicit statement of what the cut cannot prove.
Decide what the team commits to - how often the cut is rebuilt as the input drifts, who owns the recipe, and the standing rule that a cut is never evidence for a volume or cost claim.
## Why chance will not give you the cut you need The two things that make a real input hard are at opposite ends of the same frequency distribution: a handful of keys carrying a very large share of the records, and shapes that occur so rarely that most days contain none. A uniform draw is the worst possible tool for keeping both. It shrinks the head toward the average - the dominant key stops dominating - and it misses the tail almost every time, because a one percent draw over a shape occurring once in a million rows is not a small sample of that shape, it is usually zero of it. So the cut is **assembled**, in named parts, each kept for a stated reason. That is also what makes it reviewable: a colleague can ask why a key is in the cut and get an answer. ## The three strata 1. **The head - the keys carrying the most records.** Keep them because their absence makes the cut a different problem from the real input. The uneven case is the one the job has to survive, and a cut where every key has comparable weight tests a world that does not exist. What one dominant key then does to the run itself, and what is done about it, belongs to a different subject; the rule here is only that the cut must not quietly drop it. 2. **The middle - ordinary keys, drawn by whole key.** Use a deterministic rule computed from the key value, so the same keys are chosen every rebuild and so every joined input selects the same ones without coordinating. This stratum supplies bulk and ordinary behaviour, and its rate is whatever brings the total runtime down to seconds. 3. **The tail of shapes - found by search, not by draw.** Enumerate the awkward cases you know of and query for one example of each. A working starting list: a null in the join column, a null in a column the job assumes is populated, an empty string against a genuine absence, a negative or zero amount, a value at the maximum length, an unusual encoding or a non-ASCII value, a key with exactly one record, two records identical except for arrival order, a timestamp outside the period the input claims to cover. ## Capping a head key honestly A dominant key can carry more records than the whole cut is allowed to hold, so keeping it whole is sometimes impossible. There is an honest way to cap it and a dishonest one. | approach | effect on shape | effect on assertions | |---|---|---| | keep the head key whole | shape preserved exactly | per-key totals are the real ones | | keep a stated fraction of its rows, and record the factor | shape preserved in kind, reduced in degree | totals for that key are known to be scaled, not real | | drop the head key | the uneven case disappears | the cut silently tests a different world | | split the head key into several invented keys | distribution destroyed | per-key totals meaningless | The second row is the working compromise: the key is still much larger than the others, so the cut still has an uneven shape, and because the reduction factor is recorded nobody mistakes its total for a production number. The fourth row is worth naming explicitly because it resembles a runtime remedy for uneven work - which is a different subject entirely, and applying it to the *input* rather than to the *run* corrupts the very thing the cut exists to preserve. ## Keeping the cut honest over time An input's shape drifts: a new source is onboarded, a key that was ordinary becomes dominant, an encoding changes upstream. A cut assembled a year ago is a faithful picture of last year. Two habits keep it useful: - **Rebuild from a recipe, not by hand.** The cut should be the output of a script that anyone can re-run: the frequency aggregate, the deterministic key rule, the named head keys, the shape queries. Then rebuilding is cheap enough to do, and the cut is reproducible. - **Keep a manifest beside it.** Which keys are in the head stratum, which reduction factor was applied to each, which shapes were searched for and found, which were searched for and *not* found. That last line is the one that saves an afternoon: a shape absent from the cut is absent from your testing, and you want that written down rather than assumed. ## What the cut still does not cover The cut preserves shape, so it exercises logic. It does not preserve volume, so it exercises nothing that volume causes. Running out of memory does not happen, and neither does **spill** - writing part of the working set to local disk when memory runs out. The cost of moving records between workers so a step can see all the records for a key does not appear either, and runtimes differ in what that even costs: some write the records out and have every worker fetch its share, others push them across the network as they are produced. Losing a worker and recovering is invisible. State the limit when you hand the cut to someone: it answers *is this logic right*, never *will this run finish*.
- How do you find the rare shapes if you do not already know what they are?Profile rather than guess: per column, count nulls, empty strings, minimum and maximum values and lengths, distinct counts, and the frequency of the top values. Anything with a tiny non-zero count is a candidate shape. Then add every shape that has ever caused an incident - the incident list is the best source there is, because those shapes are known to break this job rather than merely being unusual.
- Why record the reduction factor when you cap a dominant key?Because the cut is otherwise assertable and that one key is not. Per-key totals for whole keys equal the production totals, so a reader will reasonably assert on them. If the heavy key was reduced and nobody wrote it down, someone eventually asserts its total, gets a number that means nothing, and either chases a phantom defect or bakes the scaled figure into an expected result.
- Should the cut include shapes that were searched for and not found?It cannot include them, which is exactly why the manifest records the attempt. A shape you looked for and did not find is a gap in your testing, not a case that does not exist - it may simply be rarer than the period you searched. Note it, and consider hand-writing a record of that shape rather than pretending the gap is closed.
saying these in an interview costs you the question
- A random draw large enough will contain the rare cases eventually
- Drop the heaviest key, it makes the cut too big
- Even out the key sizes so the cut runs predictably
- Rare shapes are data-quality problems, not the job's concern
- If the cut is representative, the real run will behave the same way
- One cut built at project start serves the whole project