skip to content

The first 100,000 rows of an input file are used as a development cut. What does that cut hide about the whole input?

level: juniorimportance: must knowfreq 56%

answer

  1. order of writing, not shape of content
  2. the head is one day or one writer
  3. rare shapes sit in the tail
  4. distribution, not row count
  5. shape yes, volume no

basics

~20 s

The first rows of a file are whatever was written first - one day, one source, one writer's share - so they carry neither the whole input's key distribution nor the rare record shapes that actually break the job.

solid answer

~50 s

A file's head is a slice of the order it was written in, not a slice of its content. Taking it gives you one day, one source system or one writer's output, so the mix of keys is wrong, the heaviest key is usually absent or ordinary-sized, and the rare shapes - a null in the join column, a negative amount, an unusual encoding, a key with a single record - sit somewhere in the tail you never read. The job then runs green on the cut and fails on the real input, and it fails on a case nobody chose to leave out. What you want instead is a cut whose shape resembles the whole: keys taken whole, the dominant key kept on purpose, and the rare shapes added deliberately rather than hoped for.

go deeper

for a junior

Recall that the head of a file reflects the order it was written in, not the content of the whole, so it carries the wrong mix of keys and none of the rare shapes.

for a middle

Explain the mechanics: which properties of the whole a contiguous slice cannot preserve, and why cutting by whole keys restores per-key totals that a row-level draw destroys.

for a senior

Show that you construct the cut deliberately - dominant key kept, rare shapes searched for, the same key set applied to every joined input - and that you say out loud what the cut still cannot prove.

for a principal

Frame it as a standing rule: what a shared development cut guarantees, how often it is rebuilt as the input drifts, and the fact that it is never evidence for a volume or runtime claim.

## What a development cut is, and what it is not A **development cut** is a small stand-in for a large input - a file, a table, or a day of records - reduced until a person can run the job over it in seconds and correct a change in one sitting. It is deliberately not the same object as a **statistical sample** used to estimate an aggregate cheaply. Those two pull in opposite directions: an estimator needs enough rows to bound its error, while a development cut needs to be as small as it can be *while still containing everything that makes the job behave interestingly*. Judging a cut by its row count is therefore the wrong measurement; the question is always **what shape did it keep**. The cheapest cut to produce is the head of the file: read the first N rows, stop. It is one command, it is fast, and it is almost always the wrong cut. ## Why the head of a file is a slice of the wrong dimension - **A file has a writing order, and that order is information about production, not about content.** Inputs are commonly written a day at a time, a source system at a time, or in the order records were extracted. The head is then one day, one source, or one extract - a slice along a dimension the job may not even model. - **Not every input is written in time order.** An input produced by many writers at once gives you a head that is one writer's share rather than the oldest rows. That is still a slice of the writing process rather than a slice of the content, so the conclusion does not change, but the reason does: do not assume the head is the past. - **Rare shapes live in the tail of a frequency distribution, not in a contiguous block.** A record shape that occurs once in a million rows has, in any single contiguous block, roughly the chance its rarity implies. You will not meet it, and it is exactly the shape that raises an error at three in the morning. - **Keys are spread through the file.** The key carrying the most records in the whole input contributes a handful of rows to the head, so in the cut it looks like any other key. Every property that depends on one key being much larger than the rest is gone. - **Counterparts in other inputs are accidental.** If you cut two joined inputs by taking each one's head, the keys in one head have very little to do with the keys in the other, and the join in your development run quietly returns almost nothing. ## What the head keeps and what it loses | property of the whole input | what the first rows give you | why it matters to the job | |---|---|---| | distribution of records per key | one contiguous slice, flattened | the run's behaviour on an uneven input is untested | | the key carrying the most records | present only by luck, at ordinary size | the case the job must survive is missing | | rare record shapes | almost never any | the error path is never exercised | | counterparts across joined inputs | accidental, usually poor | joins return far too few rows, or too many nulls | | value ranges and lengths | narrow, one period's worth | boundary handling untested | | repeats and duplicates of a key | under-represented | grouping and deduplication untested | ## What to take instead 1. **Aggregate the key frequencies once over the whole input.** One pass, one grouped count. This is the only expensive step and you do it rarely. 2. **Cut by whole keys, not by rows.** Choose a set of keys and keep every record belonging to them. Per-key totals in the cut then equal the real per-key totals for the keys kept, which makes the cut usable for checking a number and not only for checking that the code runs. 3. **Keep the key carrying the most records on purpose.** What one dominant key then does to a run, and what is done about it, is a different subject; here the rule is only that the cut must not quietly drop it. 4. **Add the rare shapes by searching for them,** one example of each, rather than hoping a draw contains them. 5. **Apply the same key set to every input the job joins,** so counterparts survive the cut. ## What even a good cut cannot tell you A cut is about *shape*, never about *volume*, and the honest half of this subject is naming what it leaves uncovered. Nothing about running out of memory shows up, and neither does **spill** - writing part of the working set to local disk because memory ran out. Neither does the cost of a **redistribution**, the point where a step needs records currently held by other workers: runtimes differ here, since some write the exchanged records out and have every worker fetch what is addressed to it, while others push records across the network as they are produced, so a cut understates a different cost depending on which you are on. Losing a worker mid-run and recovering from it is invisible too. A cut proves the logic is plausible; it never proves the job is operable.

  • Is a flat random one percent of rows a good enough cut instead?
    Better than the head, and still wrong in two ways. It cuts the dominant key down to ordinary size, so the uneven shape of the input disappears, and it misses rare shapes almost every time, because one percent of something that occurs once in a million rows is nothing. It also breaks joins, since the rows kept on each side rarely refer to each other. Cut by whole keys and add the rare shapes deliberately.
  • How small should a development cut be?
    Small enough that a full run finishes in seconds, because the value of the cut is the length of the edit-run-correct loop. Size is a consequence of that target, not a percentage chosen up front. Once the keys you must keep are fixed - the dominant one, the rare shapes, the counterparts - you sample the ordinary middle at whatever rate brings the total down to that runtime.
  • The job ran clean on the cut and failed on the real input. What is the first thing to check?
    Which property of the failing record the cut did not contain. Usually it is a shape rather than a volume: a null where the job assumed a value, an unexpected encoding, a key with one record, a duplicate. Add that example to the cut permanently, so the next run of the cut reproduces the failure, and then ask what else of that class the cut is still missing.

A jar filled in layers - fruit at the bottom, syrup, cream on top. Tasting the top spoonful tells you what was poured last, not what the jar tastes like. You have to draw from the whole depth to learn the mixture.

saying these in an interview costs you the question

  • A cut just needs to be big enough; size is the thing that matters
  • The first rows of a file are a random enough selection
  • If the job runs on the head of the file it will run on the whole file
  • Rare shapes can be ignored, they are a tiny fraction of rows
  • A cut built once is good forever, whatever the input does afterwards
  • A development cut also tells you how long the real run will take