skip to content

Testing and Changing Jobs

Proving a distributed job is right when its output is far too large to read, and changing that logic safely once people already depend on the numbers it produces.

on this pageshow

explore

questions

22

The first 100,000 rows of an input file are used as a development cut. What does that cut hide about the whole input?

level: juniorimportance: must knowfreq 56%

answer

  1. order of writing, not shape of content
  2. the head is one day or one writer
  3. rare shapes sit in the tail
  4. distribution, not row count
  5. shape yes, volume no

basics

~20 s

The first rows of a file are whatever was written first - one day, one source, one writer's share - so they carry neither the whole input's key distribution nor the rare record shapes that actually break the job.

solid answer

~50 s

A file's head is a slice of the order it was written in, not a slice of its content. Taking it gives you one day, one source system or one writer's output, so the mix of keys is wrong, the heaviest key is usually absent or ordinary-sized, and the rare shapes - a null in the join column, a negative amount, an unusual encoding, a key with a single record - sit somewhere in the tail you never read. The job then runs green on the cut and fails on the real input, and it fails on a case nobody chose to leave out. What you want instead is a cut whose shape resembles the whole: keys taken whole, the dominant key kept on purpose, and the rare shapes added deliberately rather than hoped for.

go deeper

for a junior

Recall that the head of a file reflects the order it was written in, not the content of the whole, so it carries the wrong mix of keys and none of the rare shapes.

for a middle

Explain the mechanics: which properties of the whole a contiguous slice cannot preserve, and why cutting by whole keys restores per-key totals that a row-level draw destroys.

for a senior

Show that you construct the cut deliberately - dominant key kept, rare shapes searched for, the same key set applied to every joined input - and that you say out loud what the cut still cannot prove.

for a principal

Frame it as a standing rule: what a shared development cut guarantees, how often it is rebuilt as the input drifts, and the fact that it is never evidence for a volume or runtime claim.

## What a development cut is, and what it is not A **development cut** is a small stand-in for a large input - a file, a table, or a day of records - reduced until a person can run the job over it in seconds and correct a change in one sitting. It is deliberately not the same object as a **statistical sample** used to estimate an aggregate cheaply. Those two pull in opposite directions: an estimator needs enough rows to bound its error, while a development cut needs to be as small as it can be *while still containing everything that makes the job behave interestingly*. Judging a cut by its row count is therefore the wrong measurement; the question is always **what shape did it keep**. The cheapest cut to produce is the head of the file: read the first N rows, stop. It is one command, it is fast, and it is almost always the wrong cut. ## Why the head of a file is a slice of the wrong dimension - **A file has a writing order, and that order is information about production, not about content.** Inputs are commonly written a day at a time, a source system at a time, or in the order records were extracted. The head is then one day, one source, or one extract - a slice along a dimension the job may not even model. - **Not every input is written in time order.** An input produced by many writers at once gives you a head that is one writer's share rather than the oldest rows. That is still a slice of the writing process rather than a slice of the content, so the conclusion does not change, but the reason does: do not assume the head is the past. - **Rare shapes live in the tail of a frequency distribution, not in a contiguous block.** A record shape that occurs once in a million rows has, in any single contiguous block, roughly the chance its rarity implies. You will not meet it, and it is exactly the shape that raises an error at three in the morning. - **Keys are spread through the file.** The key carrying the most records in the whole input contributes a handful of rows to the head, so in the cut it looks like any other key. Every property that depends on one key being much larger than the rest is gone. - **Counterparts in other inputs are accidental.** If you cut two joined inputs by taking each one's head, the keys in one head have very little to do with the keys in the other, and the join in your development run quietly returns almost nothing. ## What the head keeps and what it loses | property of the whole input | what the first rows give you | why it matters to the job | |---|---|---| | distribution of records per key | one contiguous slice, flattened | the run's behaviour on an uneven input is untested | | the key carrying the most records | present only by luck, at ordinary size | the case the job must survive is missing | | rare record shapes | almost never any | the error path is never exercised | | counterparts across joined inputs | accidental, usually poor | joins return far too few rows, or too many nulls | | value ranges and lengths | narrow, one period's worth | boundary handling untested | | repeats and duplicates of a key | under-represented | grouping and deduplication untested | ## What to take instead 1. **Aggregate the key frequencies once over the whole input.** One pass, one grouped count. This is the only expensive step and you do it rarely. 2. **Cut by whole keys, not by rows.** Choose a set of keys and keep every record belonging to them. Per-key totals in the cut then equal the real per-key totals for the keys kept, which makes the cut usable for checking a number and not only for checking that the code runs. 3. **Keep the key carrying the most records on purpose.** What one dominant key then does to a run, and what is done about it, is a different subject; here the rule is only that the cut must not quietly drop it. 4. **Add the rare shapes by searching for them,** one example of each, rather than hoping a draw contains them. 5. **Apply the same key set to every input the job joins,** so counterparts survive the cut. ## What even a good cut cannot tell you A cut is about *shape*, never about *volume*, and the honest half of this subject is naming what it leaves uncovered. Nothing about running out of memory shows up, and neither does **spill** - writing part of the working set to local disk because memory ran out. Neither does the cost of a **redistribution**, the point where a step needs records currently held by other workers: runtimes differ here, since some write the exchanged records out and have every worker fetch what is addressed to it, while others push records across the network as they are produced, so a cut understates a different cost depending on which you are on. Losing a worker mid-run and recovering from it is invisible too. A cut proves the logic is plausible; it never proves the job is operable.

  • Is a flat random one percent of rows a good enough cut instead?
    Better than the head, and still wrong in two ways. It cuts the dominant key down to ordinary size, so the uneven shape of the input disappears, and it misses rare shapes almost every time, because one percent of something that occurs once in a million rows is nothing. It also breaks joins, since the rows kept on each side rarely refer to each other. Cut by whole keys and add the rare shapes deliberately.
  • How small should a development cut be?
    Small enough that a full run finishes in seconds, because the value of the cut is the length of the edit-run-correct loop. Size is a consequence of that target, not a percentage chosen up front. Once the keys you must keep are fixed - the dominant one, the rare shapes, the counterparts - you sample the ordinary middle at whatever rate brings the total down to that runtime.
  • The job ran clean on the cut and failed on the real input. What is the first thing to check?
    Which property of the failing record the cut did not contain. Usually it is a shape rather than a volume: a null where the job assumed a value, an unexpected encoding, a key with one record, a duplicate. Add that example to the cut permanently, so the next run of the cut reproduces the failure, and then ask what else of that class the cut is still missing.

A jar filled in layers - fruit at the bottom, syrup, cream on top. Tasting the top spoonful tells you what was poured last, not what the jar tastes like. You have to draw from the whole depth to learn the mixture.

saying these in an interview costs you the question

  • A cut just needs to be big enough; size is the thing that matters
  • The first rows of a file are a random enough selection
  • If the job runs on the head of the file it will run on the whole file
  • Rare shapes can be ignored, they are a tiny fraction of rows
  • A cut built once is good forever, whatever the input does afterwards
  • A development cut also tells you how long the real run will take
open as a page

Which part of a distributed job can a plain unit test call without a cluster, and what stops it?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The per-record rule can: a plain function taking ordinary values and returning ordinary values. It stops being callable once its signature mentions a runtime type, or it reads configuration, storage or the clock itself instead of taking them as arguments.

open as a page

Why is sleeping in a test a poor way to prove a job groups records by when they happened?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Sleeping proves only that the machine's wall clock advanced. Instead hand the job records stamped with chosen moments and advance from the test the job's own claim that nothing older will arrive, so groups close on command.

open as a page

Why does comparing a distributed job's output line by line against a stored expected result fail even when every value is right?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A distributed run promises the right records, not an order, a file layout or bit-identical arithmetic. A line-by-line comparison asserts all three at once, so it fails on arrangement while every value is correct.

open as a page

What does running two versions of a live job's logic over the same input, with only one publishing, prove that a test cannot?

level: middleimportance: must knowfreq 62%

basics

~20 s

Running both versions over the same production input, with only one publishing, measures the change against real volume, real key distribution and real dirty records — evidence a test cannot give, because a test holds only the cases its author imagined.

open as a page

Two joined inputs are each cut to a random one percent independently, and the join comes back almost empty - why?

level: middleimportance: must knowfreq 51%

basics

~20 s

Each row survives on its own side only, so a kept row on one side keeps its counterpart on the other with about one percent probability - roughly one in ten thousand pairs survives. Draw one key set and restrict every input to it instead.

open as a page

Which totals do you reconcile against the input to judge a nightly aggregation job, and what does a match still miss?

level: middleimportance: must knowfreq 62%

basics

~20 s

Carry a small set of totals through the job: rows in against rows out with the relation the transform implies, the grand total of an additive measure, the distinct count of the grouping key, and the null and rejected counts. A match proves the job faithful to its input, not the number true.

open as a page

Two versions of a job's logic ran over the same input and thousands of grouping keys differ — how do you decide which differences are the change working?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Write down which keys the change should move, and why, before reading anything. Then classify every differing key against that prediction: the number that matters is how many moved for a reason you did not predict, not how many moved.

open as a page

Every lifted function is tested and the whole job runs green in one process - which defect classes stay invisible?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Four: anything appearing only when records are redistributed between workers; anything appearing only when a worker is lost and its share recomputed; anything appearing only when one key holds most of the records or the input is large; and anything appearing only when the function crosses a process boundary.

open as a page

A per-record function passes when the whole job runs in one test process but fails on a real cluster - why?

level: middleimportance: should knowfreq 54%

basics

~20 s

Because the function has to reach worker processes on other machines, and whatever it captured has to go with it. In one process the capture is the same object in the same memory; distributed, it must be transportable, rebuilt per worker, or already present there.

open as a page

Which parts of a job's logic cannot be lifted into a plain function, and what do you do about them?

level: middleimportance: should knowfreq 46%

basics

~20 s

Compositions do not lift: which records land in the same group, how many times a partial merge runs and where, what a join does with an unmatched key. The fragments inside them do lift, so keep the wiring free of business rules and test the fragments exhaustively.

open as a page

Which records, in what arrival order, would you hand a job to prove its grouping tolerates out-of-order input?

level: middleimportance: should knowfreq 48%

basics

~20 s

Feed records whose event moments deliberately do not ascend, and move the job's completeness claim in steps you choose: one advance that leaves the group open, one that closes it, and one record fed afterwards.

open as a page

When no independent expected value exists for an aggregate, which invariants and bounds still catch a wrong result, and which errors survive?

level: middleimportance: should knowfreq 55%

basics

~20 s

Assert what must be true of any correct result rather than the value itself: uniqueness at the declared grain, parts summing to the whole, keys present in the reference set, measures inside their physical range, and volume inside a band drawn from history. A plausible wrong number survives all of them.

open as a page

Where do you place the cutover between old and new logic so no published period is produced half by each version?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Cut over at a boundary in the output's own grouping — the start of a whole published period — not at the moment someone deploys. Record the first period the new version owns, stored with the numbers, and keep the old version runnable.

open as a page

When two versions of a job's logic run side by side over the same input and only one publishes, which effects of the silent arm still reach the outside world?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Everything the silent arm does besides writing its own output: a shared destination, an advanced reader position on the input, side effects fired from inside the logic, telemetry keyed by job name, and doubled load on shared capacity.

open as a page

How do you cut an input so that the dominant key and the rare record shapes both survive into the development sample?

level: seniorimportance: should knowfreq 39%

basics

~20 s

Measure key frequencies once, then build the cut in three deliberate parts: the heaviest keys kept whole or at a recorded fraction, a stable draw over the ordinary middle by whole key, and one searched-for example of every rare shape. Chance supplies none of these.

open as a page

Your test asserts a group's total is 7, but sees that group emitted twice with different values — what is wrong with the assertion?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Probably nothing is wrong with the job. Groupings differ in when they emit, and many legitimately publish early and revise. Assert the sequence of emissions and the settled value after a stated claim position, rather than one final number.

open as a page

How do you decide whether a 0.4% move between last night's published total and the previous run's is real or a defect?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Decompose the delta before judging it: attribute the move per key and per period, then check whether the input moved too. A few keys shifting is usually the source; every key shifting proportionally is usually the job. Explain the remainder rather than widening the tolerance.

open as a page

How do you diff last night's output against the previous run's when both hold billions of unordered rows at the same grain?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Match on the declared grain rather than on position, classify each key as only-previous, only-latest, differing or unchanged, and compare measures with an absolute floor plus a relative allowance. Aggregate per bucket first and descend only into the buckets that disagree.

open as a page

A correct fix makes today's figures disagree with every number already published — do you restate the history or fork the series?

level: principalimportance: should knowfreq 39%

basics

~20 s

Restate when the old figures were wrong and the past inputs can still reproduce them; fork with a labelled break when the definition changed rather than the answer being wrong. The deciding question is what readers compare across the seam.

open as a page

A crafted test of a job's time logic passes every run, yet production groups are wrong — which time-specific things did the test never exercise?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Three, mainly: where the event moment comes from on a real record, how the completeness claim behaves when several inputs feed it instead of one, and how wide real disorder is compared with the disorder somebody imagined when writing the fixture.

open as a page

A cut that keeps every counterpart across three chained joins grows toward the whole input - where do you stop, and what do you tell the team that absence in the cut means?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Close upward always and downward selectively: every kept row's parent must be present, while children are kept only where the job actually joins them. Then publish a manifest saying which joins are complete, so a thin result during development is read as the cut, not a defect.

open as a page