skip to content

Across dozens of long-running jobs, what standing rule would make each team predict its retained set before first deployment?

level: principalimportance: should knowfreq 38%

answer

  1. prediction before deployment, written down
  2. one row per held category
  3. counts as products, not totals
  4. cardinality needs a source and a date
  5. re-declare on re-key or horizon change

basics

~20 s

Require a written retained-set budget per job before first deployment: entry categories, live key cardinality with its source, the horizon per category, bytes per entry, and the product. Fix the units so numbers compare, and require re-declaration when the grouping key or horizon changes.

solid answer

~50 s

Make the estimate a required artifact, not a habit. Each job declares, before it first runs, one line per held category — running accumulators, records buffered awaiting a group, join sides, deduplication entries, pending per-key wake-ups — with the live key cardinality and where that number came from, the retention horizon it assumes, bytes per entry, and the resulting total. Fix the definitions centrally so numbers from different teams mean the same thing: live means not yet removed rather than seen recently, and the horizon is stated next to every cardinality. Set an escalation threshold above which the design gets a second reviewer; the value is an organisational choice, but having one written down is not. Then tie re-declaration to the two events that invalidate the number — a change of grouping key, and a change of horizon. The rule buys prediction, nothing else: which store holds the entries, and what removes them, are separate decisions with their own owners.

go deeper

for a junior

Recall that the size of what a job holds is something predicted before it runs, and that the prediction rests on distinct key counts rather than on how fast records arrive.

for a middle

Be able to fill the table for your own job: one row per held category, each count written as a product, with the horizon and the cardinality source stated beside it.

for a senior

Defend the factors, not the total. Source the cardinality, keep total retained bytes separate from the working set, and update the declaration in the same change that re-keys the job or extends a horizon.

for a principal

Own the definitions and the escalation threshold across teams, resist folding store choice and removal policy into this rule, and measure the rule by whether wrong numbers become traceable rather than by whether they stop happening.

## What the rule is actually buying One team getting a size estimate wrong is a production incident. A fleet getting it wrong is a pattern, and patterns are only fixable with a standing rule. What you want from the rule is a **prediction made before the first deployment, by the people who know the data, in a form a reviewer who does not know the data can check**. Everything in the rule's design follows from that sentence. It is worth being honest about what the rule does not buy. It does not decide where the entries sit, what removes them, or how much memory a worker gets — those are separate decisions with their own owners, and a rule that quietly absorbs them will be argued about instead of followed. ## The artifact Require one table per job, one row per held category, with these columns: 1. **Category** — running accumulator, records buffered awaiting a group, join side, deduplication entry, pending per-key wake-up. Naming the category forces the team to notice the ones that are not accumulators. 2. **Entry count, written as a product** — not a number, a product of quantities, so a reviewer can attack a factor rather than a total. 3. **Cardinality source** — the system and query the distinct-key number came from, with its date. A cardinality with no source is the single most common reason these estimates are wrong. 4. **Horizon** — how long an entry stays before anything removes it, stated per category rather than once for the job. 5. **Payload bytes per entry**, and a **per-entry overhead** line treated as a range until measured. 6. **Twelve-month figure** — the same arithmetic at expected cardinality growth, because cardinality is what moves this number, not traffic. ## Fix the definitions or the numbers will not compare A fleet-level rule fails quietly when two teams use one word differently. Three definitions have to be central: - **Live keys**, meaning keys whose entries have not yet been removed — *not* keys seen in a recent interval. The difference is often an order of magnitude. - **Total retained bytes**, meaning everything held across the job, kept strictly separate from the subset touched in a given interval. They answer different questions and must never be reported in one column. - **Bytes**, meaning a stated unit with a stated base, because a fleet dashboard that mixes two conventions is worse than no dashboard. ## The threshold and the second reviewer Set a total above which a job's design gets a second reviewer before it ships. Do not look for a universal number — what is unremarkable for one organisation's hardware and restart expectations is alarming for another's, so the value is a local choice. What matters is that it is written down, that it is expressed in total retained bytes rather than in a vague sense of "large", and that crossing it triggers a conversation rather than an automatic refusal. The conversation is where the decisions this rule deliberately excludes get made. ## Keeping it from rotting A declaration made once and never revisited is worse than none, because it carries authority it no longer earns. Two events invalidate the number, and both are visible in a code review: - **The grouping key changes.** A re-key changes the key population outright, so every cardinality in the table is stale. - **A horizon changes.** Extending how long entries are kept multiplies the categories that are counted per record. Require the table to be updated in the same change that does either, and require the twelve-month line to be recomputed annually against measured cardinality rather than the original estimate. ## What good looks like when it lands | Before the rule | After the rule | |---|---| | "It is a streaming job, it should be fine" | a product of sourced quantities per category | | Discovered when a restart takes hours | predicted before the first deployment | | Each team invents its own units | one definition of live, total and bytes | | Cardinality growth noticed after it bites | a twelve-month line recomputed each year | The measure of the rule is not that every number turns out right. It is that when a number turns out wrong, the table says which factor was wrong and who sourced it — and the next job's estimate is better because of it.

  • A team cannot source its distinct key cardinality. What should the rule say?
    That the job may not ship on a guess. Accept a measured cardinality from a replayed slice of real input over the full horizon, or accept an explicit upper bound with the reasoning for it. An unsourced number is the factor most likely to be wrong by an order of magnitude, and it is the cheapest one to go and measure.
  • Why require the count as a product rather than a total?
    Because a total cannot be reviewed. A product exposes the assumptions — cardinality, horizon, bytes per entry — so a reviewer who knows the domain can challenge one factor without redoing the estimate, and so a later miss can be attributed to the specific assumption that failed rather than to the estimate as a whole.
  • Does the rule apply to jobs over finite input?
    Mostly not, and saying so keeps the rule credible. A run over a bounded input holds entries only for its own duration, and the two-phase disk-to-disk batch model keeps nothing between runs at all. Apply the rule to long-running jobs whose entries outlive a single run, and exempt the rest explicitly rather than by neglect.

saying these in an interview costs you the question

  • Relies on teams remembering to estimate
  • Accepts a total with no factors behind it
  • Quotes a universal byte threshold as if it were a standard
  • Lets each team define live keys its own way
  • Reports total retained bytes and touched working set as one number
  • Treats the declaration as permanent after the first deployment