skip to content

You add a row count and a distinct-key count either side of ten steps and the run slows badly — which of those claims is free and which is not?

level: seniorimportance: should knowfreq 44%

answer

  1. not every claim is free
  2. a row count is bookkeeping
  3. uniqueness is a full pass
  4. a deferred plan runs per question

basics

~20 s

On a materialised in-memory table the row count is bookkeeping the object maintains and costs nothing, while a distinct-key count must read every key value, so it is a full pass. On a deferred pipeline even the row count forces the plan to run.

solid answer

~50 s

The two claims look alike and cost nothing alike. A materialised table already knows how many rows it holds, so asserting the row count is free; the number of distinct combinations of the key columns is not kept anywhere, so it has to be derived from every value, costing a pass over those columns and memory in proportion to the number of distinct keys. The bigger factor is the evaluation model. Under deferred evaluation — where operations build a plan and nothing computes until a result is demanded — asking for a count is a demand, and the next count can be another one, so ten claims can mean ten executions. The same is true of a reader walking a file in chunks, where a count means another pass over the input. So: keep the free claims everywhere, put the expensive ones at boundaries, and collect the numbers you need in as few passes as you can.

go deeper

for a junior

Take away the distinction rather than the detail: asking how many rows a table holds is usually free, while asking how many distinct key values it holds means reading them all.

for a middle

Explain why one number is maintained by the object and the other has to be derived, and mention the memory the distinct set costs on a large, nearly-unique key.

for a senior

Show that you know the evaluation model decides the cost: on a deferred pipeline each count is a demand that can run the plan again, so the claims are collected into as few passes as possible and placed at boundaries rather than between every step.

for a principal

The call is a budget: how many passes over the data the team spends on evidence, and where. Expensive claims scattered everywhere get the whole practice deleted after the first complaint about run time.

## Two claims that look alike and are not Both of these read like one line of arithmetic in a file: - `row_count(t) == expected` - `distinct_count(t, keys) == row_count(t)` The first asks the table for a number it already keeps. The second asks a question nobody has answered yet, and the only way to answer it is to look at every value in those columns and remember what has been seen. | claim | on a materialised in-memory table | on a deferred pipeline or a chunked read | |---|---|---| | row and column counts | bookkeeping the object maintains; effectively free | a demand that forces the plan to execute, or another pass over the file | | a total over a column | one pass over that column | one execution, fused with whatever else the plan does in that pass | | distinct count of the key columns | one pass, plus memory proportional to the number of distinct keys | one execution, and the same memory for the distinct set | ## Metadata against measurement Calling the row count *metadata* is not loose talk: a materialised table maintains its size as part of being a table, so reading it is a lookup rather than a computation. Nothing maintains the number of distinct key combinations, because it changes with every write and nobody would pay to keep it current. That is the real dividing line between the two claims, and it is why one can be sprinkled everywhere and the other has to be placed. The distinct count also costs **memory**, not just time. Establishing that no combination repeats means holding the combinations seen so far. On a key with a hundred distinct values that is nothing; on a key that is nearly unique across a very large table it is another structure roughly the size of the key columns. ## Deferred evaluation changes the arithmetic completely Under **deferred evaluation** — a design where each operation records what to do and nothing is computed until something asks for a result — a count is not a question, it is a trigger. Two consequences follow, and both surprise people: 1. **Each claim can run the work.** Nothing obliges an unmaterialised plan to keep intermediate results for the next question, so a count between each of ten steps can mean ten executions of a growing plan rather than one. 2. **The claim can change what is measured.** A plan executed to answer a count may not be the same physical work the final result triggers, so the number is about a run, not about an object sitting in memory. The same applies to a reader that walks a file in chunks to stay inside memory: a count over the whole input is another pass over the file. ## The form the check is written in One widely repeated sentence needs qualifying. *"A check written as a function handed to the table is called once per record, so it costs more than the step it guards"* is true of the **per-record callback** form, where the tool really does invoke your function for every record and pays the host language's call cost each time. It is false of the same surface applied a **whole column at a time**, which several tools offer under the same name and which runs at full speed — and it is beside the point for most claims in this leaf, which are plain comparisons and aggregates over a column and never involve a callback at all. State the form before you state the cost. ## Budgeting the claims A workable policy, in descending order of cheapness: - **Keep the free claims everywhere.** On materialised tables, counts either side of a step are not the thing that slowed your run down. - **Place the expensive claims at boundaries.** A uniqueness claim belongs where an input is loaded and where a result is published, not between each of ten intermediate steps that cannot affect it. - **Collect, then assert.** If several numbers are needed from one deferred pipeline, compute them together as one result rather than as separate demands. - **Assert at the widest point, not the narrowest.** Checking a key on the small reference table is a fraction of the cost of checking the large one, and it is usually the reference table whose uniqueness the code assumes. - **Do not downgrade a claim to a sample.** Checking uniqueness on a sample is not a weaker version of the claim, it is a different and much less useful statement: one repeated key in a million rows breaks the match just as thoroughly and will not be in your sample. Use a sample as a smoke test if you like, but do not let it retire the claim. The point of knowing the costs is not to write fewer claims. It is to stop the expensive ones from discrediting the cheap ones, because when evidence gets a reputation for slowing the run down, the first thing that gets deleted is the free count that would have caught the defect.

  • How do you keep the evidence when every claim on a deferred pipeline costs an execution?
    Collect rather than sprinkle: compute the numbers you need as one result in a single pass and assert against them afterwards, and put the remaining claims at the points that justify an execution — where an input is loaded and where a result is published. A claim between two steps that cannot affect it buys nothing.
  • Is checking uniqueness on a sample an acceptable substitute?
    No, and it is worth being explicit about why: a repeated key that occurs once in a million rows damages the match exactly as much as one that occurs everywhere, and it will not be in the sample. A sample is a smoke test during development, never the claim the file relies on.

saying these in an interview costs you the question

  • Assumes any count is free because tables know their size
  • Thinks a uniqueness claim costs the same as a row count
  • Forgets a count makes a deferred plan run, and run again
  • Puts a distinct-key claim either side of every step
  • Assumes any check handed a function is slow whatever its form
  • Samples the key columns and calls it an assertion