The nightly rollup's runtime tripled after a one-line change to its description - how do you find out why?
answer
- the text did not get bigger
- ask the engine, not the source
- diff the plan, both versions
- counters over wall clock
- which freedom did the line remove
basics
~20 sYou cannot read the cause off the description; you ask the engine what it did. Compare the execution plan and the work counters before and after the change, then identify which freedom the new line took away from the planner.
solid answer
~50 sReading the text harder will not answer it - the text is one line longer and the work is three times bigger, which is the whole leak. Capture the plan the engine produced for both versions and diff them, and compare work counters - records and bytes read, partitions, spilled intermediates - rather than wall clock, which only tells you something got slower. Then ask what the new line denied the planner: a stage it can no longer see through, a grouping key that widened and exploded the group count or collapsed the parallel width, a step that forces the whole input to be materialised, or a narrowing stage it can no longer perform early. The fix is usually to restate the intent in terms the engine recognises, and then to assert on the plan so the regression is caught next time rather than on the bill.
code
pseudocode · 12 linesbefore:
orders.keep(recent)
.project(function(o) return pair(o.region, o.total))
.groupAndSum(by = region)
after:
orders.keep(recent)
.project(function(o) return pair(lookupRegion(o.customerId), o.total))
.groupAndSum(by = region)
// one line longer; lookupRegion() is a remote call.
// one call per surviving record, and a stage the planner cannot cost.go deeper
The takeaway to carry away is simply that you cannot tell what a description costs by reading it. When something gets slower, the engine has a record of what it did, and that record is where you start.
Be able to describe the diff: two plans, compared structurally, plus work counters rather than durations. Know the usual culprits - an opaque stage, a widened key, a forced materialisation.
Show the discipline of ruling out the non-edit causes first - input volume, data skew, a re-plan, an engine upgrade - and of expressing the cause as a freedom the planner lost rather than as a slow line of code.
Make it systemic: the reason this hurt is that the plan was invisible until it was expensive. Argue for plans as reviewed artifacts and budgets denominated in work, so cost regressions land in review rather than in the invoice.
## Why the description cannot tell you This is the failure the intent-first style exists to be honest about. When you write the mechanism, the work is on the page and a reviewer can count it. When you write the intent, the work performed is produced by the engine from three inputs you did not write: the description, what the engine knows about the shape of the data, and the engine's own version and heuristics. A one-line edit can move any of the three, and the text gives no signal about the size of the move. Cost is not proportional to stage count, line count, or how simple a phrase reads. So the first thing a senior engineer does here is stop reading the source and start reading the run. ## The method 1. **Confirm the input did not change.** Before anything else, check the record count and byte volume of the source for both nights. A tripled runtime on tripled input is not a regression in the description. 2. **Get the plan for both versions.** Whatever form the engine exposes - a printable plan, a stage graph, a description of what it decided - capture it for the old description and for the new one, over comparable data. 3. **Diff the plans, not the durations.** You are looking for a structural difference: a stage that used to be performed with its neighbour and now is not, a partition count that fell, an ordering step that appeared, an intermediate that is now written down. 4. **Compare work counters rather than wall clock.** Records read, bytes moved between stages, spilled intermediates, per-partition skew. Wall clock on a shared machine tells you that something is slower; counters tell you what got bigger. 5. **Name the freedom that was lost.** Every cost regression of this kind reduces to the planner being able to do less than before. ## What usually caused it | Symptom in the plan | Likely cause | |---|---| | A stage is no longer performed with its neighbour | the new line is opaque to the engine, so it must stand alone | | Partition count collapsed | the split key now has far fewer distinct values | | Group count exploded | the grouping key widened, so each group is tiny and there are many | | An intermediate is now written down | a step was introduced that needs the whole input at once | | A narrowing stage moved later | the engine can no longer establish that moving it ahead is safe | The most common of these is the first, and it is the one that looks most innocent in review: a short phrase that calls out to something the engine cannot inspect. Three words of text, one remote call per surviving record, and a plan that now treats the middle of the pipeline as a wall. A second family has nothing to do with the edit at all, and you must rule it out before blaming the diff: the engine may simply have re-planned. Data shape moves, the statistics behind the plan move with it, and the same description is executed by a different strategy tonight. An engine upgrade does the same thing. A candidate who insists that cost cannot change without a code change has not operated one of these systems. ## Fixing it, and keeping it fixed The repair is usually not cleverness but **restating the intent in terms the engine recognises**, so that the thing you meant is visible to the planner instead of hidden inside something it must treat as opaque. Where an engine offers a way to state a preference about strategy, that is the second lever, and it should be used knowingly: every pinned decision is optimiser freedom you have given back, and it will age as the data changes. The durable part is making the regression visible before the bill is. Two habits do most of the work: - **Treat the plan as an artifact.** Capture it on every run and assert on the properties you care about: the stage count, the partition width, whether the narrowing stage still happens early. A plan diff in review is far cheaper than a plan diff at 3 a.m. - **Budget in work, not time.** Records read and bytes moved are stable across a noisy machine; wall clock is not. A threshold on work done catches the regression on the first night rather than the first invoice. ## What the interviewer is listening for Three things, in order. First, that you know the description cannot answer the question - that is the whole point of the leaf. Second, that your first move is to obtain what the engine actually decided, in both states, and diff it. Third, that you express the cause as *lost freedom*: the planner used to be allowed to do something, and one line took it away. Anyone who goes straight to profiling the host, or who declares the engine buggy without comparing two plans, is guessing.
- What could triple the cost with nobody changing the description at all?The data's shape: a key that grew, skew that appeared, a source that got slower. Or the engine's picture of that data changed and it re-planned - a different split, a different grouping algorithm. An engine upgrade does the same. The description is stable; the plan is not, and that is by design.
- How do you stop the same regression recurring?Capture the plan as an artifact and assert on its properties - stage count, partition width, whether the narrowing stage still happens early - and budget the job in work counters rather than wall clock. Then a plan-changing edit fails in review, and a plan-changing data shift alerts on the first night instead of on the invoice.
saying these in an interview costs you the question
- Estimates cost by counting lines or stages in the description
- Assumes tripled runtime must mean tripled input
- Says an engine's chosen strategy cannot change without a code change
- Profiles the host machine before ever looking at the plan
- Declares the engine buggy without comparing the two plans
- Compares wall clock on a shared machine and calls it evidence