Your platform standardises nightly batch work on descriptions a planner executes - how do you keep cost accountable?
answer
- the plan is an artifact
- budget work, not wall clock
- opaque stages declare their cost
- diff plans in review, not at 3 a.m.
- pinning is narrow, owned and dated
basics
~20 sMake the plan a reviewed artifact rather than a runtime surprise: budget each job in work done, not wall clock; capture and diff plans automatically; require opaque stages to declare their cost; and decide which jobs may pin a strategy.
solid answer
~40 sThe bet the platform made is that jobs describe intent and the engine picks the strategy, which buys portability and improvement over time at the price of a cost model nobody can read off the source. Accountability means restoring visibility without taking the freedom back. Capture the plan on every run and diff it, so a strategy change shows up as a reviewable event rather than an invoice. Budget jobs in units the engine reports - records read, bytes moved, partitions - because wall clock on shared capacity is noise. Require a stage the engine cannot see through to carry a declared per-record cost, since those are where the surprises hide. And write down which jobs are allowed to pin a strategy, who approves it, and when that pin is revisited.
go deeper
The idea to take away is that somebody has to own the cost of a job whose strategy is chosen by a machine, and that the numbers to watch are about work done, not minutes elapsed.
Be able to argue why work counters beat wall clock on shared capacity, and why a captured plan is worth storing next to each run rather than printed once during debugging.
Show the operational form: budgets on counters with a named response, plan assertions in the pipeline, and a habit of ruling out data shape before blaming an edit.
Own the trade explicitly. Say what the standardisation bought, what visibility you add to make the cost legible, and where you will accept less optimiser freedom - narrowly, with an owner and a review date - rather than everywhere or nowhere.
## The bet the platform has already made Standardising on descriptions is a deliberate trade. The organisation gains that one job description outlives engine versions and substrate moves, that improvements to the planner lift every job at once, and that teams write what they mean rather than how to compute it. It loses the property that made the old jobs governable: you could read a step-by-step job and know roughly what it would cost. Nobody can do that with a description, because the work is decided later, from data the author never saw. So the lead's problem is not whether the bet was right. It is how to make cost legible again **without** taking back the freedom that justified the bet. ## Make the plan an artifact The root cause of every ugly surprise here is the same: the plan is invisible until it is expensive. Treat it as an output of the job, like a schema or a test report. - Capture the plan on every run and store it with the run. - Assert on the properties that matter rather than on the whole text: the number of stages, the parallel width, whether the narrowing step still happens early, whether an intermediate is materialised. - Surface the diff in review when a change alters the plan, so an edit that quietly turns the middle of a pipeline into a wall is caught by a human who can ask why. This is not the same as freezing the plan. The plan is *allowed* to change - it changes as data moves, which is the mechanism working. What is not allowed is for it to change silently. ## Budget in work, not in time Wall clock is the number everyone reaches for and the worst one to govern with: it moves with cluster load, retries and neighbours. Denominate budgets in what the engine reports about the work itself - records read, bytes moved between stages, partitions used, intermediates spilled. Those numbers are stable enough that a threshold means something, and they are the numbers that track the bill. A budget is only enforceable if it is: expressed in units the engine actually emits, attached to a named owner, checked on every run, and paired with a stated response when it is breached. A budget with no response is a dashboard. ## Make opacity visible The recurring failure is a stage the planner cannot see through: it cannot cost it, cannot execute it alongside its neighbour, and cannot move work across it. A platform can require those stages to announce themselves - a declared per-record cost, a note on whether the stage performs I/O - and can report how much of each job's pipeline is opaque. A job that is mostly opaque has all the unpredictability of a description and none of the optimisation benefit, and the platform should be able to say so without anyone reading the source. ## The escape-hatch policy Every serious engine offers some way to express a preference about strategy, and the question is not whether to allow it but who may, and for how long. | Control | What it buys | What it costs | |---|---|---| | Plan captured and asserted | change becomes reviewable | a gate to maintain, some false alarms | | Budget in work counters | regressions found on night one | someone must own the thresholds | | Declared cost on opaque stages | surprises become visible in review | friction on every such stage | | Pinned strategy | variance removed for that job | optimiser freedom given back; ages with the data | The defensible policy is narrow and dated: pinning is for the handful of jobs where **variance itself is the risk** - a hard external deadline, a contractual cost ceiling - it is approved by a named owner, and it carries a review date, because a pin is a decision taken against a snapshot of the data that will not hold. Letting every team pin by default dissolves the reason the platform standardised on descriptions at all. ## Where this bet is the wrong one A principal should also be able to say where not to make it. If a job must finish by a fixed hour and a re-plan could miss it, predictability is worth more than portability. If the data's shape is stable and small and one hand-chosen strategy is obviously right, the optimiser has nothing to contribute. And if nobody on the team can read a plan, the platform has handed out a tool whose failure mode it cannot diagnose - which is an argument for training and tooling before adoption, not for abandoning the approach. The answer that lands is the one that refuses both extremes: not banning descriptions because their cost is hard to predict, and not shrugging at variance because the engine knows best, but making the engine's decisions observable and budgeted so the freedom keeps paying for itself.
- Where is predictability worth more than the optimiser's freedom?Where variance itself is the risk: a job with a hard external deadline, or a contractual cost ceiling, where a re-plan that is usually faster but occasionally twice as slow is unacceptable. Also where the data is small and stable, so there is nothing for a planner to exploit. Elsewhere, keep the freedom - the data moves and a pinned strategy ages.
- What makes a cost budget enforceable rather than decorative?Four things: it is expressed in units the engine actually reports, it is attached to a named owner, it is evaluated on every run rather than reviewed monthly, and a breach has a stated response - block, page or file. A threshold with no owner and no response is a chart nobody reads.
saying these in an interview costs you the question
- Bans descriptions outright because their cost is hard to predict
- Governs only by wall clock, which hides what work was done
- Lets every team pin a strategy, discarding the optimiser's value
- Treats any plan change as an incident rather than expected behaviour
- Assumes a cost budget set once stays valid as data grows
- Buys a planner-based platform with nobody able to read a plan