A suite's blocking pipeline invocation has a ten-minute budget but takes thirty-five; how do you restructure what each invocation promises?
answer
- Attribute the time before cutting it
- Fixed cost multiplies when you split
- One promise cannot serve two needs
- Deferred work needs an owner and a date
basics
~20 sFirst attribute the thirty-five minutes between fixed invocation cost and per-case cost. Then split one entry point into several with declared budgets and blocking rules, and make each enforce its own deadline rather than being killed.
solid answer
~50 sStart with attribution, not surgery. Measure how much of the thirty-five minutes is fixed cost the invocation pays before any case runs - process start, seeding, warm-up - and how much scales with the cases selected. Teams routinely reach for deleting cases when a third of the time is preamble nobody attributed, because a run's record times cases and not the invocation. Then change the contract rather than argue about it. One entry point can make only one promise, so declare several: a strict invocation inside the budget whose failures stop the stage, and a broader one whose failures are reported against the change without stopping it. Make each enforce its own deadline - a run that blows its budget should stop itself and return a budget-breach status with partial results, rather than being killed from outside with no evidence of where the time went.
code
yaml · 17 linesentry_points:
gate:
budget_minutes: 10
selection: "blocking"
on_budget_breach: "stop_and_report" # returns a breach status, not a kill
verdict: "blocks_the_stage"
broad:
budget_minutes: 45
selection: "everything_else"
on_budget_breach: "stop_and_report"
verdict: "reports_against_the_change"
owner: "payments"
response_within_hours: 24
shared:
fixed_setup_minutes: 6 # paid once per entry point - attack before splittinggo deeper
Know that a stage running a suite has a time limit people care about, and that the limit is a decision someone made rather than a property of the machine. Be able to say roughly how long your team's blocking run takes.
Explain how a run's time divides between fixed cost paid once per invocation and cost that scales with the cases selected. Expect to be asked what a run should do when it exceeds the time it was allotted.
Show that you would measure before restructuring, and describe how work moved out of a blocking run stays visible and owned. Be able to name the signals that a split moved risk rather than work.
Own the trade-off between one simple promise and several tuned ones: fewer entry points mean coarse budgets, more mean coverage nobody watches. Be ready to state the conditions - revert cost, catch rate, ownership, shared fixed cost - that decide which side you take.
A feedback budget is a promise about how long a stage's invocation of the suite may take before the waiting itself becomes the problem. When the invocation outgrows it - ten minutes promised, thirty-five delivered - the instinct is to start deleting cases. That is usually the third move, not the first. ## Attribute before you cut A run's duration is not one number. It divides into: - **Fixed cost paid once per invocation** - process start, dependency resolution, data seeding, warm-up, and any wait for a target to be ready. - **Cost that scales with the selection** - the cases themselves, plus whatever each one sets up and tears down. - **Cost that scales with contention** - waiting on shared resources, queueing behind neighbours, throttling. Most teams have never measured the split, because a run's record times *cases* and not the *invocation*: the results say the cases took nine minutes; nobody accounts for the other twenty-six. A third of a blocking run is often preamble. Fixing that is cheaper than any conversation about coverage, and it has to come first. ## The contract change: one promise becomes several One entry point can make only one promise. If a single invocation must be both fast enough to block a change and broad enough to be worth trusting, it will fail at one of the two. The structural move is to **declare several entry points, each with its own budget and verdict rule**: - a **gate** invocation with a strict selection and a hard budget, whose failures stop the stage; - a **broad** invocation with a longer budget, whose failures are reported against the change without stopping it. The pipeline wires both, but the difference between them is written in the suite, not in the stage. Note what this is *not*: not a decision about which trigger each one runs on, and not the mechanics of dividing a run across parallel workers. This is about what each invocation promises and what its result may mean. ## Enforce the budget inside the entry point A budget the entry point does not know about is enforced by something killing the job from outside, producing the worst possible artefact: no results, no attribution, a rerun. Give the entry point its own deadline instead: as it approaches, the run stops itself and returns a **distinct budget-breach status** carrying partial results, the list of what had not run, and phase timings. A breach is a *result*. It tells you whether to buy hardware, cut cases or move the line; a silent kill tells you nothing. ## The trade-off | | One blocking entry point | Several with different promises | | --- | --- | --- | | What the pipeline wires | One command, one meaning | Several, each with a budget and a verdict | | Wait before a change advances | The whole suite | The gate selection only | | Coverage at the merge point | Everything the suite has | Only what the gate selects | | Who owns a failure | Whoever is blocked by it | Must be assigned explicitly | | Characteristic failure | The gate is overridden or skipped | The broad set turns red and nobody looks | | Total machine time | Lower | Higher, by the duplicated fixed cost | Four conditions decide which side you take: 1. **Revert cost.** If a bad change is cheap to unwind, deferring verification is a reasonable bet. If unwinding is expensive - data written, notifications sent, a release cut - more stays inside the budget. 2. **Catch rate of the deferred set.** Measure how often the broad invocation finds something the gate did not. Never, over months, means it is cost rather than insurance. Often means the line is drawn in the wrong place. 3. **Ownership.** A non-blocking invocation with no named owner and no stated response time is not deferred coverage; it is deleted coverage with a report nobody reads. 4. **Shared fixed cost.** This is why attribution comes first. If every invocation pays the same six-minute preamble, splitting one run into two adds six minutes of pure overhead and may raise total machine time while lowering the gate's latency. ## What moves, what is cut, and what is never on the table Whatever leaves the blocking invocation takes three things with it: an owner, a response time, and visibility at the place the change is reviewed. Some work should be cut rather than moved: coverage duplicated more cheaply at a lower level, and cases whose failures nobody has acted on in a year. Which specific cases to delete is a separate decision with its own criteria; the budget only says how much has to go. What is never on the table is meeting the budget by weakening what a result means: letting the gate's failures stop blocking, or reporting success on a run that did not finish. That converts a latency problem into an evidence problem, which surfaces later and costs more. ## Signs the split went wrong - Budget met but overrides rising - you moved risk, not work. - The broad invocation has been red for three weeks - the deferral had no owner. - Total pipeline time went up - duplicated fixed cost, unmeasured. - Nobody can state what the gate covers - the promise was never written down.
- You split the suite into a gate and a broader invocation, and total pipeline time went up. Why?Almost always shared fixed cost. Each entry point pays its own process start, seeding and warm-up, so a six-minute preamble paid twice adds six minutes of pure overhead before a single extra case moves. Attack the fixed cost first - share prepared data and warm caches where it is safe, or make the preamble proportional to the selection - then decide whether the split still earns its keep.
- How do you know whether the deferred set is earning its keep?Count what it catches that the blocking set did not, over a real period. A deferred invocation that has not found a unique failure in months is not insurance, it is cost, and it should be cut and said so out loud. One that catches regularly is telling you the budget line is drawn in the wrong place and some of that work should be paid before the change advances.
- What should happen when a run exceeds its declared budget?It should stop itself and return a distinct budget-breach status with whatever results it has, rather than being killed from outside with no record. A breach is a result: it names which cases had not run and how the time was spent, so the next decision is informed. Silent kills produce reruns and a suite whose cost nobody can attribute.
saying these in an interview costs you the question
- Cuts cases before measuring where the time actually goes
- Splits the suite without noticing setup cost is paid twice
- Defers work to an invocation with no named owner
- Meets the budget by making failures stop blocking
- Treats a budget breach as a rerun rather than a result