Your platform can change the worker count of a job over endless input on demand. What makes an aggressive change policy cost more than it saves?
answer
- saving is capacity times duration
- cost is a pause, sized by state
- the two are unrelated
- reaction interval longer than the pause
- dwell, hysteresis, cap, schedule
basics
~20 sEvery change stops the job while its retained state is relocated, and that cost follows how much the job holds rather than how much capacity moves. A policy that reacts faster than the pause is long spends its savings on pauses and provokes further changes.
solid answer
~50 sThe saving from a width change is capacity returned multiplied by the time it stays returned. The cost is a pause whose length follows the volume of retained state, plus the time spent recovering the ground lost during it. Those two quantities are unrelated, so a policy that reacts to every movement in load can pay a forty-second pause to release capacity it takes back three minutes later — and each pause disturbs the very signal it is reacting to, which is how oscillation starts. Sane policies damp it: a minimum dwell between changes, different thresholds for widening and narrowing, widen quickly and narrow slowly, a cap on changes per period, and a fixed width through known peaks. The workload that earns this is one holding little state with a large, predictable swing in load.
go deeper
Recall that changing a running job's worker count is not free, so doing it often can cost more than the capacity it hands back.
Explain why several small changes cost more than one large one: each relocates a large share of the retained state, so the cost follows the number of changes, not their size.
Show how a reactive policy oscillates — the pause makes the job look busier, which triggers the next change — and name dwell time and separate up and down thresholds as the damping.
Set the policy: decide which workloads may change width at all, from retained volume and load profile, and accept a fixed width sized for the peak where the pause would breach the latency budget.
## The two quantities a policy is trading A **mid-run capacity change** is a running job acquiring worker processes or handing some back. For a job over endless input, the ledger has exactly two sides: - **The saving** — worker processes released, multiplied by how long they stay released. Measured in machine-hours. (What that converts to on a bill, and whose budget it lands on, is a different subject.) - **The cost** — a pause while the job's retained per-key state is relocated to match the new worker count, plus the time afterwards spent working through what arrived while it was stopped, plus the startup of any new processes. The two sides have nothing to do with each other, which is why an aggressive policy can be confidently, arithmetically wrong. ## Why the cost does not scale with the size of the change Ownership of keys is computed from the worker count, so almost any change in that count reassigns a large share of keys. Moving from eight worker processes to nine can relocate nearly as much state as doubling. The consequences for policy: - **Few large changes beat many small ones.** Ten increments of one cost roughly ten pauses; one increment of ten costs roughly one. - **Narrowing is not cheaper than widening.** Both relocate state; a policy that treats release as free because it involves fewer machines has the sign of the cost wrong. - **The cost is a property of the job, not of the moment.** A job retaining very little can be adjusted almost freely. A job retaining a great deal has an expensive move whenever it is asked for, and no amount of scheduling makes an individual move cheaper. ## How oscillation starts This is the failure that turns a reasonable idea into a worse system than a fixed width: 1. Load rises; the policy asks for more worker processes. 2. The job pauses to relocate state, and during the pause it processes nothing. 3. When it resumes, the work that arrived during the pause makes the job look even busier. 4. The policy reads that as more load and changes width again. Each round costs a pause and makes the reading worse. The rule of thumb worth stating in an interview: **the reaction interval must be much longer than the pause**, and if it cannot be, the width should be fixed. ## The levers that damp it - **Minimum dwell** — no change within N minutes of the last one, chosen from the observed pause length rather than a round number. - **Different thresholds each way** — widen at one level of utilisation, narrow at a distinctly lower one, so the job does not sit on a boundary flipping. - **Widen fast, narrow slowly** — being briefly over-provisioned is cheap; being under-provisioned on a latency-sensitive job is not. - **A cap on changes per period** — a blunt but effective bound on total pause time, and a useful safety net over any signal. - **Scheduled width instead of reactive width** — for a load profile that is genuinely diurnal, a planned change at two known times a day captures most of the saving with two pauses. - **A floor and a ceiling** — the policy should never narrow below what the job needs to keep up at its quietest, nor widen beyond what the pool can grant. ## Which workloads earn it | Workload | Verdict | Why | |---|---|---| | Little retained state, large day-night swing | Good candidate | Cheap move, big and predictable saving | | Large retained state, flat load | Fix the width | Every move is expensive and there is nothing to chase | | Latency budget tighter than one pause | Fix the width at the peak | Any change breaches the budget by itself | | Unpredictable, spiky load | Scheduled or manual only | A reactive policy will chase the spikes and oscillate | ## What varies between engines The policy that is right depends on a mechanism the platform does not control. Where a runtime relocates state in place, a change is a pause of seconds to minutes and reactive policies are at least arguable. Where a runtime accepts a new width only when the job is brought back up, every change is far more expensive, and the only defensible policy is a small number of planned changes. A design that runs endless input as a rapid succession of small jobs with an end has natural boundaries and, if it keeps little across them, can change width far more often than the others. An answer that names one of these as the way the class works has picked an engine; an answer that says which one the platform is on, and sets the policy from that, is the one being asked for.
- Which levers actually damp the oscillation?A minimum dwell between changes derived from the measured pause, different thresholds for widening and narrowing, widening faster than narrowing, a cap on changes per period, and a width held fixed through known peaks. Each buys fewer pauses with some idle capacity, which is almost always the right side of the trade.
- Which jobs should simply be left at a fixed width?Those retaining a lot, because every change is a long pause; those whose latency budget is tighter than one pause; and those whose load barely moves, because there is nothing to capture. Size those for the peak and spend the review effort on jobs with a genuine swing.
- Does a scheduled change avoid the cost?No — it pays the same pause, but at a time you choose, with the number of pauses per day known in advance. For a diurnal profile that captures most of the saving with two disruptions, and it removes the oscillation risk entirely because no signal is in the loop.
saying these in an interview costs you the question
- Assumes narrowing is cheap because it hands machines back.
- Thinks smaller, more frequent changes are gentler than one large one.
- Treats capacity matching as valuable even while the job is paused.
- Ignores that a pause distorts the signal the policy reacts to.
- Applies one reactive policy to every job regardless of retained state.
- Believes a width change reduces how much the job retains.