You own an autoscaling pool of single-use CI runners. How do you decide between scaling to zero and keeping warm capacity, and what would you measure to know the choice was right?
answer
- idle cost versus wait cost
- price both in the same currency
- queue wait, not total duration
- fix cold start before buying warmth
- every extra pool has its own floor
basics
~20 sTrade idle machine cost against developer wait time, and price both. Measure queue wait separately from job duration, at high percentiles, alongside provisioning latency. Fix slow provisioning before buying warm capacity, then size a minimum pool to absorb the usual arrival burst.
solid answer
~60 sFrame it as two costs in the same currency. Idle capacity is money you can read off an invoice; queue time is engineers waiting, multiplied by every run, and it compounds because a slow pipeline changes how people batch work. Measure them separately: **queue wait** — enqueue to first step — at p50 and p95, never mixed into total job duration, plus provisioning latency broken into its parts (instance boot, image pull, agent registration, checkout, cache restore). That breakdown usually decides the question, because a pool that takes four minutes to produce a usable runner needs warm capacity to hide it, while one that takes forty seconds may not. So attack provisioning first: pre-bake the toolchain into the runner image, mirror base images locally, keep the image small. Then size a minimum warm count against your typical arrival burst rather than your peak, cap maximum concurrency to bound spend, and remember that each additional pool — per trust level, per hardware class — carries its own minimum, which is where fleets quietly get expensive.
go deeper
Know that runners in a pool are created and destroyed on demand, and that a job waiting for one to appear is queue time — separate from how long the build itself takes.
Explain what a cold start actually consists of — instance boot, image pull, agent registration, checkout, cache restore — and which of those a pre-baked runner image removes.
Diagnose a slow fleet with the right telemetry: queue wait percentiles, utilisation and arrival shape, then decide whether the fix is faster provisioning, a warm floor, or a concurrency cap that is currently being hit.
Own the economics and the governance: convert wait time into money to justify the floor, set maximum spend, and hold the line on how many separate pools exist, since each one carries its own idle cost and maintenance.
## Two costs, one currency An autoscaling runner pool sits between two costs that are usually measured in different units, which is why the argument goes in circles. **Idle capacity** is visible: instances running with no job, on an invoice, attributable to your team. **Queue time** is invisible: engineers waiting for a runner to pick up their job. It is larger than it looks, because it is paid on every run of every pipeline by every engineer, and because it changes behaviour — when feedback is slow, people batch more changes into a branch, review later, and context-switch, which costs more than the waiting itself. The first move as an owner is to put both in the same currency. Multiply median queue wait by runs per day by a loaded engineering rate and compare it with the monthly cost of the warm instances that would remove it. The number is usually decisive in one direction or the other, and it converts a taste argument into an arithmetic one. ## Measure the right things Most fleets are reported on with the wrong metric — average total job duration — which mixes waiting with working and hides the problem. - **Queue wait**, from job enqueued to first step executing, at p50 and p95. The tail is what people remember and what they complain about. - **Provisioning latency**, decomposed: instance boot, runner image pull, agent registration and job assignment, source checkout, cache restore. This is the cold-start cost of being single-use, and its decomposition tells you what to fix. - **Utilisation**, as busy runner-minutes over paid runner-minutes. Low utilisation with low queue wait means you are over-provisioned; high utilisation with rising queue wait means you are at the cliff. - **Arrival shape**, not just volume: jobs per minute over a day and a week. Almost every fleet is idle overnight and spiky at mid-morning, after lunch, and around merge activity, and a monorepo that fans out fifty jobs per pull request arrives as a step function rather than a stream. - **Rejected or timed-out queue entries**, which are the failure mode of a maximum-concurrency cap and get misreported as flaky CI. ## Fix provisioning before buying warmth The most common mistake is buying warm capacity to hide a slow cold start. If a fresh runner takes four minutes to become useful, a warm pool is a permanent tax paid to conceal a fixable defect. Attack the parts first: - **Pre-bake the runner image** with the toolchain, language runtimes and common base images already present, so no job spends its first two minutes installing what every job needs. - **Keep that image small enough to pull quickly**, and pull it from a registry close to the instances or from a local mirror. - **Make cache restore fast**, with a cache service or shared remote cache in the same region as the pool. - **Cut registration latency** so an instance that has booted is assigned work immediately rather than on the next poll interval. Once a fresh runner is useful in tens of seconds rather than minutes, scale-to-zero becomes viable for most of the day and the warm pool only has to cover bursts. ## Then size the warm floor Size the minimum against the *typical* burst, not the peak — the peak is what the maximum and the scale-out rate are for. A workable shape is: a small always-on floor during working hours that absorbs the ordinary arrival burst; scale to zero outside them; a scale-out step large enough to answer a monorepo fan-out in one move rather than ten; and a scale-in cooldown long enough that a lull between two pull requests does not destroy capacity you are about to need. Overshoot deliberately on scale-out and be lazy on scale-in — the asymmetry is cheap and matches how demand actually arrives. Cap maximum concurrency so a runaway pipeline or a retry storm cannot spend without limit, and make sure hitting that cap is visible as queueing rather than as failed jobs. ## The hidden multiplier: how many pools Every separate pool has its own floor, its own image to maintain, and its own idle waste. Pools multiply for good reasons — trust separation between internal and outside-contributor jobs, hardware classes such as GPU or a second CPU architecture, licensed operating systems, region or data-residency constraints — and for bad ones, mostly one team wanting guaranteed capacity. This is where fleet cost grows quietly, and it is a governance decision rather than a technical one: keep the number of pools small and justified, and give teams fairness within a shared pool rather than a private pool of their own. Finally, use cheaper capacity where the workload tolerates it. Interruptible or spot instances suit single-use runners well when jobs are short and retry-safe, and badly when a two-hour job loses its work at ninety minutes — so the policy should follow job duration, not a blanket preference.
- Why measure queue wait separately from job duration?Because they have different owners and different fixes. Queue wait is a capacity and provisioning problem for the fleet owner; job duration is a build problem for the team that owns the pipeline. Reporting a single blended number lets a capacity shortage hide behind a slow test suite, and vice versa, so neither gets fixed.
- When are interruptible or spot instances a bad fit for CI runners?When jobs are long and not cheaply retried. A single-use runner that is reclaimed ninety minutes into a two-hour job throws away the work and the queue position. Match the policy to job duration: interruptible capacity for short, retry-safe jobs, on-demand for long or release-critical ones, and make the retry automatic so a reclaim is invisible.
- A team asks for its own dedicated runner pool to guarantee capacity. How do you respond?Ask what problem they are actually solving. If it is queue time, fairness within the shared pool — concurrency limits per pipeline, priority for release jobs — usually fixes it without a second floor to pay for and patch. Grant a separate pool for reasons a shared one cannot satisfy: different trust level, different hardware, licensing, or a residency constraint.
saying these in an interview costs you the question
- Just autoscale, capacity planning is unnecessary
- Average job duration is the metric that matters
- Keep a large warm pool because cold starts are slow
- Scale in immediately whenever a runner goes idle
- One pool per team is the fair way to allocate capacity