Your team keeps a pool of machines up for its jobs. Beyond the hourly bill, what does that pool cost, and what stays when you shrink it?
answer
- two bills, one invoice
- sized for peak, idle most of the day
- people cost is per cluster
- halving machines changes one column only
- the lever is count, not size
basics
~20 sCapacity billed while idle, plus a standing surface: version consistency across every worker, coordinated upgrades, an access and credential model, on-call for infrastructure nobody on the team wrote, and the expertise to run it. That surface is per cluster, so halving the machines does not halve it.
solid answer
~60 sTwo bills run at once and only one of them arrives as an invoice. The first is capacity: a pool that is up bills whether or not work is on it, and because it is sized for the peak it is idle most of the time — real utilisation of a pool with a daily shape is often a small fraction. The second is the surface the team inherits. Every worker must carry a compatible runtime and library set, so upgrades become a coordinated exercise rather than a deploy. Processes on many machines need credentials and an access model. Somebody must be on call for a component nobody on the team wrote. And every engineer now has to understand a second execution model to read a failure. The distinction that matters when you cut cost is that the first bill scales with machine count and the second does not: the upgrade path, the access model, the rota and the expertise cost the same at ten machines as at a hundred.
go deeper
Know that machines that are up cost money even when nothing runs on them, and that somebody has to keep them patched, credentialed and monitored.
Separate the two bills and say which parts scale with machine count. Be able to explain why a pool sized for the peak is idle most of the day and why that idleness is charged.
Quantify it: utilisation over a week, the share of wall clock that is set-up and tail, and the hours the team actually spends on upgrades, access and on-call. Bring both columns to a cost review rather than the invoice alone.
This is the call you own. Decide how many execution models the organisation supports, what a new pipeline gets by default, who may stand up another pool and what the escape hatch costs — because the standing surface multiplies per runtime, not per machine.
## Two bills, one invoice A standing pool — machines your team keeps up so that work can be handed to them at any time — produces two quite different costs, and only the first appears on a statement. The **capacity bill** is the one everybody sees: machines that are up are charged for, regardless of whether anything is running on them. Because a pool is sized for its peak, it is idle for most of the day by construction. A pool built to absorb a nightly load carries that width for twenty-four hours, and even during a run, utilisation is well below the headline: capacity sits waiting through set-up, through the serial parts of a plan, through the tail of each stage while the last pieces finish, and whenever the run is waiting on its source rather than on its processors. The meter does not care. The **standing surface** is the one that gets left out of comparisons, and it is usually the larger of the two once a team is past its first year. ## What the surface consists of - **Version consistency.** Every worker must run a compatible runtime and a compatible library set, and so must the program. A dependency change is no longer a deploy; it is a change to a fleet. - **Upgrades.** Engine and platform versions move whether or not you want them to. Upgrading means a compatibility pass over every pipeline that runs on the pool, a rehearsal, and a window. - **Access and credentials.** Processes on many machines read storage and write to destinations on someone's authority. That authority has to be granted, scoped, rotated and audited — for machines, not just for people. - **On-call.** Somebody must be reachable when the pool is unhealthy at three in the morning, for software the team did not write and cannot patch. - **Expertise.** Every engineer who reads a failure now needs a mental model of distributed execution. That is a hiring filter, an onboarding cost and a bus-factor risk, and it is paid in people rather than in machines. - **Reliability of the pool itself.** The pool is now a dependency of everything that runs on it: when it is down, nothing runs, including the small jobs that would have been fine anywhere. ## What shrinking the pool does and does not touch | cost | halves when the pool halves | stays the same | |---|---|---| | hourly capacity charge | yes | — | | total memory and processor available | yes | — | | upgrade path and compatibility testing | no | one path per cluster | | access and credential model | no | one model per cluster | | on-call rota | no | one rota per cluster | | the expertise the team must hold | no | one model to learn | The practical consequence is that cost work aimed only at the machine count reaches a floor quickly, and the floor is the per-cluster half of the list. The lever that moves the second column is not size; it is **count** — how many distinct pools, runtimes and execution models the organisation is on at once. ## Comparing honestly against one large machine The alternative that keeps this list honest is one large machine running one process. Its capacity bill can look worse per unit of work, and it is often still cheaper overall, because the second column collapses: one thing to patch, one place logs land, one failure mode, one mental model, and no possibility of half the work finishing. A comparison that prices only the hourly rates against each other is not a comparison. Put both columns in it, and state the exposure each way — a single machine has a ceiling you can hit and an outage that stops everything, and those belong in the same table. ## What varies, and what to do with it The arithmetic of the first column changes with how capacity is supplied: capacity raised for one run and torn down afterwards charges for a shorter window but still charges for its own start-up, and capacity you never see charges per statement or per byte read rather than per hour. None of those variations touch the second column much — the upgrade, access, on-call and expertise costs follow the *runtime you are on*, not the way its machines are billed. So the decision a lead actually owns is about count and defaults: how many execution models the organisation supports, what a new pipeline gets by default, who is allowed to introduce a second pool, and what the escape hatch costs when someone needs one. Get that wrong and the standing surface multiplies while every individual cost review stays focused on the one column that was never the problem.
- Why is utilisation of a standing pool usually far below what a team expects?Because it is sized for the peak and the peak is brief, and because even during a run the capacity is not all busy: it waits through set-up, through the serial parts of the plan, through the tail of each stage while the last pieces finish, and whenever the run is limited by its source rather than its processors. The meter runs through all of it.
- Which of these costs does adopting a second engine multiply?All of the second column. A second runtime means a second upgrade path, a second access model, a second set of failure modes on the same rota and a second mental model for every engineer, while the capacity bill may not move at all. That is why the number of execution models an organisation supports is a bigger lever on standing cost than the size of any one pool.
- What belongs in the comparison against one large machine besides money?The exposure on both sides. A single machine has a hard ceiling and an outage that stops everything at once; a pool removes those but adds partial failure, a fleet to keep consistent and a rota. Price the obligations alongside the rates, and state which risk the team would rather carry.
Running a vehicle fleet instead of owning one car. Selling half the vans halves the fuel and the parking, and changes nothing about needing a mechanic on call, an insurance policy, a licensing regime and somebody who knows the rules.
saying these in an interview costs you the question
- Prices a cluster as the hourly rate of its machines only
- Assumes a standing pool is busy most of the time
- Thinks halving the machines halves the total cost
- Believes a support contract removes the operational surface
- Adds a second runtime without counting the second rota
- Compares against one large machine on rate alone