skip to content

Raising a workload's desired copy count from twenty to sixty left only twenty-six running - what is missing, and which loop supplies it?

level: seniorimportance: nice to knowfreq 33%

answer

  1. a count is a request, not capacity
  2. each copy carries a reservation
  3. pending is not failed
  4. a second, slower loop buys machines
  5. requested minus running is the real alarm

basics

~20 s

Machines. Replica autoscaling changes a number, not capacity: each copy needs a host with enough unreserved room for its reservation, and the pool ran out. Growing the pool is a second, much slower loop, so host supply is the real ceiling on copy count.

solid answer

~40 s

The desired count is a **request**, and the scheduler can only satisfy it where a host has enough capacity left after the reservations already placed on it. Once the pool is full the remaining copies stay pending: requested, not running, not serving, and usually not reported as an error. A second loop grows the pool by acquiring hosts, and it is far slower than the replica loop - acquire, boot, join the cluster, pull images - minutes rather than seconds. So the two loops have to be configured against each other: a copy ceiling the pool can never satisfy is a promise the platform cannot keep, and a pool allowed to grow without bound is an unbounded bill. The symptom to recognise is silence, not failure.

go deeper

for a junior

Know that asking for more copies does not create machines. If there is no room in the pool, the extra copies wait rather than run, and nothing reports an error.

for a middle

Explain placement as capacity arithmetic - each copy's reservation subtracted from a host's remaining capacity by the scheduler - and why a copy with nowhere to go stays pending indefinitely.

for a senior

Show that you monitor requested against running, that you size the copy ceiling against real pool limits, and that you account for the machine-acquisition lag stacking on top of the scaling lag.

for a principal

Own both bounds together: the copy ceiling is an availability promise and the pool ceiling is a spend control, and a ceiling the pool can never satisfy is a commitment the platform cannot honour.

## Two loops, one number between them Replica autoscaling and host-pool growth are separate control loops that meet at a single point: the copies the first one asks for have to land on machines the second one supplies. | | Replica loop | Host-pool loop | |---|---|---| | Changes | The desired copy count | The number of machines in the pool | | Reacts to | A load signal against a target | Copies that cannot be placed | | Typical latency | Seconds to a couple of minutes | Several minutes | | Bounded by | Its configured floor and ceiling | Quota, budget and how fast machines can be acquired | The replica loop is the fast one and it is the one everybody configures. It is also the one that cannot keep its promise alone: raising the count from 20 to 60 creates 40 requests for placement, not 40 machines' worth of capacity. ## Why the extra copies simply sit there Placement is capacity arithmetic before it is anything else. Each copy declares a **reservation** - the processor and memory the scheduler must set aside for it - and the scheduler subtracts that from a candidate host's remaining capacity when it places the copy there. Note the direction: the reservation is what the *scheduler* subtracts at placement time; the *ceiling* is a separate number the runtime enforces on the running process. A copy can only be placed where the unreserved remainder is at least its reservation. When no host has that much left, the copy is not rejected and nothing fails. It stays **pending** - the platform keeps trying, because desired state is level-triggered and the request does not expire. That is why the symptom is so easy to miss: the count on the dashboard says 60, the number serving is 26, and there is no error anywhere. Anyone reading only the desired count believes the workload scaled. A second, sharper version of the same arithmetic: if a copy's reservation is a large fraction of a whole machine, the pool grows in a very coarse staircase. Each added copy may need an entire new host, so the fleet's ability to grow is quantised by machine size rather than by load. ## What the slower loop has to do Growing the pool is not one step: 1. Notice that copies are pending for want of capacity. 2. Acquire a machine - subject to quota, availability in the target failure domain, and whatever the infrastructure takes to hand one over. 3. Boot it and join it to the cluster so the control plane will place work on it. 4. Pull the images the newly placed copies need, on a host whose cache is empty. 5. Start those copies and wait for their readiness checks. Steps 2 and 4 are the expensive ones, and they stack on top of the replica loop's own lag. The practical consequence: a workload that can grow within its existing pool scales in about a minute, and the same workload growing past the pool scales in five or ten. Those are two very different products, and the difference is invisible in the scaling configuration. ## Configuring the two loops against each other - **A copy ceiling the pool cannot satisfy is a lie.** If the maximum count needs 40 hosts and quota allows 25, the extra ceiling buys nothing but pending copies at the worst possible moment. - **A pool with no bound is an unbounded bill.** The pool loop needs a maximum too, and it is the last line of defence against a runaway signal. - **Standing room is the only fast path.** Copies that fit in capacity that is already joined start in seconds. Everything past that waits for machines. How much standing room to keep is capacity planning and is owned elsewhere - what belongs here is knowing that it is the boundary between the two regimes. - **Alert on pending copies, not just on the count.** The only honest signal that scaling failed is the gap between requested and running, held for longer than a start-up normally takes. ## Platform variation Designs differ in how tightly the two loops are coupled. Some watch for unplaceable work and add machines reactively; some let you declare a standing pool and never grow it; some hand out capacity per copy so that the distinction almost disappears from the operator's view. What is constant is the underlying fact: a copy count is a request against a finite pool, and when the pool is the binding constraint, the replica loop's tuning is irrelevant.

  • Why is a pending copy not reported as an error?
    Because desired state is level-triggered and the request does not expire. The platform keeps re-evaluating placement and will start the copy the moment capacity appears, which is the behaviour you want for a transient shortage. The cost is that a permanent shortage looks exactly the same as a transient one, so the gap between requested and running has to be watched explicitly.
  • What does it do to scaling behaviour when one copy's reservation is a large fraction of a whole machine?
    It quantises growth. Each additional copy effectively needs a new host, so the pool grows in coarse steps, each with the full machine-acquisition lag, and the fleet's utilisation is poor because the leftover slice on each host is too small to hold anything. Smaller copies pack better and let the replica loop work inside existing capacity far more often.
  • Which alarm actually catches this failure?
    The gap between the requested copy count and the number running and ready, sustained for longer than a normal start takes. Alerting on the desired count tells you the loop reacted; alerting on the load signal tells you the workload is struggling. Only the difference between requested and serving tells you the platform could not honour what was asked.

saying these in an interview costs you the question

  • Assumes raising the desired count creates the capacity to run it
  • Reads the desired count as the number of copies actually serving
  • Expects an error when a copy cannot be placed for lack of capacity
  • Sets a copy ceiling far beyond what the machine pool can ever supply
  • Says the reservation is enforced by the runtime rather than subtracted at placement
  • Treats host-pool growth as being as fast as adding a copy