skip to content

What measurement does replica autoscaling watch, and what does it compare that measurement against to pick a copy count?

level: juniorimportance: must knowfreq 66%

answer

  1. a number, a target, a count
  2. measure, compare, multiply, clamp
  3. ratio of observed to target
  4. signal is normalised per copy
  5. floor and ceiling bound the result

basics

~20 s

Replica autoscaling watches a load signal - processor utilisation, queue backlog, requests in flight - and compares it against a configured target, adding copies while the signal sits above that target and removing them while it sits below.

solid answer

~40 s

Replica autoscaling is a control loop over one number: the desired copy count. It needs three things - a **load signal**, a **target value** for that signal, and a **floor and ceiling** on the count. On every evaluation it reads the signal, divides observed by target, and scales the current count by that ratio: roughly `desired = ceil(current x observed / target)`, clamped into the allowed range. The signal is normally expressed *per copy* (average utilisation per copy, queued items per copy) so the ratio means something arithmetically. The target is not a limit the signal may never cross - it is the set point the loop steers toward, and the gap between observed and target is exactly the error the loop is correcting.

code

pseudocode · 12 lines
pseudocode
every evaluationInterval:
    observed = average(signalPerCopy, over = measurementWindow)
    ratio    = observed / target
    desired  = ceil(currentCount * ratio)
    desired  = clamp(desired, floor, ceiling)

    if desired != currentCount:
        setDesiredCount(desired)

# currentCount = 4, observed = 80, target = 50
#   ratio   = 1.6
#   desired = ceil(6.4) = 7

go deeper

for a junior

Recall the three ingredients: a measured load signal, a target for it, and a minimum and maximum copy count. More copies while the signal is above target, fewer while it is below.

for a middle

Explain the proportional recompute - current count times observed over target, rounded up and clamped - and why the signal is denominated per copy so that adding copies actually moves it.

for a senior

Show what you check when a scaling policy misbehaves: whether the signal falls when a copy is added, whether observation and target share units, and whether the floor still holds your redundancy at idle.

for a principal

Frame the target as the estate's steady-state utilisation policy: it sets the cost floor, the burst absorption every service gets for free, and how much reaction time the platform has to buy elsewhere.

## The loop and the one number it owns Replica autoscaling sits **above** the control loop that keeps a declared number of copies running. That lower loop reads a desired count and works to make reality match it. The autoscaling loop's whole job is to decide what that desired count should be right now, from a measurement. It does not choose where copies land, does not judge whether a copy is healthy, and does not control how a new version replaces an old one - it moves one number and lets the rest of the platform react. Four pieces make the loop: 1. **A load signal** - a number that stands in for how hard the workload is being pushed. Processor utilisation per copy, queued items per copy, in-flight requests per copy, or an application-reported number such as sessions held. 2. **A target** - the value of that signal you want to hold. Choosing the target is choosing your steady-state utilisation. 3. **A recompute rule** - the arithmetic that turns the gap between observed and target into a new count. 4. **A floor and a ceiling** - the smallest and largest count the loop may ask for. The usual recompute rule is proportional: `desired = ceil(current x observed / target)`. With four copies running, an observed average of 80 and a target of 50, the ratio is 1.6, `4 x 1.6 = 6.4`, and the loop asks for **7**. If the signal then settles near the target, the ratio approaches 1 and the count stops moving. That self-correcting shape is why the signal has to be a *rate or level*, not a total that only grows. ## Why the signal is normalised per copy The ratio arithmetic only works if the signal responds to the count. Two properties make a signal usable: - it must be **proportional to unserved demand** - it rises when the workload is falling behind; - it must be **reduced by adding a copy** - otherwise the loop has no feedback and will either sit still or run to the ceiling. | Signal | What it measures | Does adding a copy move it? | |---|---|---| | Average processor utilisation per copy | How busy each copy's share of the machine is | Yes, for work that is actually processor-bound | | Queued items per copy | Unserved work divided across the fleet | Yes - the same backlog spread over more copies | | In-flight requests per copy | Concurrency each copy is carrying | Yes, when requests are routed evenly | | Total queue length (not per copy) | Unserved work, fleet-wide | Eventually, but the ratio arithmetic is wrong against a per-copy target | | An upstream dependency's error rate | A neighbour's health | No - adding copies cannot lower it | A fleet-wide total is not unusable; it just has to be paired with a target expressed the same way, or divided by the current count before the comparison. Mixing a fleet-wide observation with a per-copy target is one of the quiet ways a scaling policy ends up pinned at the ceiling. ## What the floor and the ceiling are actually for They are not safety decoration. The **floor** is the count the loop may never go below, and it is where your availability lives when the signal reads zero - a floor of one means a single copy is one restart away from no service at all. The **ceiling** is a blast-radius cap: if the signal misbehaves, or something that is not load pushes it up, the ceiling is what stops the loop from asking for a thousand copies and spending accordingly. Platforms differ in what they default these to and whether a ceiling is mandatory, so read yours rather than assuming. ## What this loop is not - **It is not a forecast.** The loop acts *after* the signal has already moved. A spike is visible only once it is happening, which is why reaction time matters so much on this material. - **It is not capacity.** Raising the desired count is a request. Machines to run those copies on come from somewhere else, and the count can rise while nothing new actually starts. - **It is not a health check.** A failing copy is replaced by different machinery; the scaling loop only sees the aggregate signal, and a broken copy that reports a low signal can even talk the loop into removing capacity. - **The target is not a hard ceiling on the signal.** Observed sitting above target is the normal running state, not a fault - it is the error term the loop exists to close.

  • Why is the load signal usually expressed per copy rather than as a fleet-wide total?
    Because the recompute rule multiplies the current count by `observed / target`, and that only means something when both sides are denominated the same way. A per-copy signal also falls as copies are added, which is what closes the loop. A fleet-wide total compared against a per-copy target produces a ratio with no useful units and usually pins the count at its ceiling.
  • What breaks if the target is set very close to the most one copy can sustain?
    Steady state moves to the edge. Every copy runs near its limit, so any burst arrives with no absorption at all and turns straight into latency while the loop reacts. It also makes the loop twitchier: near saturation the signal is non-linear, so a small change in load produces a large change in the measurement and a large change in the requested count.
  • Can the same workload be scaled on two signals at once?
    Yes, and the usual composition is to compute a recommendation per signal and take the **highest** of them. That is deliberately asymmetric: any one signal being over target is enough to justify more copies, while shrinking requires every signal to agree. It also means one noisy or badly chosen signal can hold the count up on its own.

A thermostat, not a timer. It does not know the weather forecast; it reads the room, compares it to the set point, and keeps nudging until the gap closes.

saying these in an interview costs you the question

  • Thinks the loop predicts the spike instead of reacting after it shows up
  • Treats the target as a hard limit the signal must never cross
  • Sets a ceiling but leaves the floor at one copy, losing any redundancy at idle
  • Assumes any measured number works as a scaling signal
  • Believes raising the desired count creates the machines to run it on
  • Thinks adding a copy halves every signal, including ones no copy influences