skip to content

Updates & Availability

How the running set of copies changes shape: replaced batch by batch, rolled back, grown and shrunk with load, or thinned for maintenance. Interviewers ask because that is when a service goes down.

on this pageshow

questions

24

What measurement does replica autoscaling watch, and what does it compare that measurement against to pick a copy count?

level: juniorimportance: must knowfreq 66%

answer

  1. a number, a target, a count
  2. measure, compare, multiply, clamp
  3. ratio of observed to target
  4. signal is normalised per copy
  5. floor and ceiling bound the result

basics

~20 s

Replica autoscaling watches a load signal - processor utilisation, queue backlog, requests in flight - and compares it against a configured target, adding copies while the signal sits above that target and removing them while it sits below.

solid answer

~40 s

Replica autoscaling is a control loop over one number: the desired copy count. It needs three things - a **load signal**, a **target value** for that signal, and a **floor and ceiling** on the count. On every evaluation it reads the signal, divides observed by target, and scales the current count by that ratio: roughly `desired = ceil(current x observed / target)`, clamped into the allowed range. The signal is normally expressed *per copy* (average utilisation per copy, queued items per copy) so the ratio means something arithmetically. The target is not a limit the signal may never cross - it is the set point the loop steers toward, and the gap between observed and target is exactly the error the loop is correcting.

code

pseudocode · 12 lines
pseudocode
every evaluationInterval:
    observed = average(signalPerCopy, over = measurementWindow)
    ratio    = observed / target
    desired  = ceil(currentCount * ratio)
    desired  = clamp(desired, floor, ceiling)

    if desired != currentCount:
        setDesiredCount(desired)

# currentCount = 4, observed = 80, target = 50
#   ratio   = 1.6
#   desired = ceil(6.4) = 7

go deeper

for a junior

Recall the three ingredients: a measured load signal, a target for it, and a minimum and maximum copy count. More copies while the signal is above target, fewer while it is below.

for a middle

Explain the proportional recompute - current count times observed over target, rounded up and clamped - and why the signal is denominated per copy so that adding copies actually moves it.

for a senior

Show what you check when a scaling policy misbehaves: whether the signal falls when a copy is added, whether observation and target share units, and whether the floor still holds your redundancy at idle.

for a principal

Frame the target as the estate's steady-state utilisation policy: it sets the cost floor, the burst absorption every service gets for free, and how much reaction time the platform has to buy elsewhere.

## The loop and the one number it owns Replica autoscaling sits **above** the control loop that keeps a declared number of copies running. That lower loop reads a desired count and works to make reality match it. The autoscaling loop's whole job is to decide what that desired count should be right now, from a measurement. It does not choose where copies land, does not judge whether a copy is healthy, and does not control how a new version replaces an old one - it moves one number and lets the rest of the platform react. Four pieces make the loop: 1. **A load signal** - a number that stands in for how hard the workload is being pushed. Processor utilisation per copy, queued items per copy, in-flight requests per copy, or an application-reported number such as sessions held. 2. **A target** - the value of that signal you want to hold. Choosing the target is choosing your steady-state utilisation. 3. **A recompute rule** - the arithmetic that turns the gap between observed and target into a new count. 4. **A floor and a ceiling** - the smallest and largest count the loop may ask for. The usual recompute rule is proportional: `desired = ceil(current x observed / target)`. With four copies running, an observed average of 80 and a target of 50, the ratio is 1.6, `4 x 1.6 = 6.4`, and the loop asks for **7**. If the signal then settles near the target, the ratio approaches 1 and the count stops moving. That self-correcting shape is why the signal has to be a *rate or level*, not a total that only grows. ## Why the signal is normalised per copy The ratio arithmetic only works if the signal responds to the count. Two properties make a signal usable: - it must be **proportional to unserved demand** - it rises when the workload is falling behind; - it must be **reduced by adding a copy** - otherwise the loop has no feedback and will either sit still or run to the ceiling. | Signal | What it measures | Does adding a copy move it? | |---|---|---| | Average processor utilisation per copy | How busy each copy's share of the machine is | Yes, for work that is actually processor-bound | | Queued items per copy | Unserved work divided across the fleet | Yes - the same backlog spread over more copies | | In-flight requests per copy | Concurrency each copy is carrying | Yes, when requests are routed evenly | | Total queue length (not per copy) | Unserved work, fleet-wide | Eventually, but the ratio arithmetic is wrong against a per-copy target | | An upstream dependency's error rate | A neighbour's health | No - adding copies cannot lower it | A fleet-wide total is not unusable; it just has to be paired with a target expressed the same way, or divided by the current count before the comparison. Mixing a fleet-wide observation with a per-copy target is one of the quiet ways a scaling policy ends up pinned at the ceiling. ## What the floor and the ceiling are actually for They are not safety decoration. The **floor** is the count the loop may never go below, and it is where your availability lives when the signal reads zero - a floor of one means a single copy is one restart away from no service at all. The **ceiling** is a blast-radius cap: if the signal misbehaves, or something that is not load pushes it up, the ceiling is what stops the loop from asking for a thousand copies and spending accordingly. Platforms differ in what they default these to and whether a ceiling is mandatory, so read yours rather than assuming. ## What this loop is not - **It is not a forecast.** The loop acts *after* the signal has already moved. A spike is visible only once it is happening, which is why reaction time matters so much on this material. - **It is not capacity.** Raising the desired count is a request. Machines to run those copies on come from somewhere else, and the count can rise while nothing new actually starts. - **It is not a health check.** A failing copy is replaced by different machinery; the scaling loop only sees the aggregate signal, and a broken copy that reports a low signal can even talk the loop into removing capacity. - **The target is not a hard ceiling on the signal.** Observed sitting above target is the normal running state, not a fault - it is the error term the loop exists to close.

  • Why is the load signal usually expressed per copy rather than as a fleet-wide total?
    Because the recompute rule multiplies the current count by `observed / target`, and that only means something when both sides are denominated the same way. A per-copy signal also falls as copies are added, which is what closes the loop. A fleet-wide total compared against a per-copy target produces a ratio with no useful units and usually pins the count at its ceiling.
  • What breaks if the target is set very close to the most one copy can sustain?
    Steady state moves to the edge. Every copy runs near its limit, so any burst arrives with no absorption at all and turns straight into latency while the loop reacts. It also makes the loop twitchier: near saturation the signal is non-linear, so a small change in load produces a large change in the measurement and a large change in the requested count.
  • Can the same workload be scaled on two signals at once?
    Yes, and the usual composition is to compute a recommendation per signal and take the **highest** of them. That is deliberately asymmetric: any one signal being over target is enough to justify more copies, while shrinking requires every signal to agree. It also means one noisy or badly chosen signal can hold the count up on its own.

A thermostat, not a timer. It does not know the weather forecast; it reads the room, compares it to the set point, and keeps nudging until the gap closes.

saying these in an interview costs you the question

  • Thinks the loop predicts the spike instead of reacting after it shows up
  • Treats the target as a hard limit the signal must never cross
  • Sets a ceiling but leaves the floor at one copy, losing any redundancy at idle
  • Assumes any measured number works as a scaling signal
  • Believes raising the desired count creates the machines to run it on
  • Thinks adding a copy halves every signal, including ones no copy influences
open as a page

Your order API runs six copies behind one shared name; why does replacing all six at once drop requests, while replacing them a batch at a time does not?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Stopping all six copies leaves the shared name with no copy able to serve until a replacement finishes starting, so every request in that window fails. Replacing a batch at a time keeps the rest serving throughout.

open as a page

Why does a platform replace the copies of a data-owning, individually-named workload one at a time in a fixed order?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because those copies are not interchangeable. Each owns a slice of durable data and a name the replacement must reclaim, and the group needs a majority alive, so only one may be missing at a time.

open as a page

A queue-fed rendering worker autoscales on average processor use per copy, yet its backlog keeps growing - why?

level: middleimportance: must knowfreq 56%

basics

~20 s

The signal does not track the pain. A rendering worker that spends most of its time waiting on downloads and uploads shows modest processor use even while items pile up, so the measurement never crosses its target and the copy count never rises.

open as a page

A platform keeps every applied rollout as a numbered revision - what does one revision contain, and what happens when you reapply an earlier one?

level: middleimportance: must knowfreq 62%

basics

~20 s

A revision is a snapshot of the entire declared spec at one apply - image reference, copy count, reservations, ceilings and settings together. Reapplying an earlier revision runs a normal forward rollout whose target happens to be that older spec.

open as a page

Reapplying the previous revision put the old ledger writer back in two minutes, but its bad rows are still there - why?

level: middleimportance: must knowfreq 70%

basics

~10 s

A revision records the desired spec, so reapplying one changes which version runs; it says nothing about rows already committed to a store outside the workload. The spec reverts, the written data does not.

open as a page

In a rolling replacement, why is each step gated on the new copy reporting itself ready rather than on its process having started?

level: middleimportance: must knowfreq 63%

basics

~20 s

A started process is not yet a serving one, and the step is trading away a copy that can serve. Gating on the readiness answer keeps the exchange honest; gating on process start removes serving capacity in exchange for capacity that does not exist yet.

open as a page

A host must be drained tonight for a firmware upgrade - what does that drain actually do, and how does it differ from the host failing?

level: middleimportance: must knowfreq 62%

basics

~20 s

A drain is a disruption the platform is told about first: it stops placing new work on the host, asks the copies already there to stop one at a time, and waits for replacements. A host failure announces nothing and the copies are simply gone.

open as a page

A sharded index's copies are being replaced one by one; what must hold for old and new members to keep replicating to each other?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Both directions must work at once: the new member reads what the old ones wrote, and the old ones tolerate what the new one writes. They are each other's replication source for the whole rollout.

open as a page

A host drain has sat for an hour with three copies left and a cap on how many may be voluntarily down - what is happening?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The drain is blocked, not broken: every removal is a request checked against the cap, and while the workload is already at the floor of available copies, granting one more would breach it, so the request is refused and quietly retried.

open as a page

What must an ordered stateful rollout wait for before it touches the next copy, and why is a started process not enough?

level: middleimportance: should knowfreq 46%

basics

~20 s

It must wait until the replaced copy has rejoined its group and caught up. A started process answers its port long before its store is replayed, so advancing on that gate stacks up members that are behind.

open as a page

A workload's copy count swings between four and twelve every few minutes - what causes that, and what damps it?

level: middleimportance: should knowfreq 46%

basics

~20 s

The loop is reacting to its own actions. Adding copies drives the per-copy signal below target, the loop removes them, the signal jumps back above target, and it adds again. A settling window, a tolerance band around the target, and slower shrinking than growth damp the oscillation.

open as a page

When a workload is allowed to scale to zero copies, what does the first request after an idle period wait for?

level: middleimportance: should knowfreq 42%

basics

~20 s

It waits for the whole cold path: something in the request path must notice it, raise the count from zero, and hold the request while a host is chosen, the image is pulled if absent, the process starts and the readiness check passes. That is the cold start.

open as a page

A workload runs one copy and its cap allows none down - why does every drain of its host block, and what resolves it?

level: middleimportance: should knowfreq 44%

basics

~20 s

The arithmetic can never be satisfied: taking the only copy down leaves zero available, below any floor of one, so the removal is refused and retried forever. Resolving it means more copies, a looser cap, or a deliberate outage.

open as a page

A two-minute traffic spike is over before any new copy serves a request - where did the time go?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Reactive scaling has a chain of delays: an averaged measurement window, the evaluation interval, placement and image pull, process start-up, then the readiness check and warm-up before traffic is routed. Two to five minutes is normal, so a two-minute spike finishes first.

open as a page

Reapplying the last known-good revision fixed the bad release but also undid a memory ceiling an operator had raised by hand - why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A revision is the whole spec applied atomically, so reapplying an older one restores every field it holds, including settings that changed after it. Rollback granularity is the apply, never the single field you regret.

open as a page

Your rolling replacement of an order API completed within seconds and the service went dark, because every new copy reported ready the moment its process started — what happened?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Every step's gate opened immediately, so old copies were removed at full speed and the whole set ended up on copies that could not yet serve. The batch structure survived; the property it protects did not, because the ready answer was true too early.

open as a page

Your order API's rolling replacement started two new copies twenty minutes ago, has removed no old copy since, and both specs are still present — why?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The new copies are not reporting ready, so the guard that allows an old copy to be removed keeps evaluating false and the step never advances. The rollout is stuck on purpose, holding the old copies in service rather than trading them for copies that cannot serve.

open as a page

Each copy of a five-member data group needs forty minutes to catch up after replacement — how would you judge whether that rollout is acceptable?

level: principalimportance: should knowfreq 34%

basics

~20 s

Judge it on three things: whether forty minutes is inherent or a defect, whether redundancy still covers an unplanned loss while one member is away, and whether a rollout that long can finish inside any window you can commit to.

open as a page

How would you set disruption-cap policy across a shared cluster where maintenance windows keep expiring with drains still blocked?

level: principalimportance: should knowfreq 38%

basics

~20 s

Make a cap a contract the workload owner has to fund: reject caps whose floor equals the declared copy count, prefer caps expressed against that count, default the workloads that declare none, and alert the owner - not the operator - when a drain blocks.

open as a page

After going back to revision 4 of five kept revisions, what does the platform's revision list look like?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

Revision history is append-only: going back adds a new highest entry holding the older spec, and retention drops the oldest. Nothing is popped or deleted, and the bad entry stays readable until it ages out.

open as a page

An ordered stateful rollout is held after replacing only the highest-numbered copies — what state is the group left in?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A cleanly split group: every copy above the hold position runs the new version with its own store intact, every copy at or below it is untouched on the old one. Nothing is half-replaced, and nothing is reverted.

open as a page

Raising a workload's desired copy count from twenty to sixty left only twenty-six running - what is missing, and which loop supplies it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Machines. Replica autoscaling changes a number, not capacity: each copy needs a host with enough unreserved room for its reservation, and the pool ran out. Growing the pool is a second, much slower loop, so host supply is the real ceiling on copy count.

open as a page

Three copies of a quorum-based store all reported available, the drain honoured its cap, and the store still lost quorum - how?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The cap counts copies that report themselves available, and a copy can report available before it has rejoined the store's own membership. Two of three were counted, only one was really voting, so removing the third left no majority.

open as a page