skip to content

During a management-API outage, which of your own automated loops can shrink a healthy transcoding fleet, and how do you stop them?

level: seniorimportance: nice to knowfreq 30%

answer

  1. your own loops are the second failure
  2. deletes land, creates do not
  3. scale-in, replace, deploy, scheduled teardown
  4. drain instead of terminate
  5. guard destruction on create health

basics

~10 s

Scale-in on a falling backlog, terminate-and-replace supervision, a deployment already in flight, and scheduled teardown can each destroy capacity you cannot rebuild. Suspend scale-in, switch replacement to drain-not-terminate, and pause deployments until creates succeed.

solid answer

~40 s

The second failure in a control-plane incident is usually your own automation. Destructive calls typically keep landing while creates fail, so every loop that removes a worker still works perfectly while the loop that adds one does not. Four loops do this: **scale-in**, when upstream is also degraded and the backlog briefly falls; **terminate-and-replace** supervision, which kills a worker that failed a check and then cannot create the replacement; a **rolling deployment already in flight**, which retires each batch before creating it; and any **scheduled teardown** that runs on a clock regardless of conditions. The freeze is simple: suspend scale-in, switch unhealthy-worker handling to drain and leave running, pause deployments, and disable scheduled teardown. Better still, build the condition in - have each loop probe whether creates are succeeding before it destroys anything.

code

pseudocode · 17 lines
pseudocode
every interval:
    createHealth = recentCreateOutcomes()   # healthy | degraded

    for each worker in fleet:
        if worker.consecutiveFailedChecks >= threshold:
            if createHealth is degraded:
                stopSendingWork(worker)         # drain, keep instance alive
                record("replacement deferred", worker)
            else:
                stopSendingWork(worker)
                terminate(worker)
                create(replacementFor(worker))

    if createHealth is degraded:
        suppress(scaleIn)                       # a shrunk fleet cannot grow back
    else:
        allow(scaleIn)

go deeper

for a junior

Recall that removing a worker and adding one are separate platform operations, and that the removal can succeed at a moment when the addition cannot.

for a middle

Explain the asymmetry between destructive and constructive management calls, and name which routine loops contain a destructive step.

for a senior

Show the operational move: freeze scale-in, convert replacement to draining, pause deployments, and justify each with what happens to the fleet if you do not.

for a principal

Make it structural rather than procedural - require that any loop with a destructive step carry a condition on create health, so no incident depends on somebody remembering a runbook.

## Your automation is the second failure A degraded control plane freezes your fleet. What actually shrinks it is almost always something you wrote, because of one asymmetry: **destructive and idempotent management calls tend to keep landing while creates fail or hang.** Terminating an instance is a small, local, cheap operation; creating one has to place it, reserve capacity, allocate addresses and publish it consistently. The half of the platform that survives is exactly the half your loops use to destroy things. So the fleet does not merely stop growing. It ratchets downward, one plausible automated decision at a time, and every step is unrecoverable until the incident ends. ## The four loops that shrink you | Loop | What it does normally | What it does during a control-plane outage | |---|---|---| | Scale-in on backlog or utilisation | Removes workers when demand falls | Removes workers because upstream is also degraded and the backlog dipped; they never come back | | Terminate-and-replace supervision | Kills a failing worker and creates a fresh one | Kills successfully, creates nothing, and repeats on the next worker | | Rolling deployment in flight | Retires a batch, then creates its replacement | Strands itself with the retire done and the create hanging | | Scheduled teardown or nightly shrink | Releases capacity on a clock | Fires on the clock regardless, into a window where nothing can be rebuilt | The second row is the cruellest, because it is triggered by a health signal that may itself be a symptom of the same incident. A worker that cannot reach a management endpoint at boot, or whose dependency is slow, fails its check; the loop terminates it; the replacement never arrives; the remaining workers take more load, get slower, and fail their checks too. That feedback loop can empty a fleet that was serving perfectly well. ## The freeze list 1. **Suspend scale-in** - the whole of it, not just its rate. A fleet that cannot grow must not be allowed to shrink. 2. **Switch replacement to drain-not-terminate.** Stop routing work to a worker that looks unhealthy and leave the instance running; if the check was a symptom of the incident, you have kept the capacity, and if it was real, you have lost nothing you could have replaced. 3. **Pause any deployment in flight**, and do not start one. Note which batches are already retired so you can finish deliberately afterwards. 4. **Disable scheduled teardown** for the duration, including anything that runs on a calendar rather than on a condition. 5. **Say so explicitly in the incident channel**, because these loops are usually owned by different people than the ones watching the incident. ## Build the condition in, so the freeze is not manual A freeze you have to remember, find and apply by hand, in four places, while an incident is running, is a runbook step that gets missed. The durable version is a guard inside the loops themselves: before any loop destroys capacity, it asks whether capacity can currently be created, and if the answer is no it degrades to the non-destructive branch. That probe can be as simple as the loop's own recent record of create outcomes - it does not need a special signal from the platform, and it should not depend on a status page, which usually lags the incident it describes. ## What this is not This is not an argument against scaling, replacement or rolling deployments; all three are correct in normal operation, and a fleet without them is a worse fleet. It is an argument that every loop with a destructive step needs a condition under which it stops taking that step - and that the condition is `can I currently create a replacement?`, not `is this worker unhealthy?`. It is also the practical face of **static stability**. A statically stable system keeps doing the last thing it was told while the control plane is unavailable. A fleet whose automation is free to shrink it is the opposite: it actively converts a management-API incident into a capacity incident, all by itself.

  • Why can an unhealthy-worker signal be misleading during this kind of incident?
    Because the check may be failing for the same reason the platform is degraded - a boot-time configuration read that times out, or a dependency that has slowed. Terminating on that signal destroys capacity in response to a symptom, and the replacement that would have justified the termination cannot be created.
  • How should a deployment that is already half-done be handled?
    Pause it where it stands and record which batches are already retired. Do not roll forward, because that requires creates, and do not roll back, because a rollback is also a deployment. Finish it deliberately once create calls are landing again, with the fleet size verified first.

saying these in an interview costs you the question

  • Assumes a terminated worker is automatically replaced
  • Leaves scale-in enabled because demand looks lower
  • Treats a failed health check as proof the worker is useless
  • Rolls back a stranded deployment mid-incident
  • Relies on a manual runbook step nobody remembers to run