skip to content

Your transcoding fleet's backlog is growing but every request to add a worker times out while running workers transcode normally - what now?

level: seniorimportance: should knowfreq 50%

answer

  1. capacity is frozen, not broken
  2. the queue converts loss into delay
  3. manage demand, not supply
  4. protect the workers you still have
  5. headroom is bought before, not during

basics

~20 s

Accept a frozen fleet and manage demand instead of capacity: keep the durable queue absorbing arrivals, shed or defer the lowest-value work, freeze anything that could shrink or replace workers, and communicate delay rather than failure.

solid answer

~50 s

Scaling out is a create call, so during a degraded control plane the fleet is frozen at whatever size it had when the incident began - and retrying the launch, correctly with backoff, still does not produce a worker. That makes this a demand problem, not a capacity problem. The queue is the asset: arrivals are durable, so the failure mode is delay rather than loss, and your job is to choose whose work is delayed. Shed or defer the lowest-value classes, pause any bulk or backfill producers you control, and protect throughput by freezing everything that would cost you a worker - scale-in, instance replacement, and any deployment in flight. Then report a growing completion time with an estimate, because a backlog measured in hours is a very different customer conversation from a failure.

go deeper

for a junior

Recall that adding a worker is a request to the platform's management API, so a fleet cannot grow while that API is failing even though the existing workers are fine.

for a middle

Explain which of the fleet's routine operations are management calls and why a durable queue turns this from lost work into delayed work.

for a senior

Demonstrate the live decisions: shed by value, throttle producers you own, freeze every loop that could cost a worker, and publish a completion-time forecast rather than a status.

for a principal

Argue for the standing posture that makes the incident boring - headroom sized against a create-nothing window, conservative scale-in, and a queue retention chosen deliberately rather than inherited.

## What the scenario actually is A transcoding fleet takes jobs from a durable queue and scales on backlog depth. During a provider incident the workers that are already running keep transcoding at full rate - they hold their own configuration, they read and write their own storage, they pull from the queue - while every request to add a worker errors or hangs. Nothing about the workload is broken. What is broken is your ability to change the size of anything. That single fact reframes the incident. You do not have a throughput bug to fix; you have a fixed amount of throughput and more demand than it can absorb, for an unknown window. ## Which of your operations are secretly management calls | Operation | Is it control-plane work? | Consequence now | |---|---|---| | Worker pulls and transcodes a job | No | Keeps running at full rate | | Worker writes output to storage | No | Unaffected | | Add a worker to the fleet | Yes | Fails or hangs | | Replace a worker that failed its check | Yes, for the create half | Destroys capacity you cannot rebuild | | Roll out a new worker build | Yes | Can strand itself half-deployed | | Register a new worker behind the entry point | Yes | New capacity would be invisible anyway | | A booting worker reading its stored boot configuration | Usually yes | A worker can launch and still be useless | The last two rows matter because they defeat the hopeful plan of `just get one instance up`. Even if a create eventually lands, the instance may not be able to fetch what it needs to join the fleet. ## What you can still do 1. **Let the queue do its job.** A durable queue converts an arrival you cannot serve now into work you will serve later. That is the difference between a delay and a loss, and it is the reason this shape of incident is survivable at all for a batch fleet. 2. **Shed or defer by value.** Stop consuming low-value classes - re-encodes, speculative pre-generation, backfills, internal test jobs - so the fixed throughput goes to the work that matters. If your jobs carry a priority, this is a consumer-side change that creates nothing. 3. **Throttle the producers you control.** Pausing a bulk import you own is faster and safer than any platform action. 4. **Freeze every loop that can cost you a worker.** Suspend scale-in, switch unhealthy-worker handling from terminate-and-replace to drain-and-leave-running, and pause any deployment in flight. Capacity you still have is the only capacity you will get. 5. **Re-forecast and communicate.** Compute completion time from current drain rate against current arrival rate and publish it. A backlog measured in hours with a number attached is a manageable conversation; silence is not. ## What not to do - **Do not treat retries as a plan.** Backing off politely is correct behaviour and it is not a mitigation: the call is not being rejected because you asked impolitely, and no amount of retrying creates capacity. - **Do not delete anything hoping the platform will rebuild it.** Destructive calls frequently keep succeeding while creates do not. Deleting a stuck worker converts a degraded fleet into a smaller one, permanently for the duration. - **Do not start a rolling deployment** to `fix it with a bigger worker`. Every batch begins by retiring capacity. - **Do not describe this to stakeholders as an outage of your service** if the data path is healthy. It is a capacity freeze with a growing completion time, and the wrong word triggers the wrong escalation. ## The posture that would have made this boring The answer to this incident is decided long before it, and the name for it is **static stability**: the system's correct behaviour must not require creating anything. - Run with **pre-provisioned headroom** sized to the demand you must absorb during a window in which you can create nothing, rather than to average utilisation. - Keep **scale-in conservative** so an idle hour does not leave you at the floor when the incident starts. - Make sure a **booting worker depends on as little management state as possible**, so capacity you do have is usable. - Make the **queue durable and deep**, with an explicit retention long enough to cover a plausible incident, and know what happens to work older than that. A batch fleet with a durable queue is the easy case, because the backlog is a buffer. A synchronous, user-facing service has no such buffer, which is why it is the one that needs standing headroom rather than the transcoders.

  • Why is deleting a stuck worker during this incident a bad move?
    Because the destructive half of the platform frequently keeps working while the constructive half does not. The delete lands, the replacement create does not, and a fleet that was merely frozen is now permanently smaller for the rest of the incident.
  • Would the same reasoning hold for a synchronous, user-facing service instead of a batch fleet?
    The mechanism is identical but the cushion is gone. A batch fleet has a durable queue that turns excess demand into delay; a request-serving tier has only the latency budget and then errors. That is exactly why standing headroom and a warm standby are justified there and often not for the transcoders.
  • How do you decide when to declare the backlog an incident for customers?
    From the completion-time forecast, not from the platform's state. Compute drain rate against arrival rate, project when the oldest queued item will finish, and compare that with what customers were promised. If the projection breaches the promise, communicate with the number attached.

saying these in an interview costs you the question

  • Thinks retrying harder will eventually launch the instance
  • Deletes stuck workers expecting the platform to rebuild them
  • Starts a rolling deployment to raise throughput mid-incident
  • Calls it a service outage while the data path is healthy
  • Assumes the scaling policy will recover the fleet on its own