skip to content

A queue-fed rendering worker autoscales on average processor use per copy, yet its backlog keeps growing - why?

level: middleimportance: must knowfreq 56%

answer

  1. you scale what you measure
  2. wait-dominated work hides from processor use
  3. unserved work, divided per copy
  4. does adding a copy move it?
  5. a saturating signal stops driving

basics

~20 s

The signal does not track the pain. A rendering worker that spends most of its time waiting on downloads and uploads shows modest processor use even while items pile up, so the measurement never crosses its target and the copy count never rises.

solid answer

~40 s

Autoscaling scales whatever you measure, not whatever hurts. This worker is dominated by waiting - fetching the source object, writing the result back - so per-copy processor use sits comfortably under target while unserved work accumulates. The loop is behaving correctly against a signal that is blind to the problem. The fix is to scale on a signal that *is* the problem: queued items per copy, or the age of the oldest queued item. The test for any candidate signal is two-part: does it rise when the workload falls behind, and does it fall when a copy is added? Processor use here fails the first test; an upstream dependency's error rate fails the second.

go deeper

for a junior

Remember that autoscaling acts on the one number it was given. If that number stays low while work piles up, the loop will never add copies - and it is not misbehaving.

for a middle

Explain why wait-dominated work hides from a processor signal, and state the two tests any candidate signal must pass: it rises when the fleet falls behind, and it falls when a copy is added.

for a senior

Diagnose from the signal outward: compare the chosen measurement against backlog over the same window, check whether the signal is saturated, and check whether a sick copy is dragging the fleet average down.

for a principal

Treat signal selection as a platform standard, not a per-team choice. Left to defaults, every team scales on the metric that was already being collected, and every wait-dominated service inherits the same silent failure.

## The loop is not broken - the signal is Nothing here is malfunctioning. The scaling loop read its signal, compared it to its target, found the ratio near one, and left the count alone. It did exactly what it was configured to do. The defect is upstream of the loop: the number chosen to represent "how hard is this workload being pushed" does not represent that at all for this workload. A rendering worker's wall-clock time is mostly **waiting**: pull the source object over the network, decode, resize, write the result back, acknowledge the queue item. Only the middle step burns processor cycles. A copy can be fully occupied - every worker thread assigned to an item, nothing idle in any sense the operator cares about - while per-copy processor use reads 25%. Meanwhile the queue grows, because arrivals exceed completions. The signal and the pain are decoupled. ## The two properties a scaling signal must have Check both before adopting a signal, in this order: 1. **Does it rise when the workload falls behind?** It must be a measure of *unserved demand* or of how close a copy is to its own limit - not of one resource that may or may not be the limiting one. 2. **Does it fall when a copy is added?** The loop's convergence depends on its own action moving the number. A signal the workload cannot influence gives the loop no feedback, and it will either sit still or climb to its ceiling and stay there. | Candidate signal | Rises when behind? | Falls when a copy is added? | Verdict here | |---|---|---|---| | Processor use per copy | No - the work is wait-dominated | Yes | Blind; never triggers | | Queued items per copy | Yes | Yes - same backlog, more copies | The right signal | | Age of the oldest queued item | Yes, and it is the user-visible symptom | Yes, with a lag | Good, slightly laggier | | Memory used per copy | No - it tracks buffers and caches, not backlog | Sometimes not at all | Wrong axis | | Error rate of the object store the worker reads | Yes, but it is a neighbour's fault | No - more copies make it worse | Never a scaling signal | The last row is the opposite failure and it is worth naming, because it is just as common: scaling on something the copies cannot affect. The loop reads over target, adds copies, reads over target again, adds more, and runs to the ceiling - now with a large fleet hammering a dependency that was already failing. ## Picking the signal for wait-dominated work For a worker fed by a queue, the honest signal is almost always **unserved work divided by the number of things serving it**. Two usable forms: - **Backlog per copy.** Total queued items divided by current copy count, against a target such as "ten items per copy". It composes with the proportional recompute rule directly and it responds immediately to a count change. - **Oldest-item age.** The time the head of the queue has been waiting, against a target derived from the latency you promised. It is closer to what a user experiences, but it moves later than backlog does - the age only starts climbing after the backlog has already been growing for a while. For request-serving work that is also wait-dominated, the equivalent is **in-flight requests per copy** - concurrency rather than processor use. It captures a copy that is saturated on connections, thread pool slots or downstream waits, none of which show up as processor time. ## Related traps in the same family - **A signal that saturates.** Once every copy is pinned at 100% processor use, the signal cannot go higher no matter how far behind the fleet falls. The ratio stays fixed at 100/target, so the loop adds a fixed proportion and then stops, even though the backlog is still growing. A backlog signal keeps climbing and keeps driving. - **A signal a sick copy drags down.** A copy that has stopped doing work reports a *low* value and pulls the fleet average down, which can talk the loop into removing capacity during an incident. - **Averages hiding a skew.** If routing is uneven, the average across copies can look healthy while some copies are saturated. Per-copy averages assume work is spread roughly evenly. - **Composing signals.** Where a workload is genuinely limited by different things at different times, compute a recommendation per signal and take the highest. Any one signal over target then justifies more copies, while shrinking needs all of them to agree. How far ahead of the backlog you choose to run - how much headroom to hold and what to do once you are past capacity - is capacity planning and load shedding, which live elsewhere. What belongs here is narrower and sharper: the number you scale on must be the number that hurts.

  • Why is scaling on a dependency's error rate worse than scaling on nothing at all?
    Because the loop has no feedback path. Adding copies cannot lower a neighbour's error rate, so every evaluation still reads over target and the count climbs to its ceiling. The result is a maximum-size fleet applying maximum pressure to something that was already failing, which usually deepens the outage instead of relieving it.
  • The worker is pinned at 100% processor use per copy and still falling behind. Why does the processor signal stop driving the count up?
    Because the measurement is bounded and the backlog is not. Once every copy reads 100 against a target of, say, 70, the ratio is stuck at about 1.43 - the loop asks for 43% more copies once and then sees the same fixed ratio forever. A backlog-per-copy signal keeps rising with the queue, so it keeps producing larger recommendations.
  • What should you check before trusting an average-per-copy signal at all?
    That work is actually spread evenly. The per-copy average assumes it is; with sticky connections, uneven partition assignment or a hashing scheme that skews, some copies can be saturated while the average reads comfortable. Look at the spread across copies, not just the mean, before you set a target against it.

saying these in an interview costs you the question

  • Assumes a workload that is falling behind must be short of processor time
  • Picks the signal that is easiest to collect rather than the one that hurts
  • Scales on a dependency's health, which no number of copies can improve
  • Trusts a per-copy average without checking that load is spread evenly
  • Thinks a saturated signal keeps driving the count up as the backlog grows
  • Believes a low signal always means the workload is comfortable