skip to content

A plant-wide fault makes every machine alarm at once and readings jump to 30,000 per second - what does autoscaling lag cost your 100-instance scoring fleet?

level: seniorimportance: should knowfreq 38%

answer

  1. the surge is correlated, not random
  2. onset in seconds, reaction in minutes
  3. excess rate times lag equals backlog
  4. draining is slower than filling
  5. no errors, only late readings

basics

~20 s

Roughly a five-minute backlog. A 20,000-per-second fleet offered 30,000 accumulates 10,000 readings a second, so a four-minute autoscaler lag builds 2.4 million queued readings that take minutes to drain - every one of them scored long after its operator window closed.

solid answer

~50 s

Do the arithmetic. The fleet's ceiling is `100 x 200 = 20,000` readings per second and the storm offers 30,000, so the excess is 10,000 per second. If the autoscaler's new capacity takes about four minutes to become ready - metric window, decision, instance start, artifact load, warm-up - the backlog at that point is `10,000 x 240 = 2.4 million` readings. With the pre-storm fleet alone, draining it at `20,000 - 12,000 = 8,000` readings per second takes another 300 seconds; the late instances shorten that, but they cannot un-late a reading that has already waited. Nothing was dropped and every reading was still scored minutes after the fault it describes. The deeper point is that this surge is **correlated with the event you are detecting**: an anomaly-scoring tier is loaded hardest exactly when its output matters most, so the surge has to be in the standing fleet's sizing, not delegated to an autoscaler whose reaction time is longer than the surge's onset.

code

pseudocode · 17 lines
pseudocode
ceiling      = 20000   // readings per second, 100 instances x 200
stormRate    = 30000   // readings per second during the plant-wide fault
stormSeconds = 240     // storm length, and roughly the autoscaler's lag
normalRate   = 12000   // readings per second once the storm passes
deadlineMs   = 250

excess  = stormRate - ceiling        // 10000 readings per second
backlog = excess * stormSeconds      // 2,400,000 readings queued, none dropped

// drain with the pre-storm fleet alone; late instances shorten this
drainRate    = ceiling - normalRate  // 8000 readings per second
drainSeconds = backlog / drainRate   // 300 seconds

if drainSeconds * 1000 > deadlineMs:
    report("storm class misses the deadline: size the standing fleet for it")
else:
    report("backlog clears inside the deadline")

go deeper

for a junior

Remember that adding machines is not instant. A fleet already over its ceiling keeps accumulating a backlog for the whole time new capacity is being brought up, and that backlog has to drain afterwards.

for a middle

Be able to compute it: excess rate times lag gives the backlog, and the spare capacity after the surge gives the drain rate. Show that the drain takes longer than anyone expects.

for a senior

Point out that the surge is correlated with the event being detected, so it is a sizing commitment rather than a scaling one, and note that no error-rate alarm will ever fire on it.

for a principal

Make someone own the trade. Standing capacity for a storm costs real money every idle hour, so name which surge classes must be scored inside the deadline and get that written down before the fault, not during it.

## The surge is correlated with the thing you are detecting Most capacity work assumes arrivals are roughly independent: users turn up separately, so the aggregate rate is smooth and its peaks are predictable from a daily shape. A sensor estate breaks that assumption in the worst possible way. A plant-wide fault trips thousands of machines within seconds, and every gateway forwards a burst at the same moment. The arrivals are **correlated with each other and with the event the model exists to catch**, so the scoring tier is offered its highest load in exactly the minutes its output is most valuable. That correlation, not the raw multiple, is what makes this a sizing problem rather than a scaling one. ## Doing the arithmetic | quantity | value | where it comes from | |---|---|---| | fleet ceiling | 20,000/s | 100 instances x 200 readings per second | | normal peak | 12,000/s | the design peak the fleet was sized for | | storm rate | 30,000/s | correlated arrivals during the fault | | excess | 10,000/s | `30,000 - 20,000` | | autoscaler lag | ~240 s | metric window, decision, start, artifact load, warm-up | | backlog at four minutes | 2.4 million | `10,000 x 240` | | drain time, pre-storm fleet alone | 300 s | `2,400,000 / (20,000 - 12,000)` | Against a 250 ms deadline, a reading sitting in a backlog that takes minutes to clear is not slightly late. It is late by three orders of magnitude, which for an anomaly alert means the operator's window has closed. ## Why the autoscaler is the wrong instrument here An autoscaler is a **cost** instrument. It is very good at following a daily traffic shape so that you are not paying for the night-time peak all day. It is poor at absorbing a surge, for reasons that stack: - The signal it watches is itself an average over a window, so the surge has to persist before it is visible at all. - The decision is deliberately damped, because a scaler that reacts to every spike oscillates. - New instances have to start, pull a model artifact, load it into memory and warm any caches before they can take traffic. - The capacity therefore arrives **after** the surge has already built a backlog, and in this scenario roughly when the storm ends. Onset measured in seconds against reaction measured in minutes is a losing race, and no tuning closes a gap of that shape. ## The three honest sizing responses 1. **Put the surge in the standing fleet.** Size for the storm rate rather than the daily peak. At 30,000 readings per second and the same 60% target that is `30,000 x 0.04 / 8 / 0.6 = 250` instances - two and a half times the fleet, idle most of the time. Expensive, and sometimes correct, because the alert this tier produces is the reason the tier exists. 2. **Pre-scale on a leading signal.** If something upstream sees the fault before the scoring tier does - a gateway's own alarm counter, a shift schedule, a planned test - capacity can be added ahead of the surge rather than in response to it. This only works when a genuine leading indicator exists; a lagging one just reproduces the autoscaler's problem. 3. **Decide, in advance, what the storm class gets.** Accept that a correlated storm is scored late and state which readings are scored first when the backlog forms. That prioritisation decision is a real design commitment and it belongs in the runbook, not in whatever order the queue happens to have. What you cannot do is write "the autoscaler handles it" in the design and move on. ## What "nothing was dropped" hides The most dangerous property of this failure is that it produces no errors. The gateways buffer and retry, so the count of lost readings stays at zero and every throughput dashboard shows the tier serving at its ceiling - which looks like success. The damage is entirely in a latency distribution, and only for the readings that arrived during the storm. If the tier's alarms are built on error rate and throughput, this outage is invisible until an operator says the alert arrived after the machine had already stopped. ## Sizing to the surge you must not miss The question to take back to the business owner is not "how big should the fleet be" but "which surges must be scored inside the deadline". A storm that is a nuisance can be allowed to queue. A storm that is the safety case cannot, and the fleet has to carry standing capacity for it whatever its idle cost. Write the answer down as an explicit class of traffic with an explicit commitment, because the alternative is discovering the commitment during the fault.

  • What would the standing fleet have to be to absorb a 30,000-per-second storm at the same 60% target?
    Concurrency is `30,000 x 0.04 = 1,200` readings in service, which is `1,200 / 8 = 150` saturated instances and `150 / 0.6 = 250` provisioned - two and a half times the 100-instance fleet, sitting largely idle outside the storm. That is the real price of the commitment, and it is a business decision about how much a late anomaly alert costs.
  • The autoscaler is retuned to react in 60 seconds. Does that solve it?
    It reduces the damage without removing it. A 60-second lag still builds `10,000 x 60 = 600,000` queued readings, which the pre-storm fleet drains in 75 seconds - still hundreds of times the 250 ms deadline. Tightening the reaction also makes the scaler twitchier on ordinary traffic, so it trades one problem for another.
  • Why does throughput monitoring show this tier as healthy throughout the storm?
    Because it is serving at exactly its ceiling, which is the maximum a healthy tier can do, and the gateways buffer the excess so nothing is recorded as lost. The failure lives entirely in the latency distribution of the readings that arrived during the storm, so only a per-reading age or queue-depth signal makes it visible.

saying these in an interview costs you the question

  • Says the autoscaler makes standing capacity unnecessary.
  • Sizes for the smooth daily peak and calls a correlated storm an outlier.
  • Believes the backlog disappears the moment new instances arrive.
  • Counts the incident as clean because no readings were lost.
  • Assumes a newly created instance can take traffic immediately.