A two-minute traffic spike is over before any new copy serves a request - where did the time go?
answer
- reactive, never predictive
- the window averages and delays the spike
- started is not the same as serving
- readiness plus warm-up is the last minute
- measure each stage before tuning one
basics
~20 sReactive scaling has a chain of delays: an averaged measurement window, the evaluation interval, placement and image pull, process start-up, then the readiness check and warm-up before traffic is routed. Two to five minutes is normal, so a two-minute spike finishes first.
solid answer
~40 sEvery stage between load arriving and capacity serving costs time, and they add up. The signal is an average over a window, so a spike is both delayed and damped before the loop ever sees it. The loop then evaluates on its own interval. The new desired count has to be turned into running copies: find a host, pull the image if it is not cached, start the process. Then the copy is still not serving - it takes traffic only once its readiness check passes, and warm caches and connection pools may lag that. Measure each stage separately before tuning: shortening the averaging window buys seconds, but if the real cost is a 90-second warm-up, nothing upstream of it matters.
go deeper
Know that scaling is not instant. A raised copy count still has to become started, ready and routed-to before any user benefits, and that takes minutes rather than seconds.
Name the stages in order - measurement window, evaluation interval, placement and start, readiness and warm-up - and explain why the average both delays and shrinks the spike the loop sees.
Show that you would measure each stage before tuning, identify the dominant term, and recognise that a burst shorter than the total lag can only be absorbed by capacity that is already running.
Set the expectation across services: reactive scaling tracks curves, not bursts. Decide deliberately which workloads get warm headroom, which get a queue, and what that costs the estate.
## The chain, stage by stage Reactive autoscaling cannot serve a spike faster than the sum of its own delays. Walk the chain with a concrete set of assumptions - these are illustrative numbers, not defaults, and yours will differ: | Stage | What is happening | Typical cost | |---|---|---| | Measurement window | The signal is an average over a rolling window, so a spike is visible only partly, and late | 30-60 s | | Collection and evaluation interval | The number is scraped or pushed on a period, and the loop runs on its own period | 15-60 s | | Decision to desired count | The recompute plus any gate that requires the reading to hold | 0-30 s | | Placement and start | A host with capacity is found, the image is pulled if not already on that host, the process starts | 10-120 s | | Readiness and warm-up | The copy reports ready and is added to routing; caches fill, connections are established | 5-120 s | Even a well-tuned chain lands around two minutes, and an untuned one around five. Against a spike that lasts two minutes, the capacity arrives to serve traffic that has already gone. ## The averaging window is the sneakiest term An average over a rolling window does two things to a spike, and engineers usually only account for the first: 1. **It delays it.** The window has to fill before the reading reflects the new level. 2. **It shrinks it.** A 60-second window that is half through a spike reports roughly halfway between the old level and the new one. The loop therefore computes a *smaller* ratio than reality justifies, asks for fewer copies than needed, and only reaches the right count an evaluation or two later - by which time the spike may be over. That damping is deliberate: the same averaging is what keeps ordinary noise from jerking the count around. It is the first half of a genuine trade-off, and the second half shows up on the way down as oscillation. ## Readiness is where the last, invisible minute hides A started copy is not a serving copy. It is routed to only when its readiness check passes, and that check is the right place for the warm-up to be reflected: an empty in-process cache, a connection pool that opens lazily, a large configuration or model loaded at start. A copy that reports ready at second three and then serves its first hundred requests slowly has simply moved the warm-up cost onto users instead of into the chain. Both cost the same wall-clock time; only one is visible on a dashboard. ## What actually shortens it Measure each stage before tuning any of them - the fix is worth nothing if it is not applied to the dominant term. - **Shorten the measurement window and evaluate more often.** Cheapest lever, buys tens of seconds. The cost is noise: a twitchier signal makes the count oscillate, which is what a settling window then has to damp. - **Make start-up cheap.** A smaller image pulls faster; an image already cached on the host does not pull at all; a process that opens its connections in parallel rather than serially reaches ready sooner. - **Reduce the warm-up, or admit it.** Pre-load what the copy needs before reporting ready, so the platform's own view of "ready" is honest even though it takes longer. - **Absorb the first minute with capacity that is already running.** This is the only real answer for a spike shorter than the chain. Either the fleet already has spare room at its steady-state target, or a floor of warm copies exists for this purpose. How much of that to hold is a capacity-planning decision, owned elsewhere - but recognising that *no reactive loop can beat its own lag* is the reasoning this question is testing. - **Turn the spike into a backlog.** Where the work is queued rather than served synchronously, the spike becomes a temporarily longer queue instead of a wall of failures, and late capacity still does useful work. That is the structural reason queue-fed workloads tolerate reactive scaling far better than request-serving ones do. ## The judgment underneath The interviewer is usually checking one thing: do you know that autoscaling is **reactive**, and therefore that its value is in tracking load over minutes and hours - the daily curve, the weekly one, a sustained shift - and not in catching a two-minute burst? A team that responds to this incident by making the loop more aggressive usually trades a spike problem for a flapping problem and still misses the next spike. The stage-by-stage measurement is what turns that instinct into a decision.
- Why does a rolling average make the loop ask for too few copies during the first minute of a spike?Because the reading is a blend of before and during. Halfway into a spike, a rolling window reports roughly halfway between the old level and the new one, so the observed-to-target ratio is smaller than reality justifies and the recomputed count lands short. The loop converges on the right number an evaluation or two later, which is exactly the time the spike does not give it.
- A copy reports ready three seconds after start but serves its first requests very slowly. What went wrong?The readiness check is not covering the warm-up. Cold caches, a lazily opened connection pool or configuration still loading mean the copy is running but not usefully serving. The platform routes traffic to it because it said it was ready, so the warm-up cost lands on users instead of on the rollout. Make readiness reflect the work the copy has to do before it is genuinely useful.
- Why do queue-fed workloads tolerate this lag better than request-serving ones?Because the work waits instead of failing. A queue turns a spike into a longer backlog, so capacity arriving three minutes late still completes every item - the cost is latency, not lost work. A synchronous request path has no such buffer: requests offered while the fleet is short either queue inside the server, time out, or are rejected.
- Is making the loop more aggressive a reasonable response to this incident?Only after measuring which stage dominates. Shorter windows and more frequent evaluation shave seconds off the first two stages, and if start-up plus warm-up is ninety seconds, that fix changes nothing while making the count oscillate. Aggressive tuning without stage-level measurement usually trades a spike problem for a flapping one.
saying these in an interview costs you the question
- Assumes new copies serve traffic the moment the loop raises the count
- Counts only start-up time and forgets the averaging and evaluation windows
- Treats a copy as serving as soon as its process has started
- Tunes the evaluation interval without measuring which stage dominates
- Expects a reactive loop to absorb a burst shorter than its own lag
- Thinks the averaging window only delays the signal without damping it