When extra capacity arrives minutes after a surge, how do you keep a performance run from hiding that delay?
answer
- Reactive capacity is never instant
- Name the interval between the two events
- Do not raise the size before starting
- Measure inside the window, not across the run
- Starting size belongs in the result
basics
~20 sStart at the size the system would reach under baseline demand, never pre-raised to peak, and timestamp two series: offered demand and capacity actually serving. The interval between them is the exposure window, and what happened inside it is the result.
solid answer
~50 sCapacity that reacts to demand arrives on a delay — something notices, something starts, and it must warm and register before it serves — and that delay is usually minutes against a jump of seconds. The **exposure window** runs from demand rising to added capacity actually serving, and the run either measures it or erases it. Erase it by pre-raising the system to peak size before starting, by rehearsing at peak rate first, or by starting measurement once figures have settled. Measure it by reaching the starting size through *running* at the baseline rate, timestamping offered demand and serving capacity as two independent series, and computing response times, refusals, pending depth and completions **for that window alone** rather than averaged across the run. A system pre-raised before demand rose has demonstrated that the larger configuration serves that load — a useful fact, but not a surge result.
code
pseudocode · 15 linesseries offered_demand(t) # what the run applied
series serving_capacity(t) # units actually accepting work
exposure_start = first t where offered_demand(t) > baseline_rate
exposure_end = first t where serving_capacity(t) > baseline_capacity
and is_accepting_work(t)
exposure_window = exposure_end - exposure_start
within exposure_window, report:
response_time_distribution, refused_count,
pending_depth_peak, completed_count,
unserved_demand classified as:
served_slowly | queued | refused_explicitly | accepted_then_dropped
result.starting_size = capacity at exposure_start # part of the resultgo deeper
Be ready to recall that capacity added in response to demand does not appear instantly, and that a system may therefore spend a period serving a surge with only the capacity it already had.
An interviewer expects you to name the interval between demand rising and capacity serving, and to explain why starting a run already raised to peak size answers a steady-demand question instead of a surge one.
Show that you reach the starting size by running at the baseline rate, timestamp demand and serving capacity as separate series, compute figures for the exposure window alone, and classify what happened to demand that could not be served inside it.
Own the tradeoff between provisioning ahead of a predictable surge and relying on capacity that reacts to it, and decide how much exposure the product will accept before that reliance stops being defensible.
## The gap the run has to expose Systems that add serving capacity in response to demand do it on a delay: something has to notice, something has to be started, and the new capacity has to become useful — warmed, connected, registered and actually receiving work. That delay is frequently measured in minutes, while the jump in demand it is responding to is measured in seconds. The whole point of a surge run against such a system is the interval in between, when demand is at its peak and capacity is still whatever it was before. Call it the **exposure window**: from the moment demand rises to the moment added capacity is genuinely serving. Everything that decides whether the surge was survivable happens inside it, and a run design either measures it or erases it. ## Three habits that erase the window 1. **Pre-provisioning for the run.** Raising the system to peak size before starting, "so the test is fair", turns a surge run into a steady-state measurement at a large size. It shows the system can serve that demand. It shows nothing whatever about absorbing its arrival. 2. **Rehearsing at high demand first.** A warm-up pass at the peak rate leaves the extra capacity already added when the measured pass begins, so the measured surge lands on a system that has already reacted to an earlier one. 3. **Starting the clock late.** Beginning measurement once figures have stopped moving skips over precisely the window the run exists to observe, and reports the recovered state as the whole result. ## Designing so the delay is measured - Start from the size the system would genuinely be at under the baseline arrival rate, reached by *running* at the baseline rate rather than by configuring the size directly. - Timestamp two independent series: offered demand, and serving capacity that is actually accepting work. The interval between the rise in the first and the rise in the second is the exposure window, and it is a reported number in its own right. - Compute the response time distribution, refusal count, pending-work depth and completion count **for the exposure window alone**. Averaged across a twenty-minute run, a two-minute exposure disappears into comfort. - Record which of four things happened to the demand that could not be served inside the window: served slowly, queued, refused explicitly, or accepted and dropped. These are very different outcomes and only the last is a defect on its face. - Repeat with the transition time varied. If demand arrives faster than capacity can be added, the exposure window is bounded by the capacity mechanism; if slower, it is bounded by demand. Which regime the product lives in is the useful output. ## What a pre-warmed pass has actually demonstrated | Run design | What it proves | What it cannot say | |---|---|---| | Started at baseline size, capacity added during the run | behaviour through the exposure window | nothing about sizes never reached | | Raised to peak size before demand rose | the larger configuration serves that demand | anything about absorbing the arrival | | Capacity held fixed throughout | behaviour of the system as it stands today | how the reactive mechanism behaves | A pre-warmed pass is a legitimate run and answers a real question — can the larger configuration serve this demand? — but it has to be reported as that question's answer and not as a surge result. The confusion is common enough that the size the system started at belongs in the result itself, not in the setup notes: **what the system started as is part of what was measured.** ## Reading the outcome honestly If the system survived only because capacity was raised before demand rose, the honest finding is that its surge behaviour is still unknown and its steady behaviour at that size is good. If it survived with capacity arriving during the run, the findings are the exposure window's duration, what happened to demand inside it, and whether recovery completed once capacity arrived — including whether that capacity was later released, and how the system behaved while it was being released. There is a second-order effect worth recording. Capacity that arrives late arrives *after* a backlog has already formed, so for a period the system is serving current demand and accumulated backlog together. Peak resource use can therefore occur several minutes after peak demand, and peak response time later still. A run that stops measuring when demand stops misses that entirely, which is a second reason the post-surge observation window has to outlast the capacity mechanism rather than merely outlasting the demand. ## What this run does not answer How much capacity should be added, what the reaction should be triggered by, and how large the system should be sized are separate decisions with their own analysis and their own inputs. This run's contribution is narrower and far more concrete: the measured length of the exposure window, and the measured consequence of demand that landed inside it.
- What has a system actually demonstrated if it only survived the surge because capacity was raised beforehand?That the larger configuration serves that demand steadily. That is a real and useful fact, but it is a steady-state result: nothing was learned about absorbing the arrival, because the arrival landed on a system that had already reacted. The honest reporting is that surge behaviour remains unknown, and the size the run started at is stated as part of the result rather than buried in setup notes.
- Why can peak resource use occur several minutes after peak demand in such a run?Because capacity arriving late arrives after a backlog has already formed, so for a period the system serves current demand and the accumulated backlog together. That combined load can exceed anything seen during the peak itself. A run that stops measuring when demand stops misses it entirely, which is why the post-surge window has to outlast the capacity mechanism and not merely the demand.
Extra staff called in when a queue forms only help from the moment they reach the counter. The queue that built while they were on their way is the part the run has to measure.
saying these in an interview costs you the question
- Raises the system to peak size before the run so the test is 'fair'
- Averages figures across the whole run, dissolving a short exposure window
- Rehearses at peak rate first, leaving capacity already added
- Reports a pre-warmed pass as a surge result
- Treats explicit refusal inside the window the same as accepted-then-dropped
- Stops measuring when demand stops, before the backlog is served