How do you parameterise a sudden surge in arrival rate so a performance run's result is attributable to it?
answer
- Three numbers, fixed before the run starts
- Where the surge starts from matters
- Settled baseline, applied multiple, transition time
- Hold the peak, then drop at once
- One variable changes per run
basics
~20 sFix three numbers before the run: the settled baseline arrival rate, the multiple applied to it, and the transition time over which the rate rises. Vary one per run, and hold the peak long enough to read a result.
solid answer
~50 sA surge is only attributable if the conditions around it are pinned first. Start from a **settled baseline**: run the ordinary arrival rate until response times, pending-work depth and resource use have stopped moving, so the pre-surge figures are a real reference rather than the tail of a system still warming. Then state the **multiple** — the peak arrival rate as a factor of that baseline, say eight times — and the **transition time** over which the rate travels from baseline to peak, because a rise spread over two minutes and a rise over two seconds are different experiments. Hold the peak long enough that the slowest normal unit of work could complete several times over, then remove demand in one move and keep measuring. Vary one of the three per run; a run that changes both the multiple and the transition time cannot say which caused what.
code
pseudocode · 12 linesprofile surge_run:
baseline_rate = 200 units per second
settle_until = response_time and pending_depth flat for 10 minutes
surge_multiple = 8 # peak = 1600 units per second
transition_seconds = 5 # baseline -> peak
peak_hold = 4 minutes
release = drop to baseline_rate in a single move
observe_after = 20 minutes still offering baseline_rate
record before surge: response_time_distribution, pending_depth,
resource_use, refusal_rate
vary_exactly_one_of(surge_multiple, transition_seconds) per rungo deeper
Be ready to say that a surge run raises arrival rate suddenly rather than gradually, that it starts from an ordinary rate the system was already serving, and that the size of the jump is described relative to that starting rate.
An interviewer expects the three numbers named — starting rate, multiple, transition time — plus why holding the peak and changing one variable per run are what make a result attributable to the surge rather than to something else.
Show that you pick the multiple and the transition time from a real event the product must survive, that you can defend the length of the peak hold, and that you record the pre-surge figures the recovery assertions will later be compared against.
Own the argument about how much sudden demand a product is designed to absorb at all, and the tradeoff between rehearsing one plausible surge thoroughly and sampling many shapes shallowly across a limited testing budget.
## What the run is trying to isolate A surge run applies a sudden, large increase in arrival rate to a system that is already serving traffic, holds it briefly, and then removes it. Almost everything it can tell you happens in two short windows: the seconds after demand jumps, and the minutes after it falls away. A result is only worth reading if those windows can be attributed to the increase itself rather than to whatever the system happened to be doing when it landed. That attribution is decided when the profile is written, not when the numbers come back. Five things belong in the profile, and the first three are what an interviewer is asking about: 1. **The baseline arrival rate** — the ordinary rate the system is serving when the increase lands, and the state it has reached at that rate. 2. **The multiple** — the peak arrival rate expressed as a factor of that baseline, for example eight times. 3. **The transition time** — how long the arrival rate takes to travel from baseline to peak: two seconds, thirty seconds, two minutes. 4. **The peak hold** — how long the peak rate is sustained before it is released. 5. **The release** — how demand is removed, and how long measurement continues afterwards. ## Why the starting state decides everything "Settled" is a testable condition, not a feeling: caches populated, connection and worker pools grown to their working size, background and scheduled work in its normal rhythm, and response time and pending-work depth flat over a defined period. Apply the increase before that and the run has two independent causes moving at once, and no amount of analysis afterwards can separate them. The settled baseline is also the reference that the recovery half of the run is compared against. "Back to baseline" is a meaningless claim if the baseline was itself a moving number, because you are comparing a post-surge state against a pre-surge state that never really existed. Recording the baseline figures explicitly — response time distribution, pending-work depth, resource use, refusal rate — turns the second half of the run into an assertion instead of an impression. ## Why a multiple rather than an absolute peak A peak stated as an absolute rate is only meaningful next to the rate already being served. Two thousand units per second is an enormous jump from two hundred and a modest one from fifteen hundred, and the system's response is entirely different. Stating the multiple keeps the run comparable across environments of different absolute capacity, and across the same environment before and after a change, and it names what actually stresses the system: the amount of *sudden extra* demand relative to what was already in flight. | Parameter | What it controls | What it decides in the result | |---|---|---| | Baseline rate | the state the system is in when demand jumps | whether "recovered" has any meaning | | Multiple | how much extra demand arrives | whether pools, buffers and queues are exceeded | | Transition time | how quickly that demand appears | whether anything reactive has time to respond | | Peak hold | how long the demand persists | whether a backlog builds or is absorbed | | Release | how demand is removed | whether the drain is observable at all | ## Transition time is the parameter people forget The same multiple applied over two seconds and over two minutes are different experiments. A gradual arrival gives pools time to grow, gives capacity that reacts to demand time to arrive, and lets a queue absorb the difference; a near-instant arrival tests only what the system already had. Neither is more correct — they answer different questions — but a result that does not state which one was applied cannot be interpreted, and two runs with different transition times cannot be compared with each other. Choose both the multiple and the transition time from an event the product genuinely has to survive: a broadcast message going out to every user at once, a scheduled opening, a promotion starting, an upstream system releasing a batch it had held, or a dependency recovering and replaying what it buffered. Where that event's real shape has been measured, derive the numbers from it; where it has not, write the assumption into the run's description so a later reader knows what was assumed rather than measured. ## Hold, release, and one variable at a time Hold the peak long enough that the slowest unit of work the system normally handles could complete several times over. A hold shorter than that measures the depth of a buffer rather than the behaviour of the system: everything is absorbed, nothing is revealed, and the run reports a comfortable pass. Then remove the demand in a single move rather than tapering it, so the drain that follows can be timed, and keep measuring for a defined window afterwards — the post-surge window is where the actual result lives. Finally, change one of these numbers per run. A run that raises the multiple and shortens the transition time together produces a difference that cannot be assigned to either. Keep a small set of profiles with one varied dimension each and the dimension a system is genuinely weak in becomes visible; vary everything at once and you get a story instead of a measurement.
- Why hold the peak arrival rate rather than dropping back the moment it is reached?A momentary touch of the peak is absorbed by whatever buffering already exists, so the run measures buffer depth instead of system behaviour and reports an easy pass. Holding the peak until the slowest normal unit of work could complete several times over is what lets a backlog form, pools saturate and any reactive capacity respond, which is where the findings actually are.
- What does it cost you if the surge starts before the baseline has settled?Two causes move at once and neither can be separated afterwards. Rising response times might be the surge or a cache still filling, pools still growing, or scheduled work still starting. You also lose the reference figures the recovery half of the run depends on, so 'returned to baseline' becomes a claim against a number that was never stable.
- How would you choose the multiple and the transition time for a product with no measured surge history?Derive them from a concrete event the product must survive — a broadcast to every user, a scheduled opening, an upstream system releasing a held batch — and estimate its shape explicitly. Then write the estimate into the run's description as an assumption rather than a measurement, so a later reader can revise it when real data arrives instead of trusting a number nobody can source.
saying these in an interview costs you the question
- Describes the surge only as 'a lot of traffic' with no multiple stated
- Applies the increase while the system is still warming up
- Changes the multiple and the transition time in the same run
- Reaches peak arrival rate and immediately drops back down
- Reports an absolute peak rate without saying what baseline it multiplied
- Treats transition time as an implementation detail rather than a parameter