How do you establish that a second, larger request peak after an applied surge was self-inflicted rather than real demand?
answer
- Compare what was sent with what arrived
- The applied rate is on record
- Look at identifiers, not just counts
- Synchronised arrival at a fixed delay
- Recovery cannot finish while it self-feeds
basics
~20 sCompare what the load generator offered against what the service received: the offered rate is known and unchanged, so any excess was manufactured inside the system. Repeats of identifiers already submitted, arriving synchronised at a fixed delay, confirm it.
solid answer
~50 sThree lines of evidence settle it. **Offered against observed**: the applied arrival rate is a known quantity because the run produced it, so plotting it against the rate the service actually received shows immediately whether requests exist that nobody offered. **Identity**: if each unit of work is tagged at submission, the second peak can be shown to consist of identifiers already submitted, answered or refused, with a repeat count per identifier. **Shape**: demand from real people is dispersed, while a manufactured wave is sharply synchronised at a near-fixed delay after the failures that produced it, and is often taller than the applied peak because it compounds at each hop. This belongs to the recovery half of the run: while the system feeds itself, drain time is not merely long but unbounded, so a run that stops measuring with the applied demand reports a false pass.
code
pseudocode · 16 linesseries offered(t) # rate the run applied, known exactly
series observed(t) # rate the service received
amplification(t) = observed(t) / offered(t)
suspect_wave = interval where amplification(t) > 1 + measurement_slack
and offered(t) is flat
within suspect_wave:
repeat_share = fraction of requests whose identifier
was already submitted before the interval
arrival_delay = suspect_wave.start - last_failure_burst.end
dispersion = spread of arrival times within the interval
report amplification_peak, repeat_share, arrival_delay,
originating_hop, effect_on_drain_timego deeper
Be ready to recall that requests reaching a service can outnumber the ones a run actually sent, because failures can cause the same unit of work to be sent again, and that both figures therefore need recording.
An interviewer expects the comparison named — applied rate against received rate — and the reason a manufactured peak arrives synchronised at a fixed delay rather than dispersed the way demand from real people is.
Show that you tag work so repeats are countable, instrument every hop because the effect compounds along a call path, keep measuring long past the applied peak, and connect the wave to why recovery never completed.
Own the position on how much amplification a platform will tolerate before calling behaviour becomes a platform-level concern rather than each team's, and what evidence a run must produce to justify that call.
## The observation The applied profile of a surge run finishes and demand is removed. Minutes later the number of requests reaching the service rises again — often higher than the applied peak — with no corresponding rise in what the load generator sent. That second wave was manufactured somewhere between the caller and the service, and recognising it as manufactured is the entire finding. Until it is named, the run reads as a system that failed at a load nobody applied. ## Three lines of evidence **1. Offered against observed.** The applied arrival rate is a known quantity, because the run produced it. Plot it against the rate the service actually received. Through a healthy run the two track each other with a small constant difference. A divergence means requests exist that nobody offered. This one comparison usually ends the argument, and it is why a surge run should always record both series rather than only the system-side count. **2. Identity.** Requests in the second wave carry identifiers that were already submitted, already answered or already refused during the peak. If each unit of work is tagged at submission, the wave can be shown to consist of repeats rather than new work, and the repeat count per identifier says how many times each unit was re-sent. **3. Shape.** A wave produced by real people is smooth and dispersed. A manufactured one is sharply synchronised: it appears at a near-fixed delay after the failures that produced it, because everything that re-sent waited the same interval, and it is frequently taller than the original peak because one failed unit yields several repeats, multiplied again wherever a further hop re-sends independently. | Feature | Real demand | Manufactured wave | |---|---|---| | Offered rate | rises together with the observed rate | flat while the observed rate rises | | Identifiers | new units of work | repeats of units already submitted | | Arrival shape | dispersed across the window | sharply synchronised at a fixed delay | | Peak height | bounded by how many people act | can exceed the original applied peak | | When demand stops | subsides on its own | can sustain itself indefinitely | ## Why this belongs to the recovery half of the run A self-inflicted wave is the most common reason a system that "survived the peak" never recovers. The peak causes failures, the failures produce repeats, the repeats cause more failures, and the loop holds the system in the failed state after the real demand is long gone. Recovery cannot complete while the system is feeding itself, and the drain time is not merely long — it is unbounded. A run that stops measuring when the applied demand stops will report a comfortable pass and miss the whole effect, which is the concrete argument for a post-surge observation window several times longer than the peak. ## Designing the run so it is visible at all - Record the offered arrival rate as a first-class series, not just the count the service saw. - Tag units of work at submission with identifiers that survive being re-sent, so repeats can be counted rather than inferred. - Instrument every hop rather than only the entry point. Amplification compounds along a chain, so a wave that is barely visible at the edge can be several times the original volume at the deepest dependency. - Keep measuring long enough for the delayed wave to arrive. Its delay is set by whatever interval the re-sending layer waits, which can be far longer than the peak itself lasted. - Separate the three possible sources: the original callers, an intermediate layer re-sending on their behalf, and people whose own request appeared to fail and who tried again by hand. All three produce a similar shape and each needs a different owner. ## Reporting it The run's output here is a description, not a design. Report the second wave's volume as a multiple of the applied peak, where in the call path it originated, the delay at which it arrived, and its effect on each recovery assertion — drain time, resource release, and whether the recovered state held. What the calling behaviour ought to become is a decision for whoever owns that behaviour; keeping the run to what it measured is what makes the finding credible when it lands on another team. Two reporting traps are worth naming. The first is treating the wave as a fault of the run itself — "the generator misbehaved" — when the applied rate is on record and never changed. The second is treating it as the system failing at a higher load than intended: the applied load was not higher, the system created the difference, and the number that matters is the ratio between what was offered and what arrived. ## What it changes about the result Strictly, a run in which the observed rate exceeded the offered rate did not test the profile that was written. The system experienced a larger and differently shaped event than the one designed, so the amplification ratio becomes a required part of the result. Without it recorded, a later reader comparing this run against another is comparing two different experiments and will attribute the difference to whatever changed in the code.
- Why can the second peak be several times taller at a deep dependency than at the entry point?Because the multiplication compounds along the call path: each hop that re-sends independently multiplies what the hop above it already multiplied. A modest excess measured at the edge can therefore be a very large one at the deepest dependency. That is why the run instruments every hop rather than only the entry point, and why the amplification figure is reported per hop rather than as a single number.
- What do you actually report as the finding, and what do you leave to someone else?Report what was measured: the wave's volume as a multiple of the applied peak, where in the call path it originated, the delay at which it arrived, and its effect on drain time, resource release and whether the recovered state held. What the calling behaviour ought to become is a decision for whoever owns that behaviour, and keeping the run to its measurements is what makes the finding land.
saying these in an interview costs you the question
- Blames the load generator for a rate rise it never applied
- Reads the wave as the system failing at a higher applied load
- Only records the service-side count, never the offered rate
- Ends measurement when the applied demand ends
- Assumes a second peak means real users returned
- Instruments the entry point alone, so amplification stays invisible