After a surge in demand subsides, what must be true before you call the system recovered?
answer
- Demand falling flatters the obvious number
- Four claims, not one
- Drain time is the headline figure
- Watch for a resource settling higher
- Hold the state across a defined window
basics
~20 sRecovery needs more than response times falling. The backlog must drain back to its pre-surge depth, resources must be released rather than settling higher, accepted work must still complete, and that state must hold across a defined observation window.
solid answer
~50 s"Response times came back" can be true while the system is still broken, because response time is driven by demand — remove the demand and the figure falls whatever state the system is in. Assert four things instead. **Backlog**: every queue, buffer or store of deferred work that grew during the peak returns to its pre-surge depth, and the time that took is recorded, because drain time is how long the system stayed impaired after the event ended. **Correctness**: work accepted during the peak still completes, and work that could not be taken was refused explicitly rather than accepted and abandoned. **Resources**: connections, worker slots and memory fall back to the pre-surge band instead of settling on a new plateau. **Stability**: that recovered state holds across a defined observation window, not for one sample. A run with no post-surge window has produced no recovery result at all.
code
pseudocode · 10 linesassert_recovery(pre_surge, post_window):
backlog_drained = pending_depth_returned_to(pre_surge.pending_depth)
drain_duration = time_from(demand_removed) to backlog_drained
work_intact = submitted == completed + refused + pending
resources_freed = post_window.resource_use within pre_surge.band
state_held = every_moved_quantity flat for longer than peak_duration
report drain_duration, resources_freed, work_intact, state_held
if measurement_ended_before(state_held):
result = "no recovery result" # not a passgo deeper
Be ready to recall that a surge run measures both the peak and the period afterwards, and that the system is not considered recovered simply because it stopped being slow once demand went away.
An interviewer expects the separate claims spelled out — backlog drained with a recorded duration, work neither lost nor repeated, resources released, and the state sustained — plus why response time alone repairs itself automatically.
Demonstrate that you size the observation window from the slowest recovery mechanism in the system, recognise a new resource plateau as a finding rather than a pass, and can name which quantity failed to return when recovery is incomplete.
Own how much post-surge impairment the product is willing to carry, and the tradeoff between long observation windows that catch delayed effects and a testing cadence a team can actually sustain.
## The trap in "response times came back" Response time is a function of demand. Take the demand away and the number falls whatever state the system is in — an exhausted worker pool serving one request per second reports excellent response times, and so does a system that has quietly discarded half of the work it accepted. The recovery half of a surge run exists precisely because the quantities that look worst during the peak are the ones that repair themselves automatically once the peak ends, while the quantities that actually matter are the ones that do not. So recovery is asserted, not noticed in passing. Asserting it needs three things: a defined observation window after demand is removed, a set of pre-surge figures to compare against, and four separate claims. ## The four claims **1. The backlog drained.** Anything that grew because demand exceeded service rate — a pending-work queue, an inbound buffer, a store of deferred work, a batch waiting to be written — returns to the depth it had before demand rose. Just as important, the time it took to get there is recorded. Drain time is the closest thing a surge run has to a headline number, because it says how long the system stays impaired after the event is over. A system whose response times recover in ten seconds but whose backlog takes forty minutes to clear was unavailable to the people in that backlog for forty minutes. **2. Nothing was silently lost or repeated.** Work accepted during the peak still completes afterwards, and work that could not be taken was refused explicitly rather than accepted and abandoned. This is established from counts and identifiers, never from the absence of failures in a log. **3. Resources were released.** Connections, worker slots, memory, open files, cached entries and background tasks return to the pre-surge band rather than settling on a new plateau. A permanent step up in resource use after demand is gone is the classic surge finding: the system survived but did not let go, and the next surge starts from a worse position than the last one. **4. The recovered state held.** A single sample after demand falls is not a recovery. The claim is that the system stayed inside the pre-surge band for a defined observation window — long enough for deferred work, delayed waves, scheduled flushes and background compaction to arrive and pass through. | Claim | Evidence to record | What a false pass looks like | |---|---|---| | Backlog drained | pending depth over time, drain duration | depth still falling when the run stopped | | Nothing lost or repeated | counts and distinct identifiers on both sides | totals matched but were never itemised | | Resources released | pre- and post-surge resource figures | a new higher plateau read as "stable" | | State held | the whole post-surge window, not one sample | measurement ended with the applied peak | ## Why the observation window has to be designed The window has to outlast the slowest recovery mechanism in the system, which is not a fixed number: it depends on the deepest queue, the longest scheduled interval and the slowest downstream dependency. A workable rule is to keep measuring until every quantity that moved during the peak has been flat for longer than the peak itself lasted, and to treat a run that ended before that as having produced *no* recovery result rather than a passing one. Where recovery does not complete, the run still produces something useful: it names which quantity failed to return, and how far from returning it was when measurement stopped. ## What the run itself does not decide Whether the recorded drain time, the residual resource use and the post-surge response time distribution are acceptable is a separate decision, agreed before the run by whoever owns the pass rule for that product. The run's job is to produce the four claims with evidence attached, in a form that decision can actually be made from. Blurring the two — looking at the numbers at the end of a run and deciding they seem fine — is how a surge result quietly becomes an opinion. ## Four common shapes of failure - **The sawtooth.** Response times recover, then rise again with no new demand — usually work the system deferred during the peak arriving later. - **The plateau.** Every rate returns to normal but one resource does not, so each successive surge begins from a higher floor than the one before. - **The silent drop.** Everything recovers beautifully because the system stopped taking work partway through the peak and never recorded that it had. - **The short window.** The run ends minutes after the peak, before the deepest queue has drained, and reports the fastest-recovering quantity as though it were the whole result. Naming which of these a run produced is more useful to a team than any single recovered number, because each one points at a different mechanism and a different fix.
- How long should the post-surge observation window be?Long enough to outlast the slowest recovery mechanism in the system, which depends on the deepest queue, the longest scheduled interval and the slowest downstream dependency. A workable rule is to keep measuring until every quantity that moved during the peak has been flat for longer than the peak itself lasted, and to treat a shorter run as producing no recovery result rather than a passing one.
- The system recovered on every measure except memory, which settled higher and stayed there. What does that mean for repeated surges?Each surge then starts from a worse position than the last. The finding is not a single run's result but a trend, so the run is repeated back to back at the same profile: if the new floor rises with every repetition, the system releases nothing and a sequence of ordinary surges eventually reaches exhaustion that no single run would ever have shown.
A river dropping back inside its banks is not the same as the flood being over. The water still standing in the fields is the backlog, and how long it takes to drain is the number worth reporting.
saying these in an interview costs you the question
- Calls the system recovered because response times fell after demand stopped
- Stops measuring at the moment the applied demand is removed
- Reads a new higher resource plateau as a stable, healthy state
- Ignores work still queued when reconciling what completed
- Reports one post-surge sample instead of a sustained window
- Treats refused work and dropped work as the same outcome