skip to content

questions

13

Past a system's intended load limit, what must "degrades gracefully" become before a performance run can assert it?

level: middleimportance: must knowfreq 62%

answer

  1. a wish is not an assertion
  2. split the offered work into outcomes
  3. served, refused, never-acceptable
  4. bound the refusal, not only success
  5. name the rate each clause holds at

basics

~20 s

"Degrades gracefully" must become outcome rules written before the run: which requests are still completed and inside what response-time bound, which may be refused, how fast a refusal must arrive, and which results fail outright.

solid answer

~40 s

"Degrades gracefully" is a wish, not an assertion, so turn it into three written clauses before the run. **An accepted-work clause**: the share of offered requests the system must still complete, and the response-time bound those completions stay inside. **A refusal clause**: the rest may be turned away, but each rejection must come back quickly and reach the caller as something it can recognise as a refusal. **A never-events clause**: results that fail whatever the timings say, such as work applied in part, requests lost with nobody told, or a backlog with no ceiling. Fix the numbers before the run so the result cannot be renegotiated afterwards, and write in the offered rate each clause holds at, because behaviour at twice the intended limit is a different promise from behaviour at ten times it.

code

pseudocode · 17 lines
pseudocode
# agreed before the run: what "degrades gracefully" means at 3x the intended rate
overloadRule:
  offeredRate           = 3 * intendedArrivalLimit
  measuredOver          = final 10 minutes of the pressure

  served.share          >= 0.60 of offered requests
  served.latency95th    <= 800 ms

  refused.latency99th   <= 100 ms      # a rejection must be cheap
  refused.signalled     = true         # the caller is told, explicitly

  neverEvents           = [ partiallyAppliedWork,
                            acceptedThenLostWithoutNotice,
                            backlogWithNoDeclaredCeiling ]

  outcome = FAIL if any neverEvent observed
            else PASS if every clause above holds

go deeper

for a junior

Recall what "past the intended limit" means: the system is being offered more work than it was built to take. The run still needs an expectation written down beforehand, not just a picture of what happened on the day.

for a middle

Explain the three clauses out loud: what still gets completed and inside what response-time bound, what may be turned away and how fast that rejection must come back, and what fails the run regardless. The numbers are agreed before the run, never after.

for a senior

An interviewer expects you to defend the numbers. Say where the accepted-work floor and the refusal bound came from, what users and calling systems actually tolerate, and why the same rule cannot be reused unchanged at a different offered rate.

for a principal

Own how much behaviour past the limit the product owes at all. A promise of graceful behaviour at any pressure commits engineering budget forever; decide which pressures are worth defending and say plainly that above them nothing is promised.

Under its intended load a system has one job: serve everything, inside the response-time bound the product promised. Past that load the job changes, and the promise has to change with it. "It degrades gracefully" is the sentence teams reach for to describe the new promise, and it is the single most common reason an overload run produces no finding: it cannot be compared against anything. ## Why the phrase is not an assertion A test case asserts by comparing an observation against a value fixed in advance. "Gracefully" fixes nothing. Two engineers reading it picture different systems. One imagines a service that turns away a third of its traffic instantly and keeps the rest fast; the other imagines a service that slows down for everybody but eventually answers every request. Both are defensible designs, and a phrase rules out neither. So the run produces graphs and the judgement moves to whoever is in the room when the graphs appear. That is a demonstration rather than a run with a result, and nothing is recorded that a later change can be caught regressing against. The repair is mechanical. Before the run, split the offered work into outcome classes and put a bound on each class. ## The three clauses a testable rule needs **1. The accepted-work clause.** Name the share of offered requests the system is still expected to complete, and the response-time bound those completions must stay inside. Inside the intended limit this clause is trivially "all of them"; past the limit it becomes a floor, because turning some work away is now permitted. The floor is what makes the run falsifiable in the useful direction: a change that quietly cuts how much work survives overload breaks it. **2. The refusal clause.** The work that is not completed has to go somewhere, and "refused" is a legitimate destination. Bound it: how quickly a rejection must come back, and the requirement that the caller receives something it can recognise as a refusal rather than silence. The refusal bound is normally far tighter than the served bound, because a rejection that costs as much waiting as a success buys the caller nothing. **3. The never-events clause.** Some results fail the run whatever the timings say, because they are damage rather than slowness: work applied in part and left inconsistent, a request accepted and then lost with nobody told, a queue or in-flight set that grows without any ceiling. These are not traded against response times, and writing that down is what stops them being traded in the room afterwards. | Clause | What it fixes in advance | What a violation looks like | | --- | --- | --- | | Accepted work | a floor on the share completed, and their response-time bound | the floor is missed, or surviving requests exceed the bound | | Refusal | how fast a rejection returns, and that it is explicit | callers wait a full timeout before being told no | | Never-events | outcomes that are not tradeable against timing | half-applied work, silent loss, a backlog with no ceiling | ## The number everyone forgets: the pressure the rule holds at Behaviour past the limit is not one behaviour. A service may hold the accepted-work floor comfortably at twice its intended arrival rate and miss it badly at ten times that rate. A rule that does not say which pressure it applies to can be satisfied by almost any run, and two runs at different pressures get compared as though they were the same test. Write the offered rate into the rule, and accept that you are writing a small family of rules, one per pressure the product is actually worth defending. Deciding that nothing is promised above some pressure is a legitimate and often correct answer; leaving it unsaid is not. ## Writing one in ten minutes 1. State the intended limit you are exceeding, and this run's offered rate as a multiple of it. 2. Fix the accepted-work floor, and the response-time bound for the requests that survive. 3. Fix the refusal bound, and state that a refusal must be explicit to the caller. 4. List the never-events, and name the check that would detect each one. 5. Say which window the figures are read over, so a slow opening minute is not argued about later. 6. Get the owner of the user-facing promise to agree the rule before the run starts, not after. ## What the rule deliberately does not do It does not try to discover the pressure at which the system stops coping. That is a capacity question with its own method and its own answer, and folding it in makes one run answer two questions badly. It also does not prescribe how the system achieves the behaviour. The rule states outcomes; whether they are met by refusing early, by reducing what each response contains, or by queueing behind a ceiling is a design decision, and a rule that names the mechanism will fail the next perfectly good implementation for the wrong reason.

  • Why must the rule name the offered rate it applies at?
    Because behaviour past the limit is not one behaviour. A service can hold its accepted-work floor comfortably at twice its intended arrival rate and miss it badly at ten times that rate, and both can be acceptable. Without the rate written in, a run at any pressure can be argued into a pass, and two runs at different pressures get compared as though they were the same test.
  • How does this rule differ from the pass criteria for a run inside the intended limit?
    Inside the limit the rule is one-sided: everything should succeed inside the bound, and any rejection is a defect. Past the limit rejection becomes an expected outcome with a bound of its own, and "all requests succeed" is replaced by a floor on the share that survives. Reusing the inside-the-limit rule past the limit fails every run for the wrong reason and teaches the team to ignore the result.
  • The team cannot agree a response-time bound for the requests that survive. What do you do?
    Write the rule with the clauses you can agree, that rejections must be fast and explicit and that the never-events must not happen, and record the survivors' bound as an open decision with a placeholder taken from current behaviour. Three assertable clauses still make the run a test. Then take the measured spread to whoever owns the user-facing promise and close the placeholder before the next run.

saying these in an interview costs you the question

  • Treats "degrades gracefully" as self-evident and never writes it down
  • Sets a rule only for successful requests, none for rejections
  • Decides what counted as acceptable after seeing the results
  • Assumes any error past the intended limit is automatically acceptable
  • Writes one rule without saying at what offered rate it holds
open as a page

How do you decide how many hours a sustained load run must hold?

level: middleimportance: must knowfreq 58%

basics

~10 s

Derive it from the build-up you expect to surface: the hold must be long enough for that accumulation to clear measurement noise and cover a readable fraction of the headroom. Calendar convenience decides nothing.

open as a page

After a surge in demand subsides, what must be true before you call the system recovered?

level: middleimportance: must knowfreq 55%

basics

~20 s

Recovery needs more than response times falling. The backlog must drain back to its pre-surge depth, resources must be released rather than settling higher, accepted work must still complete, and that state must hold across a defined observation window.

open as a page

How do you parameterise a sudden surge in arrival rate so a performance run's result is attributable to it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Fix three numbers before the run: the settled baseline arrival rate, the multiple applied to it, and the transition time over which the rate rises. Vary one per run, and hold the peak long enough to read a result.

open as a page

In a load run held steady for hours, why assert on the growth trend rather than the peak?

level: juniorimportance: should knowfreq 47%

basics

~20 s

A peak says how high a figure got; accumulation is a direction. Measure retained memory, open handles, cached entries and disk use as a slope across the hold, then compare that slope with a period expected to be level.

open as a page

How does an overload run assert that refusing work is a correct outcome rather than a defect?

level: middleimportance: should knowfreq 44%

basics

~20 s

By asserting three things about every rejection: it comes back quickly rather than after a full wait, it reaches the caller as an explicit refusal it can act on, and it never counts as a success.

open as a page

Which results fail an overload run outright, whatever its response times show?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Three results fail regardless of timing: work applied in part and left inconsistent, a request accepted and lost with nobody told, and a queue or in-flight set that grows with no ceiling. These are damage, not slowness.

open as a page

Under overload, two systems both show a falling success rate. Which observations tell shedding apart from collapse?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Shedding holds completed work near the sustainable level, answers rejections in milliseconds, keeps survivors near normal speed, and holds backlog and memory flat. Collapse shows completed work falling, everyone waiting, and a backlog that keeps growing.

open as a page

Response times climb across a long steady-load run while per-request resource use stays flat - what do you check?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Check whether the climb advances with work done or with elapsed time. Replot it against cumulative completed work, run a low-rate control, probe a path the run never writes to, and restart the process keeping its data.

open as a page

When extra capacity arrives minutes after a surge, how do you keep a performance run from hiding that delay?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Start at the size the system would reach under baseline demand, never pre-raised to peak, and timestamp two series: offered demand and capacity actually serving. The interval between them is the exposure window, and what happened inside it is the result.

open as a page

How do you prove across a demand surge that no unit of work was lost or done twice?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Tag every submitted unit with a unique identifier before demand rises, then reconcile after the observation window: submitted equals completed plus refused plus pending. Separately compare distinct completed identifiers against total completions, because totals alone hide loss and repetition together.

open as a page

How do you decide whether a long sustained run should hit scheduled and rotational work or avoid it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Decide from what the hold must prove. Arrange the run to hit periodic work - scheduled jobs, credential rotation, cache expiry, index maintenance - when that interaction is the risk; otherwise disable it and record that the trend excludes it.

open as a page

How do you establish that a second, larger request peak after an applied surge was self-inflicted rather than real demand?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Compare what the load generator offered against what the service received: the offered rate is known and unchanged, so any excess was manufactured inside the system. Repeats of identifiers already submitted, arriving synchronised at a fixed delay, confirm it.

open as a page