Past a system's intended load limit, what must "degrades gracefully" become before a performance run can assert it?
answer
- a wish is not an assertion
- split the offered work into outcomes
- served, refused, never-acceptable
- bound the refusal, not only success
- name the rate each clause holds at
basics
~20 s"Degrades gracefully" must become outcome rules written before the run: which requests are still completed and inside what response-time bound, which may be refused, how fast a refusal must arrive, and which results fail outright.
solid answer
~40 s"Degrades gracefully" is a wish, not an assertion, so turn it into three written clauses before the run. **An accepted-work clause**: the share of offered requests the system must still complete, and the response-time bound those completions stay inside. **A refusal clause**: the rest may be turned away, but each rejection must come back quickly and reach the caller as something it can recognise as a refusal. **A never-events clause**: results that fail whatever the timings say, such as work applied in part, requests lost with nobody told, or a backlog with no ceiling. Fix the numbers before the run so the result cannot be renegotiated afterwards, and write in the offered rate each clause holds at, because behaviour at twice the intended limit is a different promise from behaviour at ten times it.
code
pseudocode · 17 lines# agreed before the run: what "degrades gracefully" means at 3x the intended rate
overloadRule:
offeredRate = 3 * intendedArrivalLimit
measuredOver = final 10 minutes of the pressure
served.share >= 0.60 of offered requests
served.latency95th <= 800 ms
refused.latency99th <= 100 ms # a rejection must be cheap
refused.signalled = true # the caller is told, explicitly
neverEvents = [ partiallyAppliedWork,
acceptedThenLostWithoutNotice,
backlogWithNoDeclaredCeiling ]
outcome = FAIL if any neverEvent observed
else PASS if every clause above holdsgo deeper
Recall what "past the intended limit" means: the system is being offered more work than it was built to take. The run still needs an expectation written down beforehand, not just a picture of what happened on the day.
Explain the three clauses out loud: what still gets completed and inside what response-time bound, what may be turned away and how fast that rejection must come back, and what fails the run regardless. The numbers are agreed before the run, never after.
An interviewer expects you to defend the numbers. Say where the accepted-work floor and the refusal bound came from, what users and calling systems actually tolerate, and why the same rule cannot be reused unchanged at a different offered rate.
Own how much behaviour past the limit the product owes at all. A promise of graceful behaviour at any pressure commits engineering budget forever; decide which pressures are worth defending and say plainly that above them nothing is promised.
Under its intended load a system has one job: serve everything, inside the response-time bound the product promised. Past that load the job changes, and the promise has to change with it. "It degrades gracefully" is the sentence teams reach for to describe the new promise, and it is the single most common reason an overload run produces no finding: it cannot be compared against anything. ## Why the phrase is not an assertion A test case asserts by comparing an observation against a value fixed in advance. "Gracefully" fixes nothing. Two engineers reading it picture different systems. One imagines a service that turns away a third of its traffic instantly and keeps the rest fast; the other imagines a service that slows down for everybody but eventually answers every request. Both are defensible designs, and a phrase rules out neither. So the run produces graphs and the judgement moves to whoever is in the room when the graphs appear. That is a demonstration rather than a run with a result, and nothing is recorded that a later change can be caught regressing against. The repair is mechanical. Before the run, split the offered work into outcome classes and put a bound on each class. ## The three clauses a testable rule needs **1. The accepted-work clause.** Name the share of offered requests the system is still expected to complete, and the response-time bound those completions must stay inside. Inside the intended limit this clause is trivially "all of them"; past the limit it becomes a floor, because turning some work away is now permitted. The floor is what makes the run falsifiable in the useful direction: a change that quietly cuts how much work survives overload breaks it. **2. The refusal clause.** The work that is not completed has to go somewhere, and "refused" is a legitimate destination. Bound it: how quickly a rejection must come back, and the requirement that the caller receives something it can recognise as a refusal rather than silence. The refusal bound is normally far tighter than the served bound, because a rejection that costs as much waiting as a success buys the caller nothing. **3. The never-events clause.** Some results fail the run whatever the timings say, because they are damage rather than slowness: work applied in part and left inconsistent, a request accepted and then lost with nobody told, a queue or in-flight set that grows without any ceiling. These are not traded against response times, and writing that down is what stops them being traded in the room afterwards. | Clause | What it fixes in advance | What a violation looks like | | --- | --- | --- | | Accepted work | a floor on the share completed, and their response-time bound | the floor is missed, or surviving requests exceed the bound | | Refusal | how fast a rejection returns, and that it is explicit | callers wait a full timeout before being told no | | Never-events | outcomes that are not tradeable against timing | half-applied work, silent loss, a backlog with no ceiling | ## The number everyone forgets: the pressure the rule holds at Behaviour past the limit is not one behaviour. A service may hold the accepted-work floor comfortably at twice its intended arrival rate and miss it badly at ten times that rate. A rule that does not say which pressure it applies to can be satisfied by almost any run, and two runs at different pressures get compared as though they were the same test. Write the offered rate into the rule, and accept that you are writing a small family of rules, one per pressure the product is actually worth defending. Deciding that nothing is promised above some pressure is a legitimate and often correct answer; leaving it unsaid is not. ## Writing one in ten minutes 1. State the intended limit you are exceeding, and this run's offered rate as a multiple of it. 2. Fix the accepted-work floor, and the response-time bound for the requests that survive. 3. Fix the refusal bound, and state that a refusal must be explicit to the caller. 4. List the never-events, and name the check that would detect each one. 5. Say which window the figures are read over, so a slow opening minute is not argued about later. 6. Get the owner of the user-facing promise to agree the rule before the run starts, not after. ## What the rule deliberately does not do It does not try to discover the pressure at which the system stops coping. That is a capacity question with its own method and its own answer, and folding it in makes one run answer two questions badly. It also does not prescribe how the system achieves the behaviour. The rule states outcomes; whether they are met by refusing early, by reducing what each response contains, or by queueing behind a ceiling is a design decision, and a rule that names the mechanism will fail the next perfectly good implementation for the wrong reason.
- Why must the rule name the offered rate it applies at?Because behaviour past the limit is not one behaviour. A service can hold its accepted-work floor comfortably at twice its intended arrival rate and miss it badly at ten times that rate, and both can be acceptable. Without the rate written in, a run at any pressure can be argued into a pass, and two runs at different pressures get compared as though they were the same test.
- How does this rule differ from the pass criteria for a run inside the intended limit?Inside the limit the rule is one-sided: everything should succeed inside the bound, and any rejection is a defect. Past the limit rejection becomes an expected outcome with a bound of its own, and "all requests succeed" is replaced by a floor on the share that survives. Reusing the inside-the-limit rule past the limit fails every run for the wrong reason and teaches the team to ignore the result.
- The team cannot agree a response-time bound for the requests that survive. What do you do?Write the rule with the clauses you can agree, that rejections must be fast and explicit and that the never-events must not happen, and record the survivors' bound as an open decision with a placeholder taken from current behaviour. Three assertable clauses still make the run a test. Then take the measured spread to whoever owns the user-facing promise and close the placeholder before the next run.
saying these in an interview costs you the question
- Treats "degrades gracefully" as self-evident and never writes it down
- Sets a rule only for successful requests, none for rejections
- Decides what counted as acceptable after seeing the results
- Assumes any error past the intended limit is automatically acceptable
- Writes one rule without saying at what offered rate it holds