Which results fail an overload run outright, whatever its response times show?
answer
- some results are not tradeable at all
- damage, not slowness
- reconcile what was sent against what is held
- half-applied, silently lost, unbounded
- ask for the ceiling, not the peak
basics
~20 sThree results fail regardless of timing: work applied in part and left inconsistent, a request accepted and lost with nobody told, and a queue or in-flight set that grows with no ceiling. These are damage, not slowness.
solid answer
~50 sResponse-time figures describe how well a system coped; these three describe damage, and damage is not traded against speed. **Partially applied work**: an operation that recorded some of its effects and abandoned the rest, leaving records that contradict each other. **Silent loss**: a request accepted, never completed and never reported as failed, so the caller believes it succeeded and cannot retry, surface or even count it. **Unbounded growth**: a queue, connection set or in-flight collection with no ceiling, so pressure is stored rather than turned away and the failure moves somewhere far less predictable. Each needs an assertion no timing figure can provide: an invariant over the whole dataset checked once at the end, a reconciliation of what the driver was told against what the system actually holds, and evidence that a declared ceiling exists rather than that the observed peak happened to be tolerable.
code
pseudocode · 16 lines# end-of-run checks, evaluated independently of any response-time rule
issued = driver.requestsOffered
ackedSuccess = driver.responsesClassifiedServed
ackedRefused = driver.responsesClassifiedRefused
unanswered = driver.responsesNeverReceived
completed = system.operationsCompleted
inconsistent = system.operationsWithSomeEffectsMissing
assert inconsistent == 0 # half-applied work
assert completed >= ackedSuccess # acknowledged but absent
assert issued == ackedSuccess + ackedRefused + unanswered
for each queue in system.queues:
assert queue.declaredCeiling exists
assert queue.rejectedOnFull > 0 when pressure exceeded the ceilinggo deeper
Recall that some outcomes are not a matter of degree. Losing a request without telling anyone, or applying half of an operation, is wrong at any pressure and cannot be excused by how much load the system was under.
Explain how each one is detected: an invariant evaluated over the whole dataset afterwards, a reconciliation of acknowledged successes against completed operations, and a check that a queue has a declared ceiling rather than a merely tolerable observed peak.
An interviewer expects you to design the check, not name the rule. Pick an invariant the data must satisfy after the pressure stops, say who computes it and against what, and explain why a passing response-time figure never substitutes for it.
Own which outcomes the organisation refuses to trade. A short published list of results that fail any overload run, independent of timing targets, is what stops each team renegotiating damage under schedule pressure.
Most of what an overload run produces is a matter of degree. Response times are worse than normal, and that may be acceptable. Some share of work is refused, and that may be acceptable too. A small number of results are not matters of degree at all, and a rule that does not name them in advance will end up trading them away against a timing target in the meeting where the results are discussed. ## What makes a result untradeable The distinction is damage versus slowness. A slow response is a worse version of a correct one: the user waited longer and got what they asked for. A half-applied operation is not a worse version of anything. It leaves the system holding records that contradict each other, and no later run and no faster machine repairs it. Silence about a lost request is the same kind of thing. It does not degrade the caller's experience, it removes the caller's ability to respond to it at all. Because these are qualitatively different, they get their own clause in the pass rule, evaluated independently of every timing figure, and stated as a failure whatever else the run reports. ## The three **Partially applied work.** An operation with several effects, a record written, a balance adjusted, a notification emitted, recorded some of them and abandoned the rest. Under pressure this happens for ordinary reasons: an internal wait expires midway, a resource is exhausted between two effects, a decision to turn work away is taken after the first effect and before the second. The result is a dataset that fails its own invariants. **Silent loss.** A request was accepted, never completed, and never reported as failed, so the caller believes it succeeded. This is worse than an outright refusal in every direction: the caller cannot retry it, cannot surface it, cannot even count it, and the loss reappears days later as a customer complaint with no trail to follow. **Unbounded growth.** A queue, a connection set, a retry buffer or an in-flight collection with no ceiling. Under overload it stores the excess instead of turning it away, which converts a load problem into a memory problem and moves the failure somewhere far less predictable. The absence of a ceiling is the never-event; the particular depth reached during one run is not the point. | Never-event | How it shows up | The check that finds it | | --- | --- | --- | | Partially applied work | records that contradict each other afterwards | an invariant over the whole dataset, evaluated once at the end | | Silent loss | acknowledged successes exceed completed operations | reconcile what the driver was told against what the system holds | | Unbounded growth | depth and memory climb for the whole period | a declared ceiling, plus evidence that entries are turned away at it | ## Checking them without inspecting every operation Hundreds of thousands of operations cannot be examined one at a time, and they do not need to be. Each never-event has a whole-dataset check: 1. **Choose operations that carry an invariant.** A total that must balance, a child row that must have a parent, a count that must equal the number of accepted requests. Evaluate the invariant once when the pressure stops; a break names the handful of operations worth opening. 2. **Reconcile the two sides.** The driver knows how many requests it issued and how many it was told succeeded. The system knows how many operations it completed. Acknowledged successes above completed operations is silent loss, and the gap is the count. 3. **Ask for the ceiling, not the peak.** Read the configured bound of every queue involved and confirm the system turns work away at it. A run that never reached the ceiling has not shown that one exists. 4. **Account for every offered request.** Issued should equal served plus refused plus unanswered, with nothing left over. An unexplained remainder is itself a finding. ## Writing them into the rule They belong in the pass rule agreed before the run, phrased as absolutes rather than budgets. Not "fewer than one operation in a thousand may be left inconsistent", but "no operation may be left inconsistent". A budget invites a negotiation about how much damage is acceptable, and under schedule pressure that negotiation has a predictable outcome. There is one honest exception, and it should be written rather than assumed. An operation the product genuinely defines as best-effort, a view counter or an optional enrichment, may be dropped, provided the drop is by design and the caller was never told otherwise. Name those explicitly. Everything not named is covered by the absolute. The reason to be strict here is that these three are the results overload uniquely produces. Ordinary functional testing exercises the same operations with no pressure at all and finds none of them, because they only appear when the system has to give something up. If the overload run does not check for them, nothing does.
- How do you detect partially applied work when the run offered hundreds of thousands of operations?Not one at a time. Choose operations that carry an invariant the whole dataset must satisfy: a total that must balance, a child row that must have a parent, a count that must equal the number of accepted requests. Evaluate the invariant once when the pressure stops. A break names the small set of operations worth opening, and the size of the break sizes the problem.
- A team argues that dropping work under extreme pressure is unavoidable, so silent loss should not fail the run. What is the answer?Dropping is not the problem; silence is. Turning a request away and telling the caller is a correct outcome the caller can retry or surface. Accepting it, discarding it and reporting success removes every option the caller had, and the loss resurfaces later as a customer complaint with no trail. The rule fails the silence, not the drop.
saying these in an interview costs you the question
- Trades data damage against better response times
- Calls a lost request acceptable because the pressure was extreme
- Checks the peak backlog depth but never whether a ceiling exists
- Assumes a returned success means the work was actually applied
- Reads only the caller's view and never reconciles system state