skip to content

Under overload, two systems both show a falling success rate. Which observations tell shedding apart from collapse?

level: seniorimportance: should knowfreq 46%

answer

  1. one figure hides two different behaviours
  2. watch finished work, not failures
  3. who waits: everyone, or only the rejected
  4. a bounded backlog versus a growing one
  5. memory and in-flight work flat or climbing

basics

~20 s

Shedding holds completed work near the sustainable level, answers rejections in milliseconds, keeps survivors near normal speed, and holds backlog and memory flat. Collapse shows completed work falling, everyone waiting, and a backlog that keeps growing.

solid answer

~50 s

A single success-rate figure cannot tell them apart; four other readings can. **Completed work over time**: shedding keeps the rate of finished requests near the level the system can sustain and turns the excess away, while collapse drives that rate down, sometimes toward zero. **Where the waiting lands**: a shedding system answers rejections in a small fraction of the served response time and keeps the work it accepted near normal, while a collapsing one makes everybody wait, including the requests it will fail anyway. **The backlog**: bounded and flat, versus growing for as long as the pressure lasts. **Resource use**: flat, because work is turned away before it is admitted, versus climbing as accepted work piles up in flight. Two of the four are visible from the caller's side alone; the other two need the system's own instrumentation.

code

pseudocode · 15 lines
pseudocode
# reading one overload run: is the system shedding, or collapsing?
for each interval in the pressure period:
  completedPerSecond   # flat near sustainable  -> shedding
                       # falling                -> collapsing
  refusedLatency99th   # small and stable       -> shedding
                       # as long as everything  -> collapsing
  servedLatency95th    # close to normal        -> shedding
                       # climbing               -> collapsing
  backlogDepth         # rises to a ceiling     -> shedding
                       # grows all period       -> collapsing
  inFlightWork         # flat                   -> shedding
                       # climbing               -> collapsing

verdictOfRun = COLLAPSE when backlogDepth and inFlightWork grow
               while completedPerSecond falls

go deeper

for a junior

Recall the difference in one sentence: a system that sheds turns work away on purpose and keeps completing the rest, while a system that collapses accepts everything and finishes less and less of it.

for a middle

Explain which series you would plot: completed work per second, response time split between the requests that were served and the ones that were turned away, and backlog depth. Say what each looks like in each of the two behaviours.

for a senior

An interviewer expects a diagnosis from partial evidence. Show which readings you can get from the caller's side, which need the system's own instrumentation, and which single reading you would trust least on its own and why.

for a principal

Own the standard: decide which of these readings a service must expose before anyone may sign it off as safe past its intended load limit, so the distinction rests on evidence rather than on argument after an incident.

A falling success rate is the symptom both behaviours share, and it is why the two get confused. One system is refusing work it cannot take, on purpose, and completing everything it did take. The other has accepted more than it can finish and is finishing less and less of it. From a single percentage they look identical. From four other readings they look nothing alike. ## Why the success rate cannot decide it The success rate compresses two independent quantities, how much work arrived and how much finished, into one number, and then discards when each thing happened. A system refusing forty percent of its offered work and a system failing forty percent of it after making everybody wait both report sixty percent. Worse, the refusing system's number gets *worse* the more aggressively it protects itself, which is exactly backwards as a health signal. ## The four readings **Completed work over time.** Plot the rate of requests actually finished, per interval, across the pressure. A shedding system holds this near the level it can sustain: excess arrivals are turned away and the finished-work line stays flat. A collapsing system's line falls, often steeply, because effort is going into work that will not complete: waits that expire, internal calls retried, contention on whatever resource ran out. Flat-and-adequate versus falling is the single most informative reading. **Where the waiting lands.** Split response time by outcome. A shedding system answers refusals in a small fraction of the served response time, and the requests it did accept stay near their normal figure, because they compete with a bounded amount of other work. A collapsing system makes everybody wait: the requests it will eventually fail wait as long as the ones that succeed, so the caller pays full latency for nothing. **The backlog.** Look at the depth of whatever holds work waiting to be processed, and ask whether it has a ceiling. A shedding system's depth rises to its ceiling and stops, because arrivals beyond it are turned away at the door. A collapsing system's depth grows for as long as the pressure lasts, and that growth is the mechanism behind the other readings: everything admitted is stored, so waiting grows, so completion falls. **Resource use.** Memory, open connections and in-flight work stay roughly flat in a shedding system, because work is refused before it is admitted and therefore before it costs anything. In a collapsing system they climb alongside the backlog. | Reading | Shedding | Collapsing | | --- | --- | --- | | Completed work per interval | flat, near the sustainable level | falling, sometimes toward zero | | Response time of refusals | small and stable | as long as everything else, or no answer at all | | Response time of served requests | close to normal | climbing throughout | | Backlog depth | rises to a ceiling and stops | grows for the whole period | | Memory and in-flight work | flat | climbing with the backlog | ## Reading it from outside the system Two of the four are available to the caller alone, which matters when the system under test is barely instrumented. The rate of successful responses per second, held against the rate being offered, gives you completed work. Splitting the caller's own recorded response times into successes and failures gives you where the waiting lands. If failures come back fast while successes stay normal, the system is refusing deliberately; if failures arrive only after a long wait, the system is queueing everything and failing it late. Backlog depth and resource growth need the system's own instrumentation, which is why an overload run is worth very little against a system that exposes neither. ## What each reading can mislead you about - A flat completed-work line proves nothing if it is flat at a level far below what the system should sustain. Flatness is good news only beside the number it settled at. - A backlog that stayed bounded over a short period may simply not have had time to grow. Bounded means a ceiling exists and entries are turned away at it, not that the observed peak happened to be tolerable. - A summary that pools refusals and successes into one response-time figure looks *better* the more work the system turns away, because refusals are fast. Split by outcome before reading any response-time figure from an overload period. - Resource use alone is ambiguous. A system can hold memory flat while failing almost everything, if its failures are cheap. The distinction is not academic. Shedding that still meets the agreed accepted-work floor is a pass. Collapse is a finding with a specific shape, that something admitted work it could not finish, and it points at a missing ceiling rather than at slow code.

  • The caller can see only its own responses. What can it still tell from that side alone?
    Quite a lot. If failures come back in milliseconds while successes stay near their normal response time, the system is turning work away deliberately. If failures arrive only after a long wait, and successes are slow too, the system is queueing everything and failing it late, so the caller pays full latency for nothing. Successful responses per second, held against the rate being offered, completes the picture.
  • Why can a system that is clearly shedding still fail the run?
    Because shedding is only correct if what it turns away is turned away cleanly. The run still fails if refused work was half applied before being dropped, if a rejection never reached the caller, or if the share still completed fell below the floor agreed beforehand. Deliberate refusal is a mechanism, never an automatic pass.

A doorman who turns people away at the entrance keeps the room usable for everyone inside. A doorman who lets everybody in and then cannot move keeps nobody comfortable, and the queue outside never stops growing.

saying these in an interview costs you the question

  • Reads only the success rate and calls both behaviours the same
  • Treats every failed request under overload as evidence of collapse
  • Assumes fast failures mean breakage, not deliberate protection
  • Ignores backlog depth because response times still look fine
  • Judges by resource use alone without looking at completed work