How does an overload run assert that refusing work is a correct outcome rather than a defect?
answer
- no is an answer, if it is fast
- three properties, not one
- bound the rejection's own response time
- the caller must be able to tell
- keep refusals out of the success tally
basics
~20 sBy asserting three things about every rejection: it comes back quickly rather than after a full wait, it reaches the caller as an explicit refusal it can act on, and it never counts as a success.
solid answer
~50 sRefusal is correct only when it is cheap, visible and honestly counted, so the run asserts all three separately. **Cheap**: rejections get their own response-time bound, far tighter than the bound on served requests, because a caller that waits its full timeout before being told no has paid the whole cost of the work and received nothing. **Visible**: the caller receives something it can tell apart from its own bad request and from an internal fault, ideally indicating whether retrying later is sensible. **Honestly counted**: refused requests sit in their own outcome class and never migrate into the success tally, because otherwise a service that turns away most of its traffic reports a healthy result. A request that simply never gets an answer is not a refusal at all; it belongs in a third class and is a finding on its own.
code
pseudocode · 14 lines# outcome classes a load driver records during an overload period
for each offered request r:
if r.answered and r.reportedComplete: class(r) = SERVED
elif r.answered and r.explicitRefusal: class(r) = REFUSED
elif not r.answered: class(r) = UNANSWERED
else: class(r) = UNCLASSIFIED
assert count(UNCLASSIFIED) == 0 # every response must be readable
assert count(UNANSWERED) == 0 # silence is not a refusal
assert latency99th(REFUSED) <= refusalBound
assert latency95th(SERVED) <= servedBound
report servedShare = count(SERVED) / count(offered)
report refusedShare = count(REFUSED) / count(offered) # never merged ingo deeper
Recall that turning work away is not automatically a bug. Past the intended load limit a quick, explicit no can be the right answer, and the run's job is to check that the no itself behaves properly.
Explain the three properties a run asserts about a rejection: its own tighter response-time bound, an explicit signal the caller can distinguish from other outcomes, and a separate class in the results that never merges into successes.
An interviewer expects the failure modes: a rejection delivered only after a full timeout, a rejection indistinguishable from an internal fault so callers retry blindly, and a summary that hides how much offered work was turned away.
Own what the product promises when it says no. How fast a rejection comes back, how it is expressed and whether a retry is advised is a cross-team interface decision, and leaving it to each service produces callers that cannot behave sensibly under pressure.
Inside its intended load, a system that refuses work is broken. Past that load, refusing work is often the only correct thing it can do, and a run that treats every rejection as a defect reports a failure for the behaviour you wanted. The question is therefore not whether the system refused, but whether it refused *properly*, and that is three separate assertions. ## Refusal is an outcome class, not an error The results of an overload run should carry at least three outcome classes for the offered work, and they behave differently. **Served** requests are bound by the ordinary response-time rule. **Refused** requests are bound by a much tighter one, and by a visibility requirement. **Unanswered** requests are not an outcome at all; they are a finding. Collapsing the middle class into either of the others is where the mistakes live. Fold refusals into successes and a system that turns away most of its traffic reports a healthy result. Fold them into failures and correct protective behaviour becomes indistinguishable from a fault, so a well-behaved system can never pass. ## The three assertions **It must be cheap.** A rejection returned only after the caller has waited its full timeout has cost the caller everything a success would have cost: a held connection, an occupied worker of its own, and its own callers waiting behind it. Worse, a very late rejection is indistinguishable from an unresponsive system, so upstream components fire retries and add pressure to exactly the system that was trying to shed it. Give refusals their own response-time bound, far tighter than the served bound, and assert it separately. **It must be visible.** The caller has to be able to tell "we did not take this work" apart from "your request was wrong" and from "we tried and it broke", because the three imply different behaviour: retry later, fix and resend, escalate. If the run cannot distinguish them from the response alone, no real client can either. Where the system can also indicate when retrying is sensible, asserting that it does is worth more than any other detail of the response. **It must be counted honestly.** Refusals stay in their own class in the results and never migrate into the success tally. The share of offered work that was turned away is one of the two headline figures of an overload run, the other being what the survivors' response times looked like, and a summary that omits it is not reporting on overload at all. | Property | Refusal done well | Refusal done badly | | --- | --- | --- | | Cost to the caller | answered in a small fraction of the served time | delivered only once the caller gives up waiting | | Signal | explicit, and distinguishable from other outcomes | indistinguishable from an internal fault or from silence | | In the results | its own class, reported as a share of offered work | folded into successes, or buried in one error figure | | Effect upstream | callers back off, or turn work away in turn | callers retry blindly and multiply the pressure | ## What the driver has to record The assertion only exists if whatever applies the load classifies every response instead of counting two buckets: - The outcome class of each response, decided from the response itself rather than inferred from how long it took. - The response time of each class separately, so refusals and successes are never averaged together. A pooled figure improves as the system does less work, which is the most misleading number an overload run can produce. - Requests that were issued and never answered, kept as a class of their own. Silence is not refusal, and the difference is exactly whether the caller was given the chance to act. - The offered rate, so the refused share means something. "Forty percent refused" is a different result at twice the intended limit than at ten times it. ## The contract behind the assertion Behind all three properties is a promise the service makes to its callers, and it is worth writing down independently of any single run: how fast a refusal comes back, how it is expressed, whether a retry is advised, and how long a caller should wait before treating silence as failure. Teams that leave this to each service end up with callers that cannot behave sensibly under pressure, because every dependency says no differently, and the aggregate effect of callers guessing is a retry storm stacked on top of an overload. That is also why this belongs to the engineer who owns the service rather than to whoever runs the load. The run can only assert what the interface makes visible. If a refusal cannot be expressed in the response, no amount of measurement makes it assertable, and the fix is in the service, before the next run.
- Why is a rejection that arrives only after the caller's full wait treated as a failure of the rule?Because the caller has already spent everything the rejection was supposed to save: a held connection, an occupied worker of its own, and its own callers waiting behind it. A very late no is also indistinguishable from an unresponsive system, so upstream components retry and add pressure to the service that was trying to shed it. Refusal earns its place only when it comes back before the work is admitted.
- What should the run assert about how a rejection is expressed to the caller?That it is distinguishable from the other outcomes and carries enough for the caller to act. The caller must be able to tell "we did not take this work" apart from "your request was wrong" and from "we tried and it broke", because the three lead to different behaviour: retry later, fix and resend, escalate. If the run cannot separate them from the response alone, no real client can either.
A venue that tells you at the door it is full lets you go elsewhere. One that leaves you queueing for an hour before saying the same thing has taken your evening and given you nothing.
saying these in an interview costs you the question
- Counts refused requests as successes because nothing broke
- Applies the same response-time bound to rejections and served requests
- Accepts a request that was never answered as a refusal
- Reports one overall error figure with no separate refused share
- Treats any rejection under load as an automatic defect