skip to content

In a performance run, why is a reply that arrives with a success status not necessarily a success?

level: juniorimportance: should knowfreq 55%

answer

  1. A status is not a body
  2. What arrived may be an apology page
  3. Truncated, empty, or degraded replies
  4. Assert one cheap invariant per reply

basics

~20 s

A success status only says something answered the request. The body may be a rendered error page, a truncated document, or an empty result where a data record was expected, so a run reading only the status overstates how much work really succeeded.

solid answer

~50 s

A success status reports delivery, not correctness. Under load the common false successes are a friendly failure page returned by an intermediary, a body truncated when the connection closed early, an empty result where the seeded input guarantees rows, and a cached or default answer served because the live path was too slow. All four are well formed and all four are fast, so they inflate the completion count and pull the reported response times down. Load makes them *more* likely, because fallbacks and degraded paths exist precisely for the moment the system is under pressure. The fix is a constant-cost check on every reply — one required marker, a length floor, or a minimum row count for an input you control — whose failures land in the error tally rather than in a log nobody reads.

code

pseudocode · 13 lines
pseudocode
# constant-cost check applied to every reply in the run
function accept(reply):
    if not reply.status_is_success:
        return outcome("rejected")
    if reply.byte_count < MIN_REAL_ANSWER_BYTES:
        return outcome("truncated_or_empty")
    if not reply.has_field("resultId"):        # only the live path emits it
        return outcome("fallback_or_error_page")
    if reply.field("servedFrom") == "cache":
        return outcome("stale_answer")
    return outcome("accepted")

# every non-accepted outcome is counted, none is discarded

go deeper

for a junior

Be ready to name the difference between a reply arriving and a reply being right, and to give two examples: an error page returned under a success status, and an empty result where data was expected.

for a middle

Explain why load raises the share of false successes — fallbacks, caches and partial results all fire under pressure — and what each one does to the completion count and the reported times.

for a senior

Demonstrate the judgement of choosing one constant-cost invariant per reply rather than full assertions, and of treating an unexpectedly good result as a signal to check the acceptance rule first.

for a principal

Own the rule that no run's figures are published without its acceptance criteria beside them, and decide how much verification the organisation requires before a performance result may inform a release.

## What a success status actually promises A reply that arrives carrying a success status tells you one thing: something in the path accepted the request and produced an answer. It does not tell you that the answer came from the code you were trying to exercise, that it contains the data the request asked for, or that it is complete. Status and content are separate claims, and a performance run that reads only the first is counting deliveries rather than results. This matters more here than in a functional test, because a performance run counts things. Every reply it accepts becomes a completed unit of work in the throughput figure and a sample in the response-time figures. A reply that is wrong but well formed inflates both. ## Four shapes of false success | Shape | What actually arrives | Cheapest check that catches it | |---|---|---| | Rendered failure page | A human-readable apology page produced by an intermediary or a shared error handler | A field or marker only the real path emits | | Truncated body | A partial answer cut short when the connection closed or a size ceiling was hit | A minimum length, plus a check that the terminating structure is present | | Empty or default payload | A well-formed answer containing zero rows where the seeded input guarantees some | A minimum result count for a request whose input you control | | Stale or degraded fallback | A cached or default answer served because the live path was too slow | A freshness marker or an identifier that changes per request | The fourth is the one that catches experienced engineers out. It is not a bug — it is the graceful degradation the team deliberately built — but during a run it converts an overloaded dependency into a fast, cheerful, entirely fictional success. ## Why load makes these more common, not less The pressure of the run is exactly the condition that triggers them: - **Fallbacks are designed to fire under stress.** The busier the system, the larger the share of answers that come from the cheap path instead of the real one. - **Caches serve stale copies for longer** when the refresh path is queueing behind everything else. - **Intermediaries substitute their own documents** when an upstream is slow, and they do it with a success status if they are configured to present a friendly page. - **Partial results become a deliberate strategy** — return what is ready, drop the rest — which is sound engineering and indistinguishable from a full answer if you only measure arrival. So the share of form-correct, content-wrong replies is at its highest precisely in the part of the run you most want to reason about. ## What it does to the numbers Three figures move, all in the reassuring direction: 1. **Success rate is overstated**, because false successes sit in the numerator. 2. **Throughput is overstated**, because a rejection produced in two milliseconds counts as one completed unit of work exactly like a real answer that took two hundred. 3. **Response times are understated**, because cheap wrong answers are fast and they pull the reported figures down. A run in this state does not look broken. It looks unusually good, which is why the finding is so often celebrated instead of investigated. ## The cheapest check worth making You do not need full functional assertions inside a performance run. You need one constant-cost check per reply that a false success cannot pass. In order of value: 1. **A required marker.** One field, header value or token that only the genuine path produces. A shared error page and a cached default both fail it. 2. **A length floor.** A number of bytes below which no real answer can be valid. Truncation and empty payloads both fall through it. 3. **A minimum count.** When the request is built from data you seeded, you know how many rows a correct answer contains; assert the floor, not the exact set. Whatever the check, its failures must land in the error tally rather than being logged and forgotten, and the report must say which check was applied. A run that asserts nothing beyond arrival is still useful — it measures whether the system accepted and answered demand at a rate — but it cannot claim that the work was actually done, and it should not be quoted as if it did.

  • Your run starts reporting faster response times than the previous one and you have changed nothing. What do you check first?
    Whether the mix of replies changed. Cheap wrong answers — cached fallbacks, rejections behind a success status, empty results — are much faster than real work, so a shift towards them lowers the reported times while the system does less. Compare the counts per acceptance outcome between the two runs before believing the improvement.
  • Why is asserting the exact expected body for every reply the wrong answer here?
    It is functional testing bolted onto a measuring tool. Comparing full documents costs far more per reply than the checks above, it consumes the capacity the run needs to apply demand, and it produces failures on every harmless variation. One invariant a false success cannot satisfy gets almost all of the value for a fraction of the cost.

An envelope can arrive on time, correctly addressed and properly stamped, and still be empty; signing for the delivery is not the same as reading the letter.

saying these in an interview costs you the question

  • Treats a success status as proof the work was done
  • Counts a rendered failure page as a completed unit of work
  • Assumes degraded and cached answers stay rare under heavy load
  • Logs content-check failures instead of counting them as errors
  • Quotes response times without saying what was verified