skip to content

questions

4

In a performance run, what does a gap between offered and achieved throughput tell you?

level: middleimportance: must knowfreq 62%

answer

  1. Two throughput numbers, not one
  2. Demand applied versus work completed
  3. Where did the missing requests go
  4. Reconcile issued against every outcome

basics

~20 s

Offered throughput is the demand the run applied; achieved throughput is the work that actually completed and was counted. A gap means requests were refused, expired or never finished, so the completed-work figure describes only part of the intended load.

solid answer

~50 s

**Offered throughput** is the request rate the run intended to apply — the arrival rate written into the workload profile. **Achieved throughput** is the rate of requests that actually completed and were counted. When the two match, the completed-work figure describes the load you meant to apply. When achieved falls below offered, some demand was refused, expired, or was still outstanding when the run stopped. A chart plotting only completed work per second is blind to this: as the system serves less, that line can stay flat or even improve, because a request that never finished has no place on it. So the first thing I read after a run is the demand accounting — issued, completed, failed, outstanding — and I only trust the response-time figures once those four reconcile against the issued total.

code

pseudocode · 16 lines
pseudocode
run_report:
  issued        = 1_080_000   # requests the profile called for
  completed_ok  =   712_000   # replies received and accepted
  completed_bad =    32_000   # replies received and rejected
  expired       =    41_000   # deadline passed, no reply
  outstanding   =   295_000   # unanswered or unsent at stop

# the accounting must close before any other figure is read
assert issued == completed_ok + completed_bad + expired + outstanding

offered_rate  = issued       / window_seconds
achieved_rate = completed_ok / window_seconds
error_rate    = (completed_bad + expired) / issued   # not / replies_received

if achieved_rate < 0.95 * offered_rate:
    report("demand shortfall", offered_rate - achieved_rate)

go deeper

for a junior

Be ready to say that a run has two throughput numbers: the demand it applied and the work that finished. Knowing that the second can be lower than the first, and that only the second appears on a completion chart, is enough at this level.

for a middle

Explain the mechanics: why a completion rate has no term for unfinished demand, why fast rejections can hold the line up, and how the issued count reconciles against accepted, rejected, expired and outstanding requests.

for a senior

Show that you read the accounting before the timings, and that you refuse a response-time summary drawn from a run whose demand did not reconcile. An interviewer expects you to catch a flattering report rather than to recite definitions.

for a principal

Own the reporting standard: which figures a run is required to publish, that the error rate is expressed over issued demand, and that no throughput number leaves the team without the offered figure beside it.

## Two numbers that are easy to confuse **Offered throughput** — also called offered load, or the applied arrival rate — is the demand a performance run *intends* to place on the system: the number of requests per second the workload profile calls for at that moment. It is a property of the plan, and you chose it. **Achieved throughput** is the number of requests per second that actually completed and were counted. It is a property of the result. On an idle system the two numbers coincide, which is exactly why they get used interchangeably in conversation. They separate at the point the run becomes interesting. | | Offered throughput | Achieved throughput | |---|---|---| | What it is | Demand the profile applied | Work that completed and was counted | | Comes from | The run plan | The result set | | Why it falls | It does not — you set it | Refusals, expiries, unfinished work | | Under pressure | Stays where you put it | Can fall, flatten, or even rise on fast failures | ## Where the missing demand goes A gap means some requests produced no counted completion. Each one sits in a small number of buckets, and a report worth reading names them separately rather than folding them into a single failure total: - **Refused before service.** The system answered at once with a rejection. These are cheap and fast, so counting them as completed units of work can make the achieved rate *rise* while the system serves less real demand. - **Expired.** The run stopped waiting after its own deadline and no reply ever arrived. - **Failed in transport.** The connection was refused outright, or closed part-way through a reply. - **Still outstanding at the stop.** Issued, unanswered, and cut off when the run ended. - **Never issued at all.** The profile called for them and nothing went out. They are unmet demand and belong in the report as such, not quietly omitted from it. ## Why a completed-work line flatters a struggling system Suppose a profile calls for **900 requests per second for twenty minutes** — 1,080,000 requests in total. The summary reports **620 completed per second** and **0.4 percent errors**. Read on its own, that is a system that looks like it is coping. Reconcile it and the picture inverts. About 744,000 requests produced a counted completion, so roughly three requests in ten produced none at all. The 0.4 percent is 0.4 percent *of the replies that arrived*, not of the demand that was applied. A completed-work chart has no term for a request that never finished, so as demand goes unserved the line does not dip — the missing requests are simply absent from it. A flat line is equally compatible with a system serving all of the load and with one serving two thirds of it. That ambiguity is structural, and it always resolves in the flattering direction. ## Reading the accounting The check is arithmetic and takes about a minute: 1. Take the **issued count** the profile called for over the measured window. 2. Add up **accepted completions**, **rejected completions**, **expiries**, **transport failures** and **requests outstanding at the stop**. 3. Confirm those categories sum to the issued count. If they do not, the run is not reporting where some of its demand went, and no other figure from it can be trusted. 4. Compute offered and achieved as rates over the same window and plot both series together, so the gap is visible rather than inferred. 5. Express the error rate over **issued requests**, not over replies received. ## What a gap does not tell you The accounting is a fact, not a diagnosis, and three temptations follow it immediately: - It does not say **which part of the path** stopped serving. That is a separate investigation with its own evidence. - It does not give you a **sustainable capacity number**. A shortfall seen once at one demand level is an observation about that run, not a limit you can publish. - It does not make the **response-time summary** safe to read. Those samples describe the requests that finished, which during a shortfall are a self-selected minority. The habit worth building is small and cheap: never quote an achieved throughput figure without the offered figure beside it. A lone number labelled *throughput* is ambiguous by construction, and a reader will always assume it means the load the system was given.

  • A run's completed-work rate holds steady while its share of rejections climbs. What is happening?
    Rejections are being counted as completed units of work. They cost the system almost nothing to produce, so as it starts refusing demand the completion rate is propped up by cheap failures. Split the achieved figure into accepted and rejected completions and the accepted line will show the real fall.
  • Why does a run that stops after a fixed number of requests rather than a fixed duration hide the same shortfall?
    With a fixed count the run cannot fall short of its demand — it just takes longer to get through it. The shortfall reappears as elapsed time instead of as missing requests, so the report has to state the wall-clock time the count took and the rate that implies, or the gap is invisible.
  • How would you present offered and achieved throughput so the gap cannot be overlooked?
    Plot both as series on the same axis over the same window, with a third series for requests outstanding. Put the reconciliation — issued against the sum of outcomes — in the summary next to the headline figure, and label the headline as accepted completions per second rather than as throughput.

A ticket office that counts only the customers it served looks busy and efficient; the people who gave up in the queue and walked away never appear anywhere in its numbers.

saying these in an interview costs you the question

  • Treats completed requests per second as the load the system was given
  • Quotes achieved throughput without stating the demand that was applied
  • Reads a flat completion line as proof the system held up
  • Counts requests still outstanding at the stop as successes
  • Computes the error rate over replies received instead of requests issued
open as a page

In a performance run, why is a reply that arrives with a success status not necessarily a success?

level: juniorimportance: should knowfreq 55%

basics

~20 s

A success status only says something answered the request. The body may be a rendered error page, a truncated document, or an empty result where a data record was expected, so a run reading only the status overstates how much work really succeeded.

open as a page

In a performance run, how do you count requests that expired, were reset, or were never sent?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Count them as outcomes, not as missing data. Each stays inside the issued total, lands in a named failure category, and the report states what time was recorded for it. Deleting them describes only the requests the system managed to serve.

open as a page

In a performance run, how much of each reply can you afford to verify in flight?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only checks cheap enough to fit inside the applying side's own budget: a length floor, a required field, a marker. Deep parsing per reply steals the capacity that generates demand, so sample it instead and report what share was verified.

open as a page