In a performance run, how much of each reply can you afford to verify in flight?
answer
- Verification spends the capacity that applies demand
- Constant-cost check on every reply
- Deep check on a small sample
- Stop the clock, then verify
- State what the run did not verify
basics
~20 sOnly checks cheap enough to fit inside the applying side's own budget: a length floor, a required field, a marker. Deep parsing per reply steals the capacity that generates demand, so sample it instead and report what share was verified.
solid answer
~40 sVerification is paid for out of the same capacity that issues requests, so heavy per-reply checking lowers the demand the run can apply and adds the measuring side's own work to the recorded times. The shape that works is two-tier: a **constant-cost check on every reply** — a required marker, a length floor, a row-count floor — plus a **deep check on a small sampled share**, one in a hundred or so, reported separately with its sampling rate. Where possible, stop the clock when the last byte arrives and verify afterwards, so the check never lands inside the measurement. And be precise about the claims: a run that verified nothing may say the system accepted demand at a rate and answered within times; it may not say the answers were right.
code
pseudocode · 19 linesSAMPLE_ONE_IN = 100
reply = await_reply(request, DEADLINE)
elapsed = now() - request.sent_at # clock stops here
# tier 1: constant cost, every reply
outcome = "accepted"
if reply.byte_count < MIN_BYTES: outcome = "too_short"
else if not reply.has_field("resultId"): outcome = "wrong_path"
# tier 2: deep check, sampled, cost paid off the timed path
if outcome == "accepted" and request.index % SAMPLE_ONE_IN == 0:
deep_checked += 1
if not matches_expected(reply, request.input):
outcome = "content_mismatch"
record(outcome, elapsed)
report.verification = { every_reply: "marker+length",
deep_sample: deep_checked / accepted }go deeper
Be ready to say that checking a reply costs the run something, and that a run which only counted arrivals can talk about rates and times but not about whether the answers were right.
Explain the two-tier shape — a fixed-cost check on every reply plus a deep check on a small sample — and why taking the timestamp before verifying keeps the measuring side's work out of the result.
Show that you match the verification level to the decision the run informs, and that you write the acceptance rule and the sampling rate into the report rather than leaving a reader to assume the answers were correct.
Own the standard: what level of verification a performance result must carry before it may support a release discussion, and where an arrival-only run is genuinely sufficient.
## The measuring side competes with the load it applies Everything a run does per reply is paid for out of the same budget it uses to issue the next request. Parsing a document, walking a structure, comparing against an expected value — each costs processor time and memory on the applying side, and each one delays the next request by a small amount. Do it once and the cost is invisible. Do it hundreds of thousands of times a minute and two things happen: the demand the run can actually apply falls, and the times it records include work the system under test never did. That is the whole tension. Verification is what turns a delivery count into a result, and verification is the one thing a run cannot afford to do generously. ## Three levels, and what each one licenses you to say | Level | Cost per reply | What the run may claim | What it may not claim | |---|---|---|---| | Arrival only | Effectively free | The system accepted demand at this rate and answered within these times | That any answer was correct | | Constant-cost check | Small and fixed | Replies had the shape a genuine answer has | That the values inside them were right | | Deep check on a sample | Large, paid on one reply in N | Correctness on the sampled share, at the stated sampling rate | Correctness of the replies outside the sample | The middle row is where most runs should live. A required marker, a length floor and a row-count floor together cost a handful of comparisons and rule out the false successes that would otherwise be counted as work. The top row is not a failure — an arrival-only run is a legitimate and useful measurement — but its claims are narrow and the report has to say so. ## Keeping the checking out of the timed section Two techniques recover most of the cost: 1. **Stop the clock before you verify.** Record the elapsed time when the last byte arrives, then run the check. The reply is still validated, the outcome is still counted, and the recorded time no longer includes the measuring side's own work. 2. **Defer the expensive part.** Keep a small sampled set of replies, or a hash of each, and do the deep comparison after the run has stopped. Nothing about the correctness finding changes; the cost simply moves out of the window where it distorts the measurement. What does not work is scaling verification with load. A check whose cost grows with reply size, or that touches shared state across the applying side, degrades exactly when the run gets busy — so the run applies less demand at the very moment you wanted the most. ## What an unverified run may and may not claim This is the part that gets overstated in reports. A run that inspected nothing beyond arrival supports a real and useful statement: under this applied demand, the system accepted requests at this rate and produced replies within these times. It does not support any statement about correctness, and in particular it does not support the sentence people actually write, which is that the feature *worked* at that load. A useful discipline for the summary: - State the applied demand and the measured window. - State the acceptance rule — what made a reply count as accepted. - State the sampling rate of any deep check, and treat unsampled replies as unverified rather than as passed by association. - Never let a check failure be swallowed. If a reply fails verification it is a counted failure, not a logging event. ## Choosing the level The honest way to pick is to ask what decision the run informs. A capacity sizing exercise can often live with arrival-only counting, because the question is about rate. A run whose result will be quoted in a release discussion cannot: a system that answers quickly and wrongly under load is worse than one that answers slowly and correctly, and only verification distinguishes them. The cost of the constant-cost check is small enough that the default should be to pay it, and to escalate to a sampled deep check when the answers are complex enough that shape alone does not tell you much.
- Why does stopping the clock before verifying change the numbers at all?Because otherwise the recorded time is the system's response time plus the measuring side's parsing time, and the second term grows with reply size and with how busy the applying side is. Taking the timestamp at the last byte and verifying afterwards keeps the measurement about the system, while the reply is still checked and its outcome still counted.
- A colleague wants full assertions on every reply so the run doubles as a functional test. What do you say?That it makes both jobs worse. Full assertions cap the demand the run can apply, so the load result is about a smaller load than intended, and functional coverage bought this way is far more expensively obtained than in a suite that is not simultaneously trying to saturate something. Keep the constant-cost check for counting honesty and put the correctness work where it belongs.
- How do you report a run where only one reply in a hundred was deeply checked?State the sampling rate beside the correctness finding and treat the other ninety-nine as unverified rather than as passed. A sample supports a statement about the sampled share and, with care, about the rate of a defect; it does not license the sentence that the replies were correct.
An inspector on a production line either examines every item and slows the line, or examines a sample and states the sample size; what is not allowed is examining a sample and reporting on the whole batch.
saying these in an interview costs you the question
- Assumes verification on the applying side is free
- Reports that the feature worked when nothing beyond arrival was checked
- Includes parsing time inside the recorded response time
- Treats unsampled replies as verified because the sample passed
- Logs a failed content check instead of counting it as a failure