A slowdown reproduces only when requests overlap. How do you design the follow-up performance run that confirms it?
answer
- Costs nothing until two requests want it
- A single user never creates overlap
- Same arrival rate, different number of clients
- Distinct records versus one shared record
- Time acquiring versus time holding
basics
~20 sVary overlap alone. Run the same arrival rate with few busy clients and then with many idle ones: if duration differs, overlap is the cause. Repeat with requests spread across distinct records to locate what is shared.
solid answer
~50 sContention costs time only while two requests want the same thing at once, so a single-user reproduction is guaranteed to be clean — the wait is always zero and the operation looks fast forever. The confirming run therefore has to **manipulate overlap directly**. The sharpest design produces the *same arrival rate* two ways: a few clients each sending often, and many clients each sending rarely. Arrival rate is identical, requests in flight differ by an order of magnitude, and any difference in duration is attributable to overlap rather than to rate. Two variants sharpen it further: run the suspected operations alone and then mixed at the same individual rates, and run all requests against distinct records and then against one shared record. Record time spent acquiring the shared thing, not just end-to-end duration, so the confirmation is read rather than inferred.
code
pseudocode · 11 linestarget_rate = 200 per second # identical in both runs
run_low_overlap = hold(clients = 20, pause = 0.1s, rate = target_rate)
run_high_overlap = hold(clients = 200, pause = 1.0s, rate = target_rate)
# same arrivals per second, about ten times the requests in flight
# then locate what is shared
run_spread = hold(clients = 200, each client uses a distinct record)
run_shared = hold(clients = 200, all clients use one record)
record per request: time_to_acquire(shared_thing), time_holding(shared_thing)go deeper
Recall that some costs appear only when two requests want the same thing at the same moment, and that running an operation repeatedly on its own never creates that situation. Know that overlap, not repetition, is what has to be varied.
Explain the manipulation: hold arrival rate constant and change how many clients produce it, so requests in flight change while work per second does not. Be able to say why that separates overlap from simply being busier.
Show a full confirming design — a concurrency ladder, a distinct-record versus shared-record pair, an alone-versus-mixed pair — and read the shapes: gradual growth, a cliff, or growth faster than overlap. Insist on acquisition time recorded separately from holding time.
Own when this investigation is worth its cost. Decide how much overlap-sensitive testing a system's traffic pattern justifies, what evidence must exist before a shared section is redesigned, and how to keep such runs comparable as the system and its data change.
## Why one user cannot see it Some costs exist only in the overlap. An exclusive section over one record, a single cache entry being refilled, a shared counter, an allowance shared across callers, a connection used by more than one request — each costs nothing at all while only one request wants it, and costs real time the moment a second one does. At a concurrency of one the wait is structurally zero: there is never a second holder, so the code path is exercised fully and still looks fast. That is why the usual reproduction attempts fail. Repeating the operation more times does not help, because repetitions are sequential. Running it on a faster machine does not help. Adding recording around the suspect section does not help either, because the section is genuinely quick when uncontended. The only variable that matters is how many requests are inside the shared thing at the same moment, and a single user cannot vary it. ## Vary overlap, hold everything else still The design principle is one manipulated variable. Overlap is the variable; the build, the dataset and its volume, the environment, the request mix and the total work must be identical across the runs being compared, or the comparison decides nothing. The strongest single design separates two things that usually move together: 1. Choose one arrival rate — say two hundred requests per second — and reach it twice. 2. **Run A:** twenty clients, each pausing briefly between requests. Few requests in flight. 3. **Run B:** two hundred clients, each pausing much longer between requests. The same two hundred arrivals per second, roughly ten times the overlap. 4. Compare duration and, if recorded, time spent acquiring the shared thing. If duration is materially worse in Run B at an identical arrival rate, throughput is not the cause — overlap is. This is the design that answers the objection "the system was simply busier", because it was not: it did exactly the same amount of work per second in both runs. ## Two variants that locate what is shared Confirming that overlap matters is half the job; the other half is finding what is being shared. | Run pair | What changes | What a difference proves | | --- | --- | --- | | One arrival rate, few clients vs many | Requests in flight | The cost depends on overlap, not on rate | | Distinct records vs one shared record | What requests collide on | The contention is over that record, not the code path | | Operation alone vs mixed with another | Which operations coexist | The two operations contend over something common | | Concurrency ladder at 1, 2, 4, 8, 16 | Degree of overlap | Whether the cost grows smoothly or has a cliff | The second row is the most decisive and the most often skipped. If the same code path at the same overlap is fast when every request touches a different record and slow when they all touch one, the shared record is the constraint and no amount of profiling the code will show it. The third row catches the case where two operations look individually healthy and only their combination is not. The fourth row distinguishes shapes. A cost that grows gently with overlap suggests a queue in front of something shared. A cost that is flat and then jumps at a particular overlap suggests a fixed allowance being exhausted. A cost that grows faster than overlap suggests the contention itself is adding work — retries, repeated attempts, or a section that gets longer as more callers arrive. ## Make the confirmation readable, not inferential End-to-end duration is a weak signal for this. Before the confirming run, add two recordings around the suspected shared thing: **time spent acquiring it** and **time spent holding it**. Contention shows up unmistakably as acquisition time growing with overlap while holding time stays flat. If instead holding time grows too, the shared section is getting longer under load and the diagnosis changes. Two practical cautions. First, run each configuration more than once and look at the run-to-run spread before believing a difference; a single pair of runs can differ for reasons that have nothing to do with your variable. Second, keep the total work identical: comparing a run that completed a hundred thousand requests against one that completed forty thousand compares two different experiments, however matched their settings looked on paper. What you should end with is a monotone relationship between overlap and cost at a constant arrival rate, and the disappearance of that relationship when the shared thing is removed from the picture. That pair of results is a confirmation. One slow run at high concurrency is not.
- Duration is worse at the same arrival rate with more clients. What alternative explanation must you rule out before calling it contention?That the two runs differed in something other than overlap. More clients can mean more connections, a different distribution of requests across records, a warmer or colder cache, or a measuring harness working harder. Hold the build, dataset, environment and request mix identical, repeat each configuration to establish run-to-run spread, and confirm the effect exceeds that spread before attributing it to overlap.
- The cost is flat up to a certain overlap and then jumps sharply. What does that shape suggest?A fixed allowance being exhausted rather than a queue growing gradually. Up to the allowance nothing waits; beyond it every additional request waits for a release. A gently rising cost points instead at a line in front of something shared, and a cost rising faster than overlap points at contention that adds work of its own, such as repeated attempts. The shape narrows the search before any code is read.
One person can walk through a revolving door all day without noticing it only fits one at a time; the constraint appears only when a second person arrives at the same moment.
saying these in an interview costs you the question
- Tries to reproduce the slowdown with a single user
- Concludes contention from one high-concurrency run
- Changes arrival rate and overlap in the same comparison
- Never varies whether requests touch the same record
- Compares runs that completed different amounts of work
- Measures only end-to-end duration around a shared section