How do you attribute a performance run's slow response tail to one step on the request path when each step's own average looks acceptable?
answer
- Averages hide a slow tail
- Slow on a few requests, ordinary on average
- Per-call cost times calls per request
- Compare slow requests against median requests
- Neutralise the suspect and re-run
basics
~20 sPer-step averages hide a step that is slow occasionally or runs many times per request. Compare each step's share of total time in the slowest requests against median ones, then confirm by neutralising the suspect.
solid answer
~50 sA per-step average is the wrong statistic for a tail. A step that is slow on one call in fifty barely moves its own mean while setting the slow end of the whole path, and a step averaging two milliseconds that runs sixty times per request costs more than one averaging forty that runs once. Two corrections fix it. First, **compare compositions rather than aggregates**: take the slowest slice of requests and a median slice, break each slice's total duration down by step, and look for the step whose *share* is larger in the slow slice. A step taking the same share in both is not the cause. Second, **weight by invocation count**, since contribution is per-call cost multiplied by calls per request. Then check the residual — the part of each request's duration no step accounts for — and confirm by re-running the identical profile with the suspect neutralised.
code
pseudocode · 13 linesslow = requests where total_duration >= slowest_one_percent_boundary
median = requests within a narrow band around the median duration
for step in path_steps:
share_slow = sum(step.time over slow) / sum(total_duration over slow)
share_median = sum(step.time over median) / sum(total_duration over median)
growth[step] = share_slow - share_median
calls[step] = mean(step.invocations per request)
residual_slow = 1 - sum(share_slow over all steps)
suspect = step with the largest growth
# confirm: rerun the identical profile with suspect returning a fixed responsego deeper
Recall that a request's total time is made up of the steps it passes through, and that a step's average duration is not the same as how much of a slow request's time it accounted for. Know that a step can be cheap per call and still dominate.
Explain why an average hides a tail: rare slowness barely moves a mean, and contribution is per-call cost multiplied by invocations per request. Be able to compute each step's summed contribution rather than ranking per-call figures.
Demonstrate the composition comparison between slow and median requests, the residual check for time no step covers, and the discipline of confirming an attribution by neutralising the suspect and re-running the identical profile.
Own what the path is instrumented to reveal. Decide which boundaries must be timed for attribution to be possible at all, weigh the recording cost against the hours a blind investigation burns, and set the bar of evidence required before a team rebuilds a step.
## Why per-step averages hide the culprit When a path is a sequence of steps and the whole path has a slow tail, the instinct is to rank the steps by their own average duration and blame the top of the list. That instinct fails in two specific and very common ways. **Rare slowness averages away.** A step that responds in 3 ms forty-nine times out of fifty and 800 ms on the fiftieth has a mean near 19 ms — unremarkable next to its neighbours. But it lands on two percent of requests, which is precisely the population that forms the slow end of the path. Its mean says nothing; its own slow end says everything. **Multiplicity averages away too.** A step's *average* is per call, while a request's duration accrues per *invocation*. A step averaging 2 ms invoked sixty times contributes about 120 ms; a step averaging 40 ms invoked once contributes 40. Ranked by average, the second looks three times worse. Ranked by contribution, the first is three times worse. Repeated small calls in a loop are the classic instance and are invisible to any per-call ranking. ## Compare compositions, not aggregates The fix is to stop comparing steps to each other and start comparing the *same* step across different populations of requests. 1. Split the run's completed requests into two slices: the slowest small fraction, and a band around the median. 2. For each slice, sum each step's time and express it as a **share of that slice's total duration**. 3. Look for the step whose share is substantially larger in the slow slice than in the median slice. | Pattern across the two slices | Reading | | --- | --- | | A step's share is much larger in the slow slice | That step is where the slow requests differ | | A step's share is the same in both slices | It is expensive everywhere, not the cause of the tail | | Every step's share is similar, totals differ | Slowness is spread evenly — suspect a shared cause | | Shares sum to well under one in the slow slice | The missing time is between steps, not inside one | The third row usually points at something outside the steps — the requests waited to start, or every step was affected at once. The fourth row is the residual check and it is worth doing every time. ## Weight by how often each step runs Alongside share, record invocations per request for each step, and separate two quantities that get confused: the **per-call cost** and the **calls per request**. A step can dominate a request for either reason, and the remedies are entirely different — make the call cheaper, or stop making so many of them. The composition comparison above naturally captures both, because it works with summed time per request rather than per-call means, but keeping the invocation count beside the share tells you which lever to reach for. ## Correlate within a request, not across the clock A seductive shortcut is to plot each step's slow moments over the run and look for steps whose spikes line up. That grouping finds steps sharing a *cause* — everything was slow at once — but it does not attribute one request's time to one step. Attribution has to be per request: for this specific slow request, where did its time go. Two steps whose slow periods coincide may both be victims of a third thing entirely. ## Mind the residual Sum a slow request's attributed step durations and subtract from its own end-to-end duration. What remains is time the request spent somewhere no step covers: waiting to be admitted before the first step ran, being handed between steps, being serialised on the way out, or inside a stretch nobody measured. A large residual in the slow slice and a small one in the median slice is a strong result in itself — it says the tail is not in any step you are looking at, and it redirects the investigation instead of letting it grind through step rankings that cannot contain the answer. ## Confirm the attribution before acting on it A correlation across two slices is a hypothesis, not a finding. Confirm it by changing the suspect and nothing else, then re-running the identical profile against the same build and data: - Replace the suspect step with a fixed, immediate response and see whether the path's slow end collapses. - Or make its slow behaviour deliberate and worse, and see whether the path's slow end grows in proportion. Either direction is decisive in a way that ranking never is. If the tail is unchanged when the suspect is removed, the attribution was wrong however convincing the shares looked, and the residual is where to look next. This confirmation step is also what separates a diagnosis from a guess when the fix is expensive: it is cheap to run one more comparison and costly to rebuild the wrong step.
- Every step's share of total time is the same in your slow slice and your median slice, yet the slow requests took five times as long. What now?Nothing on the path is disproportionately responsible, so the cause is shared: every step was stretched together. That points at something outside the steps — waiting for admission before the first step, an environment-wide pause, or contention that slowed all work at once. Check the residual first, then look for a factor common to all steps rather than continuing to rank them against each other.
- How do you separate a step that is expensive per call from one that is called too many times?Record both figures per request: the summed time in that step and the number of invocations. Their ratio is the per-call cost. A step with high summed time and one invocation needs the call itself made cheaper; a step with high summed time and dozens of invocations needs the calls consolidated or removed. Ranking on per-call means alone conflates them and points at the wrong fix.
- What is the risk of acting on a composition comparison without a confirming run?Share is correlation. A step can carry more of the slow requests' time because it is downstream of the real cause, because slow requests systematically carry heavier inputs into it, or because it is where the waiting happens to be recorded. Neutralising the suspect and re-running the identical profile turns the hypothesis into a result, and it is far cheaper than rebuilding a step that was never the problem.
saying these in an interview costs you the question
- Ranks steps by their own average and blames the top
- Ignores how many times a step runs per request
- Groups steps by when their slow periods coincided
- Never checks the time no step accounts for
- Acts on a share comparison without a confirming run
- Assumes the slowest single call caused the whole tail