Response times climb across a long steady-load run while per-request resource use stays flat - what do you check?
answer
- Does it track work or the clock?
- Flat ratios can hide a falling denominator
- Probe a path the run never touches
- Restart the process, keep its data
- Trend quantities, not only rates
basics
~20 sCheck whether the climb advances with work done or with elapsed time. Replot it against cumulative completed work, run a low-rate control, probe a path the run never writes to, and restart the process keeping its data.
solid answer
~50 sTwo stories fit that picture: the system is accumulating something as it works, or its surroundings changed while the clock advanced. Separate them by asking what the climb tracks. Replot the response-time trend against cumulative completed work instead of clock time — if it straightens, the cause advances with work and lives inside; if it still follows the clock, look outward. Support that with a control period at a much lower rate, a small constant probe on a path touching none of the state the run writes, and an independent observer on the shared infrastructure. Then split a restart: keep the data the run wrote and the slowdown persists if the cost is in persisted state; it clears if the cost was in process state. Also distrust the flat per-request figure itself, since a falling throughput keeps that ratio level while the system does less.
code
pseudocode · 19 lines# Does the climb advance with work done, or with the clock?
per_work = slope(high_percentile_response_time, against = cumulative_completed_operations)
per_clock = slope(high_percentile_response_time, against = elapsed_minutes)
control = probe_touching_no_state_the_run_writes() # tiny, constant, all hold long
observer = idle_measurement_on_shared_infrastructure()
if per_work > 0 and control.is_flat:
conclude("build-up inside the system under test")
if per_clock > 0 and (control.is_rising or observer.is_rising):
conclude("something shared or around the run changed during the hold")
if completed_operations.fell_during(hold):
warn("flat per-request figures may be a moving denominator, not stability")
# confirm from the other side
restart(process, keep = data_written_by_run) # persists -> cost is in written state
restore(fresh_data, keep = process) # clears -> confirms the samego deeper
Be ready to say that a system can get slower without any single resource figure looking wrong, and that the first useful move is to look at how the numbers changed across the hours rather than at one reading.
Explain why a per-request figure can stay level while things get worse, since the denominator moves too, and describe plotting the climb against work completed instead of against the clock.
Show a diagnosis plan that runs during the hold rather than after it: a control probe on untouched paths, an independent observer on shared infrastructure, quantity counters, and the restart that keeps the written data.
Own the default rule that an unattributed slowdown is repeated with observers in place rather than filed as environmental, and make that instrumentation standing rather than something requested run by run.
## Two families of explanation, and both fit the same picture A hold at a fixed rate where response times climb while every per-request resource figure stays level is the most commonly misdiagnosed result of this run shape, because it fits two very different stories equally well. **Something inside the system is accumulating as it works.** Growing collections and caches whose hit rate falls as the working set outgrows the limit, more entries to scan, storage fragmenting, queues lengthening, sessions or connections never released, and data written by the run itself making each subsequent lookup do slightly more work. The common signature is that the cost advances with work done. **Something around the system changed while the clock advanced.** A co-tenant's workload arriving on shared infrastructure, a scheduled batch elsewhere in the estate of environments and datasets the team tests against, a shared volume filling, throttling on shared hardware after a sustained period at high utilisation, or somebody deploying into the environment mid-hold. The common signature is that the cost advances with elapsed time, whatever the system was doing. Flat per-request figures do not settle it, and they are also frequently an artefact: a ratio hides a moving denominator. If completed work fell during the hold and total resource use fell with it, per-request use is flat while the system is doing less than it was. ## Observations that separate the two 1. **Replot the climb against cumulative completed work instead of clock time.** This is the single most informative move. If the curve straightens against work done, the cause advances with work and lives inside. If it keeps tracking the clock no matter how much work was completed, look outward. 2. **Run a control at a much lower rate for the same duration.** The same climb at a tenth of the work is a time-driven effect; no climb at a tenth of the work is work-driven. 3. **Keep a constant, tiny probe on a path that touches none of the state the run writes.** If the probe slows in step with everything else, the shared surroundings are implicated, because the accumulated state cannot reach it. 4. **Keep an idle observer on the shared infrastructure**, measuring something the system under test does not use at all. 5. **Split the restart.** Restart the process but leave the data it wrote in place: if the slowdown survives, the build-up is in persisted state. Restore fresh data and leave the process untouched: if the slowdown clears, the same conclusion is confirmed from the other side, and if it does not, the cost is in process state. 6. **Trend quantities, not only rates**: entries held, rows written by the run, on-disk footprint, queue depth, open connections. A quantity that grew all hold is a candidate cause even when every rate looks calm. 7. **Ask who else was there.** Deployment records, other teams' schedules and shared-storage usage across the same hours are evidence, and they are only available if someone asks while the hold is still recent. | Observation | Points inside the system | Points to the surroundings | |---|---|---| | Climb replotted against completed work | Straightens | Still tracks the clock | | Control at one-tenth the rate | No climb | Same climb | | Probe on untouched paths | Stays flat | Slows in step | | Restart keeping written data | Slowdown persists | Slowdown clears | | Quantity counters across the hold | One grows steadily | All level | ## Both can be true, and the cheap instruments settle it during the run These are not exclusive: a system can accumulate slowly while a shared volume also fills. The reason to name the discriminators in advance is that almost all of them cost nothing and must be running *during* the hold to be useful. A control probe, an idle observer and a cumulative completed-work series are minutes of setup; reconstructing any of them after a twenty-six hour hold means holding again. When writing the result up: - State which family the evidence supports and what observation would change that conclusion. - Keep the raw series rather than summary figures, because the next hold will be compared against this one and a summary cannot be replotted. - Treat an unattributed climb as a reason to repeat the hold with the observers in place, not as grounds to file it as environmental. Environmental is the comfortable conclusion, it requires no work, and it is the one that lets a genuine build-up walk into production wearing a note that says the staging machine was busy.
- Why are flat per-request resource figures weak evidence that nothing is accumulating?A ratio hides a moving denominator. If completed work fell during the hold and total resource use fell with it, the per-request figure stays level while the system is doing less than it was. Read the absolute series and the completed-work series alongside the ratio; flat per-request use with falling throughput is a warning rather than reassurance.
- What cheap instrumentation would you add before starting a long hold so this is answerable afterwards?A constant low-rate probe on a path that touches none of the state the run writes, an independent observer measuring the shared infrastructure, cumulative completed-work counts recorded alongside every resource series, and quantity counters such as entries held, rows written, on-disk footprint and queue depth. All are minutes of setup, and none can be reconstructed once a day-long hold has finished.
A shop that gets slower through the day might be filling with clutter, or the whole street might be getting busier. Watching a customer who never enters the cluttered aisles tells you which.
saying these in an interview costs you the question
- Blames the environment because no resource figure moved
- Treats flat per-request use as proof nothing accumulates
- Never plots the climb against completed work
- Ignores who else was using the shared infrastructure
- Concludes from a single unrepeated observation
- Files an unexplained slowdown as environmental by default