skip to content

Every route on an asynchronous service is slow; what evidence separates parked workers from one slow dependency?

level: seniorimportance: should knowfreq 50%

answer

  1. ask which routes actually got slower
  2. unrelated routes means a shared resource
  3. high latency with near-idle processors
  4. repeated stack samples, never just one
  5. queue time before the stage starts

basics

~20 s

Ask which routes degraded. A slow dependency hurts only its callers; parked workers hurt routes with nothing in common, alongside near-idle processors, growing time before a stage starts, and stack samples finding the pool waiting.

solid answer

~50 s

The discriminator is the *shape* of the degradation, not its size. A slow dependency hurts the routes that use it and leaves the rest alone; parked workers hurt routes that share no code, no data store and no dependency, because the only thing they share is the pool. Three more pieces of evidence corroborate it: processor usage sits low while latency climbs, which is waiting rather than computing; the delay appears **before** the stage runs, so time-to-start grows while the stage's own duration is flat; and repeated snapshots of every worker's stack keep finding the small pool inside the same waiting call. A blocking-call detector in the test suite settles it earlier and more cheaply. The two stories often coexist — a dependency slowed *and* its stage was never isolated — and the fix then addresses both.

go deeper

for a junior

Know the first question to ask: which routes got slower? If routes with nothing in common slowed together, the cause is something they all share rather than any one dependency.

for a middle

Explain why waiting shows as high latency with idle processors, and why splitting queue time from service time tells you whether requests waited for a worker or waited on the call.

for a senior

Walk through the investigation you would actually run: latency broken down by route, repeated worker stack samples, caller-versus-dependency timings, then the isolation and the timeout together.

for a principal

Decide what standing evidence the platform should carry so this is answered in minutes: a time-to-start objective, worker sampling always available, and a detector in every build.

## Two stories, one graph A service-wide latency chart looks identical for two different causes. Either a dependency you call became slow, or your shared workers are being spent waiting. They demand different work — one is an upstream conversation and a timeout, the other is a change to where a stage runs — so guessing is expensive. What separates them is available within minutes if you ask for the right cut of the data. ## The discriminator: which routes degraded Break latency down by route and ask what the degraded routes have in common. - If the degraded set is exactly the routes that call one dependency, and routes that do not call it are healthy, the dependency is the story. - If routes that share **no** dependency, no data store and no handler code degraded together, they are sharing something else, and in an asynchronous service the thing they share is the worker pool. - If the degradation began when one feature's traffic rose, while that feature's own latency is only mildly worse, look at that feature's stage: the damage it exports exceeds the damage it suffers. | Evidence | Slow dependency | Parked workers | |---|---|---| | Which routes degraded | Only callers of that dependency | Unrelated routes too | | Processor usage | Unchanged, usually low | Low, while latency climbs | | Time before the stage starts | Flat | Rising | | Dependency's own reported service time | Rising | May be unchanged | | Worker stack samples | Workers idle, waiting for work | Workers inside a waiting call | ## Evidence you can collect during the episode 1. **Sample every worker's stack, repeatedly.** One snapshot proves nothing — it catches whatever was running that instant. A series taken seconds apart that keeps finding the small pool's workers inside the same call is the direct signature. The complementary reading matters too: workers found idle and waiting for work mean the pool is not the constraint. 2. **Separate queue time from service time.** Record when a request was accepted and when its first stage actually ran. Time spent waiting for a worker shows up as a growing gap before the stage starts, while the stage's own duration is unchanged. An end-to-end number alone cannot make that split, which is why it cannot answer the question. 3. **Compare latency measured at the caller with latency the dependency reports.** If the dependency says it served in 20 ms and your stage measured 900 ms, the missing 880 ms was spent somewhere on your side — most often queued for a worker. 4. **Watch processor usage against latency.** Rising latency with idle processors is waiting; rising latency with saturated processors is a different failure and a different fix. ## Evidence you can collect before the episode - A **blocking-call detector** marks known waiting operations and alarms when one runs on a worker of the non-blocking pool. Running in the test suite it turns a future incident into a failing test, and it catches waits in code you did not write. It is not proof of production cleanliness: it only knows the operations it has been taught, and it usually carries enough overhead that teams do not leave it on everywhere. - **Cold-start runs** surface first-use waits that a warm load test averages away. - **A latency objective on time-to-start**, not just on end-to-end latency, gives you an alarm that fires on this failure specifically rather than on everything at once. ## When both are true The common production case is not either/or. A dependency got slower, and the stage that calls it was never isolated, so its extra latency multiplied into occupancy and spread across the service. Little's Law makes that mechanical: occupancy is arrival rate times wait, so doubling the wait doubles the workers held at the same traffic. The response has two parts, and they are independent: 1. Isolate the stage onto a worker set sized for waiting, with a bound, so the dependency's latency stops being charged to unrelated routes. 2. Put a timeout on the call and take the slow dependency up with whoever owns it, because isolation caps the blast radius without making anything faster. Doing only the first leaves a feature quietly broken behind a healthy-looking service. Doing only the second leaves the whole service hostage to the next dependency that slows down.

  • The dependency really is slower and the workers really are parked. What do you do first?
    Treat them as two fixes. Isolate the stage onto a worker set sized for waiting with a bound, so unrelated routes stop paying for that dependency, and put a timeout on the call. Then pursue the dependency's own latency with its owner. Isolation caps the blast radius; it makes nothing faster.
  • What does a blocking-call detector give you that production evidence cannot?
    Timing and coverage. It fails a test before the code ships, and it flags waits inside libraries nobody inspected. Its limits are real: it only knows the operations it has been taught, so a clean run is not proof, and its overhead usually keeps it out of production.
  • Why is a single stack snapshot poor evidence?
    It captures one instant and is easily unrepresentative in both directions — it can miss a wait that dominates the minute, or catch a rare one and overstate it. Repeated samples over the episode give a distribution: what fraction of the time the pool's workers were inside a waiting call.

saying these in an interview costs you the question

  • Reads slow routes as automatic proof a dependency is slow.
  • Dismisses blocking because processor usage is low.
  • Takes one stack snapshot and declares the pool healthy.
  • Measures only end-to-end latency, never time before the stage starts.
  • Assumes a clean detector run in tests proves production is clean.
  • Stops at isolation and never raises the dependency's latency with its owner.