Every route on an asynchronous service is slow; what evidence separates parked workers from one slow dependency?
answer
- ask which routes actually got slower
- unrelated routes means a shared resource
- high latency with near-idle processors
- repeated stack samples, never just one
- queue time before the stage starts
basics
~20 sAsk which routes degraded. A slow dependency hurts only its callers; parked workers hurt routes with nothing in common, alongside near-idle processors, growing time before a stage starts, and stack samples finding the pool waiting.
solid answer
~50 sThe discriminator is the *shape* of the degradation, not its size. A slow dependency hurts the routes that use it and leaves the rest alone; parked workers hurt routes that share no code, no data store and no dependency, because the only thing they share is the pool. Three more pieces of evidence corroborate it: processor usage sits low while latency climbs, which is waiting rather than computing; the delay appears **before** the stage runs, so time-to-start grows while the stage's own duration is flat; and repeated snapshots of every worker's stack keep finding the small pool inside the same waiting call. A blocking-call detector in the test suite settles it earlier and more cheaply. The two stories often coexist — a dependency slowed *and* its stage was never isolated — and the fix then addresses both.
go deeper
Know the first question to ask: which routes got slower? If routes with nothing in common slowed together, the cause is something they all share rather than any one dependency.
Explain why waiting shows as high latency with idle processors, and why splitting queue time from service time tells you whether requests waited for a worker or waited on the call.
Walk through the investigation you would actually run: latency broken down by route, repeated worker stack samples, caller-versus-dependency timings, then the isolation and the timeout together.
Decide what standing evidence the platform should carry so this is answered in minutes: a time-to-start objective, worker sampling always available, and a detector in every build.
## Two stories, one graph A service-wide latency chart looks identical for two different causes. Either a dependency you call became slow, or your shared workers are being spent waiting. They demand different work — one is an upstream conversation and a timeout, the other is a change to where a stage runs — so guessing is expensive. What separates them is available within minutes if you ask for the right cut of the data. ## The discriminator: which routes degraded Break latency down by route and ask what the degraded routes have in common. - If the degraded set is exactly the routes that call one dependency, and routes that do not call it are healthy, the dependency is the story. - If routes that share **no** dependency, no data store and no handler code degraded together, they are sharing something else, and in an asynchronous service the thing they share is the worker pool. - If the degradation began when one feature's traffic rose, while that feature's own latency is only mildly worse, look at that feature's stage: the damage it exports exceeds the damage it suffers. | Evidence | Slow dependency | Parked workers | |---|---|---| | Which routes degraded | Only callers of that dependency | Unrelated routes too | | Processor usage | Unchanged, usually low | Low, while latency climbs | | Time before the stage starts | Flat | Rising | | Dependency's own reported service time | Rising | May be unchanged | | Worker stack samples | Workers idle, waiting for work | Workers inside a waiting call | ## Evidence you can collect during the episode 1. **Sample every worker's stack, repeatedly.** One snapshot proves nothing — it catches whatever was running that instant. A series taken seconds apart that keeps finding the small pool's workers inside the same call is the direct signature. The complementary reading matters too: workers found idle and waiting for work mean the pool is not the constraint. 2. **Separate queue time from service time.** Record when a request was accepted and when its first stage actually ran. Time spent waiting for a worker shows up as a growing gap before the stage starts, while the stage's own duration is unchanged. An end-to-end number alone cannot make that split, which is why it cannot answer the question. 3. **Compare latency measured at the caller with latency the dependency reports.** If the dependency says it served in 20 ms and your stage measured 900 ms, the missing 880 ms was spent somewhere on your side — most often queued for a worker. 4. **Watch processor usage against latency.** Rising latency with idle processors is waiting; rising latency with saturated processors is a different failure and a different fix. ## Evidence you can collect before the episode - A **blocking-call detector** marks known waiting operations and alarms when one runs on a worker of the non-blocking pool. Running in the test suite it turns a future incident into a failing test, and it catches waits in code you did not write. It is not proof of production cleanliness: it only knows the operations it has been taught, and it usually carries enough overhead that teams do not leave it on everywhere. - **Cold-start runs** surface first-use waits that a warm load test averages away. - **A latency objective on time-to-start**, not just on end-to-end latency, gives you an alarm that fires on this failure specifically rather than on everything at once. ## When both are true The common production case is not either/or. A dependency got slower, and the stage that calls it was never isolated, so its extra latency multiplied into occupancy and spread across the service. Little's Law makes that mechanical: occupancy is arrival rate times wait, so doubling the wait doubles the workers held at the same traffic. The response has two parts, and they are independent: 1. Isolate the stage onto a worker set sized for waiting, with a bound, so the dependency's latency stops being charged to unrelated routes. 2. Put a timeout on the call and take the slow dependency up with whoever owns it, because isolation caps the blast radius without making anything faster. Doing only the first leaves a feature quietly broken behind a healthy-looking service. Doing only the second leaves the whole service hostage to the next dependency that slows down.
- The dependency really is slower and the workers really are parked. What do you do first?Treat them as two fixes. Isolate the stage onto a worker set sized for waiting with a bound, so unrelated routes stop paying for that dependency, and put a timeout on the call. Then pursue the dependency's own latency with its owner. Isolation caps the blast radius; it makes nothing faster.
- What does a blocking-call detector give you that production evidence cannot?Timing and coverage. It fails a test before the code ships, and it flags waits inside libraries nobody inspected. Its limits are real: it only knows the operations it has been taught, so a clean run is not proof, and its overhead usually keeps it out of production.
- Why is a single stack snapshot poor evidence?It captures one instant and is easily unrepresentative in both directions — it can miss a wait that dominates the minute, or catch a rare one and overstate it. Repeated samples over the episode give a distribution: what fraction of the time the pool's workers were inside a waiting call.
saying these in an interview costs you the question
- Reads slow routes as automatic proof a dependency is slow.
- Dismisses blocking because processor usage is low.
- Takes one stack snapshot and declares the pool healthy.
- Measures only end-to-end latency, never time before the stage starts.
- Assumes a clean detector run in tests proves production is clean.
- Stops at isolation and never raises the dependency's latency with its owner.