A node reports wait time and service time per request — what does each measure and what does a rise in each mean?
answer
- two clocks on one request
- queued, then handled
- waiting means capacity, handling means work
- slow work occupies handlers for everyone
- held requests may land in either
basics
~20 sWait time is how long a request sat in the node's request queue before a handler took it; service time is the handling itself. Rising wait means too little handler capacity for the arrival rate; rising service means the work got slower.
solid answer
~50 sMost broker implementations accept a request off a connection, place it in a request queue, and let a bounded pool of handlers pick it up — so one request's duration divides into **wait time** in that queue and **service time** in a handler. The split is what makes the number diagnosable: a request that took 300 ms because it queued for 290 ms has a capacity problem, while one that queued for 2 ms and was handled for 298 ms has a work problem, and the two lead to entirely different investigations. Requests that must be held until a condition is met — a write waiting for other copies, a read waiting for records to arrive — are counted inside service time on some platforms and into a separate held bucket on others, so establish which before reading a large service time as slowness.
go deeper
Recall that a request first waits for a handler and is then handled, and that the total on its own does not say which of the two took the time.
Explain what each half responds to — arrival rate and free capacity against the cost of the work — and why a service-time problem in one stream shows up as wait time for every other client.
Show the diagnosis: read the pair together, check where held requests are counted before calling anything slow, and split by request kind and client to find the traffic responsible.
Decide what the platform publishes per node so stream owners can tell a shared-capacity problem from their own workload, and what it costs to keep that split on every cluster.
## The two halves of one request A broker node does not handle a request the instant it arrives. The common shape is: the connection layer reads the request off the socket, puts it in a **request queue**, and a bounded pool of **handlers** takes requests from that queue and does the work — appending records, reading them back, answering a metadata question. That gives one request two clocks: - **Wait time** — from being enqueued to being picked up by a handler. Nothing is happening to the request; it is waiting for capacity. - **Service time** — from pickup to the response being ready. This is the work itself. Their sum is roughly what the client waited inside the node, excluding the network in both directions. ## What a rise in wait time is telling you Wait time rises when requests arrive faster than handlers free up. The causes sit on either side of that sentence: - **More arrivals** — a new client, a burst, a reader group that just restarted and is re-reading, a retry storm amplifying the original rate. - **Fewer free handlers** — the pool is occupied. And this is the coupling to watch: handlers held for a long time by *slow work* push wait time up for every unrelated request behind them, so a service-time problem in one stream shows up as a wait-time problem for everyone. The practical reading is that wait time is a **shared-resource** number. When it moves, the whole node's clients are affected, including ones doing nothing unusual. ## What a rise in service time is telling you Service time rises when the work per request got more expensive, which is much more specific: - reads that used to be answered from memory now come off the volume, because readers fell far enough behind that the records they want are no longer resident; - requests carry more data than they did; - on platforms where an accepted write is held until other copies hold it, that wait is inside the write's service time, so a lagging copy appears here rather than in the node's own work; - the node is contending for its own resources with something else on the machine. ## Where held requests land, and why you must check A request that is deliberately parked — a read told to wait until records arrive or a size threshold is met, or a write waiting on other copies — is not being handled and is not queueing for capacity either. Platforms count that time differently: inside service time on some, in a separate held bucket on others, and on rented services often nowhere you can see. A long-poll read that waited because *no records existed* is healthy; counted naively it looks like a very slow request. Establishing where held time lands is the first step before reading either half. ## Reading the pair together | Wait time | Service time | Most likely reading | |---|---|---| | High | Normal | Not enough handling capacity for the current arrival rate | | Normal | High | The work itself got more expensive for this request kind | | High | High | Slow work is occupying handlers and backing up everything behind it | | Normal | Normal, but clients report slowness | The delay is outside this node — network, the client's own queue, or further down the path | ## What varies across platforms - Not every implementation exposes both halves. Some report only a total duration; some report the queue depth but not the waiting time; rented services often publish neither. - The pool that matters may not be a single one. Some designs separate the threads that read from sockets from those that do the work, and either can be the constrained one. - Where a platform can add handling capacity dynamically, a wait-time rise may resolve itself; where the pool size is fixed at start, it will not. ## What the split cannot tell you It localises the delay to a half of one hop, and no further. It does not say which client or stream caused it — that needs the same numbers split by dimension. It does not say what the node will do if the pressure continues; that behaviour is a separate subject. And it says nothing about the time before arrival or after the response, so a healthy split is not evidence that the whole path is healthy.
- Wait time on one node is high while service time is normal. Name two genuinely different causes.Either arrivals rose — a new or restarted client, a burst, retries amplifying the original rate — or handlers are occupied by long-running work that does not show in the service time of the requests now waiting. The second is the trap: the unaffected-looking requests are queueing behind somebody else's expensive ones, so you look at per-kind service time, not only at the aggregate.
- Why can a read that waited 500 ms be perfectly healthy?Because readers commonly ask the node to hold a read until records arrive or a size threshold is met, rather than answering immediately with nothing. That wait is idle time by design, not slowness. If the platform counts it inside service time, healthy long polls inflate the number, which is why held time has to be identified before any threshold is set on it.
saying these in an interview costs you the question
- Treats one total request duration as enough to diagnose the node
- Assumes rising wait time is never caused by slow handlers
- Reads a long held read as a slow node
- Thinks service time covers the network back to the client
- Concludes the whole path is healthy from a clean split