Publish latency is normal but a node's handler-pool saturation climbed from 40% to 85% over an hour — why act on that?
answer
- occupancy fills before waiting appears
- flat, then steep
- queue depth confirms the slack is gone
- no lead when the work itself slows
- wrong pool predicts nothing
basics
~20 sBecause occupancy fills before waiting appears. Duration stays flat while a pool has spare handlers and then rises steeply once it does not, so saturation and request-queue depth usually turn first and buy lead time that latency does not.
solid answer
~50 s**Handler-pool saturation** is the share of a node's request-handling capacity in use, and **request-queue depth** is how many requests are waiting for a handler. While a pool has spare capacity, an arriving request is picked up almost immediately and total duration barely moves; once free handlers become scarce, requests start queueing and duration climbs steeply for a small further rise in arrivals. That non-linearity is the whole argument: latency is a late, cliff-shaped signal, while occupancy is a smooth one you can watch fill. A climb from 40% to 85% is an hour of warning you would not have had from the latency panel. The caveat is that saturation of a pool that is not the constrained resource predicts nothing, and where the slowdown is in the handling itself, duration and occupancy rise together with no lead at all.
go deeper
Recall that a busy-ness number and a duration number are different: the first can climb for a while before the second reacts at all.
Explain why waiting stays near zero while spare capacity exists and then rises steeply, and why request-queue depth is read next to occupancy rather than instead of it.
Demonstrate the triage: separate rising arrivals from more expensive work, check whether it is one node or the fleet, and say plainly when this signal gives no lead at all.
Weigh how much headroom the estate keeps against what it costs, and set what the platform publishes so stream owners can see the pressure they are contributing to before anyone is woken.
## Why occupancy turns before duration A node's request handling is a finite resource that requests contend for. The relationship between how busy that resource is and how long a request waits for it is not a straight line: - While free handlers exist, an arriving request is picked up almost at once. Waiting is near zero and the total duration is essentially the work itself. - As free handlers become scarce, an arriving request sometimes finds none and waits for one to finish. Waiting becomes non-zero and variable. - Near full occupancy, each further increment of arrivals produces a disproportionate increase in waiting, because the resource has no slack to absorb the ordinary burstiness of real traffic. The standing relationship between arrival rate, the number of requests waiting and the time each one waits is what makes this steep: at high occupancy, queue depth and waiting grow together for a small change in load. So **latency is a lagging, sharply-shaped signal, and occupancy is a leading, smooth one**. ## What the two numbers do for you | Signal | Shape | What it buys | |---|---|---| | Handler-pool saturation | Rises smoothly with load | Warning while duration is still flat | | Request-queue depth | Near zero, then climbs | Confirms the pool has actually run out of slack | | Request wait time | Flat, then steep | The moment clients begin to feel it | | Request service time | Tracks the cost of the work | Says the problem is the work, not the capacity | Saturation and queue depth are the pair: occupancy says how full the resource is, depth says whether anything is actually backing up behind it. Occupancy alone can be high on a node doing steady heavy work with no queue at all. ## Reading a climb from 40% to 85% The useful questions in order: 1. **Is this arrivals or the work?** Rising request rate at a steady cost per request is growth; steady request rate at a rising cost per request means the handling got more expensive and the pool is being occupied for longer. 2. **Is it one node or the fleet?** One node means something is concentrated there — traffic skew, a large client, partition leadership that happens to sit there. The whole fleet means the estate has grown into its capacity. 3. **Is there anything scheduled?** An announced change window explains a lot of one-node saturation; an unexplained fleet-wide climb does not. 4. **Has queue depth left zero?** If it has, the pool is already out of slack and the latency rise is imminent rather than hypothetical. ## What saturation does not tell you This signal is useful and easy to over-read: - **A pool that is not the constrained resource predicts nothing.** If the node is limited by its volume's throughput, watching handler occupancy fill will mislead you. - **There is no universal threshold.** The occupancy at which waiting takes off depends on how bursty arrivals are and how variable the work is; a steady, uniform workload tolerates a much fuller pool than a spiky one. - **It gives no lead when the work itself slows.** If the handling cost rises, occupancy and duration rise together and neither warns you about the other. - **It says nothing about why.** The ceilings and quotas that shape the load, and what the node does once it genuinely cannot keep up, are separate subjects with their own signals. ## What varies across platforms - Many self-run platforms expose an occupancy or idle-share number per pool, and some expose more than one pool with different constraints. - Where the pool size is fixed at process start, an operator can raise it and restart; where it adapts, saturation may resolve on its own; on rented clusters the size is usually not yours to change and the number may not be published at all. - When a tenant cannot see node occupancy, the nearest honest substitute is the client-side duration distribution plus the request rate — you lose the lead time and read the cliff after it starts. The interview point is the shape, not the number: you act on 85% because the curve between occupancy and waiting is steep at the top, and by the time the latency panel agrees with you, clients have already noticed.
- Handler-pool saturation is at 90% but request-queue depth has never left zero. What does that pair mean?The pool is busy but not short: every arriving request still finds a handler promptly, so nothing is backing up. That is a node working hard at a sustainable rate, not one in trouble. It is worth watching because the slack left is thin, but acting as though clients are queueing would be reading occupancy as if it were a backlog.
- Why is there no single occupancy percentage that means danger on every cluster?The occupancy at which waiting takes off depends on how bursty arrivals are and how variable the cost per request is. A steady, uniform workload runs comfortably much fuller than a spiky one, because slack exists to absorb bursts. The honest approach is a threshold derived from that node's own observed relationship between occupancy and waiting, not a number copied between clusters.
- When does this signal give you no warning at all?When the handling itself gets slower rather than the arrivals increasing. Each request occupies a handler for longer, so occupancy and duration rise together and neither leads. The tell is service time moving while the request rate is flat, which is why the split of a request's duration is read alongside the occupancy number rather than after it.
A road carries traffic at full speed while it has room, and the same extra cars that changed nothing at half capacity produce stop-start queues once it is nearly full. Occupancy has been climbing the whole time; the journey time only steepens at the end.
saying these in an interview costs you the question
- Waits for latency to move before acting on a filling handler pool
- Quotes a universal saturation percentage as the danger line
- Reads high occupancy as a backlog without checking queue depth
- Assumes saturation always gives warning, including when handling slows
- Watches a pool that is not the node's constrained resource