In a load test, your service shows CPU around 45% across the fleet, yet p99 latency has tripled and total throughput has stopped rising. Which saturation signals do you look at, and what would each one tell you?
answer
- utilisation is not saturation
- look for waiting, not busy
- queue depth and pool wait time
- averages hide one pinned instance
basics
~20 sUtilisation is not saturation. Look for evidence of waiting rather than of being busy: request queue depth and time-in-queue, thread-pool and connection-pool wait, downstream call latency, lock contention, and per-instance skew. Each names a different bottleneck, and each implies a different fix.
solid answer
~50 sCPU tells you how busy one resource is; saturation is about what work is *waiting*, and something is clearly waiting here. I would walk the request's path looking for queues. At the edge: accept-queue depth and time between arrival and handler start. Inside the process: thread-pool active-versus-max and its task queue, plus connection-pool checkout wait time for the database and for HTTP clients — pool exhaustion produces exactly this signature, since threads block without burning CPU. Then downstream: if dependency latency rose in step with mine, I am not saturated at all, my dependency is. Then per-instance and per-core breakdowns, because a fleet average of 45% can hide one pinned instance or a single-threaded hot path. Runtime pauses and lock contention round it out. The distinction matters operationally: if the bottleneck is a shared database or a connection pool, adding replicas makes it worse, not better.
go deeper
Understand that a service can be at its limit while CPU looks low, because threads waiting on a database connection or a slow dependency are blocked rather than busy.
Be able to name the concrete signals — queue depth, queue wait, pool checkout time, in-flight count, dependency latency — and explain what each one distinguishes. Know why an exhausted connection pool caps throughput at pool size over hold time.
Demonstrate the diagnostic walk under time pressure and, crucially, tie each finding to a different remedy: scaling out helps for per-instance limits and actively hurts when the constraint is shared. Bring a real example where the fleet looked idle.
Own the requirement that these signals exist before an incident: a platform-wide standard for pool, queue, and in-flight instrumentation, and the argument for why utilisation dashboards alone leave every team unable to answer this question when it counts.
## Utilisation is not saturation These are two different measurements and conflating them is the classic capacity-testing mistake. **Utilisation** is the fraction of time a resource is busy. **Saturation** is the amount of work that is queued because a resource was not available. A service can be fully saturated at 45% CPU if the thing it is short of is not CPU — and in practice, for request-serving services, it usually is not. Threads blocked on a connection, on a lock, or on a downstream response are not consuming CPU while they wait; they are the bottleneck and they are invisible in the utilisation figure. The symptom described — latency rising while throughput has stopped — is the signature of a queue growing somewhere. The job is to find which queue. ## Walk the request path and look for queues **At the edge.** Listen-backlog / accept-queue depth, and the time between a connection being accepted and the handler starting work. If requests are already old before your code sees them, the constraint is upstream of your handler: not enough worker threads, an event loop blocked by synchronous work, or a proxy holding connections. **Thread or worker pools.** Active threads versus maximum, and the depth and age of the pending task queue. A pool at max with a growing queue is a saturated pool, whatever the CPU says. Note the age of the oldest queued item: it converts directly into added latency for every request behind it. **Connection pools.** Time to check out a database or HTTP connection is one of the highest-value saturation signals in a service, and one of the least often instrumented. Pool exhaustion produces this exact pattern: throughput pinned at (pool size ÷ per-query time), latency climbing linearly with offered load, and CPU low because everyone is blocked. **Downstream calls.** Compare your latency increase to your dependency's latency increase. If they rose together, you are not the saturated component — a database, cache, or third-party API is, and testing harder against your own service just measures theirs. **Runtime and lock effects.** Garbage-collection pause time as a fraction of wall clock, safepoint pauses, and lock contention or single-writer serialisation. A single global lock caps throughput at 1/critical-section-time regardless of core count, and shows as low CPU with high wait. **Below the process.** Disk write latency and queue depth for anything doing durable writes; network throughput against the instance's actual bandwidth allocation, which on cloud instances is often far below the theoretical NIC rate and may be credit-based and burstable, so it degrades partway into a sustained test. ## Averages lie in two directions A fleet average of 45% can hide an instance at 100%. Check the distribution across instances: uneven hashing, session affinity, a hot shard, or one host with a noisy neighbour will saturate a single instance while the fleet looks comfortable, and if that instance serves a slice of users, their p99 is the p99 you are seeing. The same applies within a machine. A workload with one hot single-threaded path pins one core; on a 16-core box that reads as roughly 6% aggregate CPU while that path is completely saturated. Always look per-core and per-thread before concluding there is CPU headroom. ## Why the distinction changes what you do This is not diagnostic pedantry — the remedies are opposite: - Saturated on **per-instance CPU or memory**: add replicas, or make the code cheaper. Scaling out works. - Saturated on a **shared downstream** (one database, one cache cluster): adding replicas increases concurrent load on the shared resource and makes latency *worse*. The fix is at the shared tier — caching, query cost, read replicas, sharding — or reducing per-request calls. - Saturated on a **connection pool**: raising pool size only helps if the downstream can absorb the extra concurrency. Otherwise you have moved the queue from your process into the database, where it is harder to see and more dangerous. - Saturated on a **lock or single-threaded stage**: neither more instances nor bigger instances help. Only removing the serialisation does. ## Instrumenting for this before you need it The reason teams cannot answer this question during a real test is that the signals were never exported. The short list worth adding to any service: queue depth and queue wait time for every internal pool, connection-pool checkout wait and exhaustion count, in-flight request count, per-dependency latency and error rate, and runtime pause time. In-flight count is especially useful because, via Little's Law, in-flight = arrival rate × latency, so a rising in-flight count with a flat arrival rate *is* the latency problem, seen from the resource side. ## What weak answers look like Staring at aggregate CPU and concluding there is headroom. Adding instances reflexively when the constraint is shared. Reading utilisation as a proxy for how close to the limit you are, when the actual distance to the cliff depends entirely on which queue is filling and how fast.
- What throughput ceiling does an exhausted connection pool impose, and how would you compute it?Roughly pool size divided by the average time a connection is held. Twenty connections held for 50 ms each caps you at about 400 requests per second, no matter how many instances or cores you add. The signature is flat throughput at exactly that number, climbing checkout wait time, and low CPU. Raising the pool only helps if the database can absorb the extra concurrency.
- Your fleet averages 45% CPU. What would make you suspect that number is hiding the real picture?Any signal that load is not evenly spread: a wide spread of per-instance latency, one instance with a much higher request rate or error rate, session affinity or a hot partition key, or a single hot core on an otherwise idle machine. A per-instance and per-core breakdown resolves it in seconds and stops you from scaling out a fleet that is already mostly idle.
- Why is in-flight request count a particularly useful saturation signal?Because Little's Law ties it directly to latency: in-flight equals arrival rate times response time. With a flat arrival rate, a rising in-flight count is the latency increase seen from the resource side — and it is also the number that predicts resource exhaustion, since each in-flight request holds a thread, a connection, and memory.
saying these in an interview costs you the question
- Reads low CPU as proof of remaining headroom
- Adds replicas when the bottleneck is a shared database
- Never instruments connection-pool checkout wait time
- Trusts a fleet average over the per-instance distribution
- Assumes saturation always shows up as high utilisation