Which measurements taken before an asynchronous rewrite would confirm or kill a claimed capacity gain?
answer
- find the binding constraint first
- waiting time versus computing time
- in flight equals rate times duration
- occupancy and utilisation read together
- saturated processors kill the claim
basics
~20 sMeasure where a request's time goes, how many requests are in flight at peak, what the machine runs out of first, and the memory each connection holds. Parked workers beside idle processors confirm the claim; saturated processors kill it.
solid answer
~40 sThe claim is that the same hardware will carry far more concurrent work, so the evidence must show that **workers**, not processors, are what runs out. Four readings decide it. Split a request's wall-clock time into waiting and computing. Take requests in flight at peak — arrival rate times duration, by Little's Law. Read worker occupancy and processor utilisation *together*, because a pool that is fully occupied while the processors idle means workers are parked, and a pool fully occupied while processors saturate means the machine is simply busy. Finally, measure memory per in-flight request against the deployment's ceiling. A fifth reading, the share of dependencies with a genuine non-blocking path, caps whatever the first four promise.
go deeper
Learn the shape of the argument: before changing how a service is built, find out what it actually runs out of first. Workers, processors and memory fail differently.
Be able to derive requests in flight from arrival rate and duration, and to explain why an occupied worker and a busy processor are not the same thing.
Propose the cheap experiment before the migration plan: one instrumented build, one realistic load test, and the two counters read together. Name what would make you abandon the rewrite.
Insist that the capacity claim is written down with the measurement that would falsify it, so the decision can be reviewed later on evidence rather than on whoever argued hardest.
## What the claim actually says "Going asynchronous will let this service carry far more concurrent work on the same hardware" is a claim about **which resource runs out first**. It is true only if the binding constraint is units of execution held during waits, and it is false if the binding constraint is anything else — processor time, memory that the new design also holds, or a downstream that is itself the limit. So the measurements are chosen to identify the binding constraint, not to describe the service in general. ## The four readings, and what each one settles 1. **Where a request's wall-clock time goes.** Split the median and the tail into time spent waiting on network, disk or another service, and time spent computing locally. If waiting dominates, there is parked capacity to reclaim. If computing dominates, there is none, and the rewrite is answering a question the service is not asking. 2. **Requests in flight at peak.** By **Little's Law**, this is the arrival rate multiplied by average duration — not something you need to instrument directly. Single digits means no pressure on the pool at all. Thousands means the pool is the thing under pressure. 3. **Worker occupancy and processor utilisation, read side by side.** This pair is the single most diagnostic reading, and either number alone misleads: | Occupancy | Utilisation | What it means | |---|---|---| | High | Low | Workers parked waiting — the case the style is for | | High | High | The machine is genuinely busy; the style gains nothing | | Low | High | A few paths are compute-heavy; look at them, not at the model | | Low | Low | Nothing is constrained; the limit is somewhere downstream | 4. **Memory per in-flight request against the ceiling.** Where each request or connection holds an execution stack, footprint scales with concurrency and often binds before anything else, especially with many long-lived, mostly idle connections. If the deployment's memory ceiling is far away, this argument is not available either. ## The reading that caps all of them A fifth measurement bounds whatever the first four promise: the **share of dependencies reachable without parking a worker**, weighted by traffic. The capacity gain applies only to paths that release their worker at every wait. Compute the projected gain using the converted share, never the whole. ## What kills the claim outright Any one of these ends the discussion, and it is worth saying so plainly rather than negotiating: - **Processors saturate before the pool does.** The work is compute-bound. Buy capacity, make the computation cheaper, or shed load. - **Requests in flight stay small at peak.** There is no concurrency pressure to relieve, whatever the traffic graph looks like in aggregate. - **The tail is owned by a downstream.** If the slow percentile is a dependency taking a long time rather than requests queueing for a free worker, releasing workers changes when the request is admitted, not when the answer arrives. - **Most dependencies cannot be converted.** The projection collapses to the converted share. ## The reading that is easiest to misinterpret Worker-pool saturation on its own. A fully occupied pool feels like the smoking gun, and teams have rewritten services on the strength of it. Occupied is not the same as busy: a worker waiting on a socket is occupied and doing nothing. Read utilisation next to occupancy, every time, and interpret only the pair. A second trap is testing the claim on a synthetic load whose downstream is a stub that answers instantly. That workload has no waiting in it, so it measures the framework's overhead rather than the effect under study — and it usually flatters the conventional build. ## Doing it cheaply None of this needs the rewrite to exist. One instrumented build of the current service, one load test that reproduces peak arrival with realistic downstream latency, and an hour with the two counters gives you the answer. The asymmetry is the point: measuring costs days and is reversible, while rewriting costs months and is not. A candidate who proposes the measurement before the migration plan is demonstrating the judgment this question is asked to find.
- A load test shows processors near ninety percent at peak. What does that settle?That the work is compute-bound, so workers are not the resource running out and there is no parked capacity to reclaim. Releasing workers during waits cannot help when there is little waiting. The options are cheaper computation, more hardware, or shedding load — and hand-offs between workers would add a small cost on top.
- Which of these measurements is most often misread, and how?Worker-pool saturation on its own. A fully occupied pool looks like proof, but occupied is not busy: a worker waiting on a socket is both. Only the pair decides it — occupancy high with utilisation low means parked workers, while both high means the machine is genuinely at its limit and the rewrite would gain nothing.
saying these in an interview costs you the question
- Reads worker-pool saturation without checking processor utilisation.
- Projects the capacity gain across paths that were never converted.
- Load-tests against a stub downstream that answers instantly.
- Expects the rewrite to improve a tail owned by a slow dependency.
- Argues from the traffic graph instead of requests in flight.