skip to content

How would you decide which endpoints of a synchronous service are worth converting to awaitable-returning handlers, and which are not?

level: principalimportance: should knowfreq 42%

answer

  1. capacity, not per-request speed
  2. wait fraction times concurrency ranks it
  3. compute-bound endpoints gain nothing
  4. a hard downstream ceiling just moves the queue
  5. convert slices, with a stopping rule

basics

~20 s

Convert endpoints that hold many concurrent requests whose time is mostly spent waiting on remote calls that can genuinely be awaited. Leave compute-bound endpoints, and those fronting a hard downstream limit, alone: converting them only relocates the queue.

solid answer

~50 s

Start from measurement, not style. For each endpoint, compare time spent waiting with time spent computing, and multiply by concurrency: the candidates are the ones holding many requests that are mostly idle. Then check feasibility — if every dependency on that path is blocking, conversion buys you an offload pool rather than genuine non-blocking waiting, and the downstream's own ceiling still caps you, so the queue simply moves. Weigh the cost of a mixed codebase honestly: two wrapper contracts, two error paths, context that must be carried explicitly, and a single blocking call left on a converted path erasing the benefit. I would convert whole slices — handler, its wrappers, its clients — for the highest waiting-times-concurrency endpoints, set a measurable exit criterion, and stop converting when the next slice's benefit no longer covers the ongoing cost of carrying two styles.

go deeper

for a junior

Take away the framing: the change buys the ability to hold more requests at once, not faster responses for any one of them.

for a middle

Be able to say which endpoint profiles benefit — mostly waiting, high concurrency, awaitable dependencies — and which gain nothing, such as compute-bound work.

for a senior

Show that you would convert whole slices and verify the result operationally, including checking that no blocking call has crept back onto a converted path.

for a principal

Own the tradeoff end to end: rank by measured benefit, price the cost of carrying two styles, state a stopping rule, and compare against scaling out, isolating, or shedding load before committing the organisation to a migration.

## Frame it as capacity, not modernisation An awaitable-returning handler buys one thing: the process can hold more requests in flight for the same resources, because a waiting request is not pinning one. It does not shorten the waiting any individual request has to do. So the question is never whether the shape is better; it is whether an endpoint's requests spend enough time waiting, at enough concurrency, for that trade to pay for the complexity it imports. ## The ranking inputs For each endpoint, gather four numbers before arguing about any of them: 1. **Wait fraction** — share of request duration spent on remote calls rather than on the CPU. A high fraction is the precondition; without it there is nothing to reclaim. 2. **Concurrency** — requests in flight, not requests per second. A slow endpoint at low volume ties up little; a moderately slow one at high volume dominates. 3. **Dependency awaitability** — whether the clients on that path can genuinely wait without holding a resource, or would have to be offloaded to a pool. 4. **Blast radius** — how many other endpoints suffer today when this one is slow, which is what the extra headroom actually protects. | endpoint profile | convert? | reasoning | |---|---|---| | high wait fraction, high concurrency, awaitable clients | yes, first | the whole benefit, at full strength | | high wait fraction, low concurrency | later | real but small; not worth leading with | | compute-bound | no | cores are the limit, and the shape does not add any | | fronts a dependency with a hard concurrency ceiling | no | admitting more only lengthens the queue in front of it | | every client on the path is blocking | rarely | you get an offload pool, with the pool's bound as the new ceiling | ## The costs a decision has to price in - **Two wrapper contracts.** Stages written for the synchronous chain behave wrongly on the awaitable one, so shared wrapper code either forks or is rewritten to compose. - **Two error paths.** A failure can arrive by a throw or by a settled handle; mappings that cover only one drift apart across a surface. - **Context by hand.** Per-request values kept in ambient storage may need to be carried explicitly across wait points, which touches logging, identity and transaction handling. - **Fragility of the result.** One blocking call left on a converted path can consume the very resource the conversion freed, so converted endpoints need a standing check that nothing blocking has crept back in. - **Knowledge cost.** Every future change on that path has to be made in the style the path uses, by whoever is on call at three in the morning. ## Convert slices, not signatures A converted handler whose wrappers still run their after-work inline, or whose clients still block, is converted in name only — usually with instrumentation that has quietly started lying. The unit of conversion is the slice: handler, every wrapper on its path, and the clients it calls. That makes conversion lumpier and slower than a signature sweep, and it is the reason to rank ruthlessly rather than convert everything. ## Stopping rules and alternatives Define the exit criterion up front — for example, sustained in-flight requests at the target load fitting inside the resource budget with headroom to spare — and stop when the next slice cannot show it. Consider the alternatives honestly before starting: - **Scale out.** More instances add capacity with no code change. It costs money instead of complexity; sometimes that is the right trade, especially for a service with a short remaining life. - **Isolate instead.** Separate pools, or a separate deployment for the endpoints that starve the others, can remove the coupling that motivated the conversion. - **Shed load.** A concurrency limit with a clear busy response protects the system better than a higher in-flight ceiling that just fails deeper. - **Wait for the platform.** Where lightweight execution contexts make blocking-style code cheap, the same headroom can arrive without changing the handler shape at all. Whether that is available is a platform question worth asking before committing. The defensible answer names the numbers it would gather, converts a small ranked set of slices end to end, measures against a stated criterion, and is explicit that a permanently mixed codebase is an accepted cost rather than an oversight.

  • An endpoint waits on a dependency that accepts only twenty concurrent calls. Is it a conversion candidate?
    Not for capacity. Converting lets the service admit far more requests, but they queue in front of the same ceiling, so latency rises while throughput does not. The useful change there is an explicit concurrency limit with a busy response, so the queue is bounded and visible, rather than a handler shape that hides it.
  • How would you stop a half-converted codebase from becoming permanent accidental debt?
    Make the split deliberate and legible: decide which endpoints live in which style, keep each style's wrapper chain separate, and record the decision where the next engineer will find it. Then set a stopping rule in advance so the remaining synchronous endpoints are a stated choice rather than the ones nobody got to.
  • What measurement tells you a completed conversion actually worked?
    Sustained in-flight requests at target load against the resource budget, plus the tail latency of unrelated endpoints during a slow-dependency episode. The second is the real prize: the conversion succeeded if one slow dependency no longer drags everything else, not if a dashboard shows faster individual responses.

saying these in an interview costs you the question

  • Converts everything on principle rather than by measurement
  • Expects individual requests to get faster after conversion
  • Ignores that a blocking dependency caps the whole gain
  • Converts handler signatures without their wrappers or clients
  • Treats a permanently mixed codebase as free
  • Never considers scaling out or shedding load as alternatives