skip to content

Several pipelines fan out to one shared classification service: how would you decide the concurrency bound each flattening step uses?

level: principalimportance: should knowfreq 34%

answer

  1. a capacity decision, not a style
  2. budget the dependency, then divide
  3. multiply by the instance count
  4. rate times latency gives in-flight
  5. isolation costs utilisation

basics

~20 s

Size it from the dependency's safe concurrency rather than each pipeline's appetite: divide one fleet-wide budget across callers with headroom, check it against arrival rate times call latency, and keep bounds per pipeline so a single backfill cannot consume the service.

solid answer

~50 s

Treat the bound as a capacity decision, not a style choice: it is the peak number of in-flight calls one pipeline places on a shared dependency, multiplied by however many instances of that pipeline run. Start from what the dependency can serve inside its latency objective, keep headroom, and divide that budget across callers rather than letting each pick a number - two pipelines each sized to the whole capacity will together exceed it, and rising latency raises in-flight time, which raises concurrency further at the same arrival rate. Sanity-check each share against Little's Law: arrival rate times call latency is the in-flight count needed to keep up, so a share below that makes the step the bottleneck and lag grows without limit. Then decide where the limit lives and whether pipelines get separate bounds, which isolates a backfill from live traffic at the cost of utilisation. Finally make it configurable and revisit it when latency moves.

go deeper

for a junior

Recall that the bound is the number of calls the step can have in flight at once, and that leaving it unstated or unbounded means a burst of arrivals becomes a burst of simultaneous calls.

for a middle

Explain how to estimate it: arrival rate times call latency gives the in-flight count needed to keep up, and the bound also caps the load this step places on the service it calls.

for a senior

Demonstrate that you measure the dependency's latency against concurrency and set the bound below the knee, and that you account for instance count and for retries inside the call.

for a principal

Own the budget rather than the number: publish the dependency's safe concurrency, divide it across callers by importance, choose isolation against utilisation deliberately, and make every step's share explicit and reviewable.

## The bound is a capacity decision A flattening step's concurrency bound looks like a local tuning knob and is not. It is the maximum number of simultaneous calls this step places on a shared dependency, and the number that matters is not the one in the code: it is that bound multiplied by the number of running instances of the pipeline, summed over every pipeline that calls the same service. A bound of 16 in a step running on twelve instances is 192 concurrent calls from one caller. Teams discover this during an incident, not during a review. So the decision is made in the opposite direction from the way it is usually approached. Not *how much concurrency does my pipeline want*, but *how much concurrency can the dependency serve, and how is that budget divided*. ## Two numbers to reconcile **The budget (top-down).** Measure the dependency's latency against offered concurrency and find the knee - the point past which latency rises faster than throughput. Set the fleet-wide budget below that knee with headroom for bursts and for whatever is not in the budget yet, then divide it among callers by importance rather than equally. **The need (bottom-up).** Little's Law gives the other number: the items in flight equal arrival rate times the time each spends in the system, so **arrival rate x inner-call latency** is the concurrency required just to keep up. A queue receiving 40 items a second against a 250 ms call needs 10 calls in flight to break even. A share below that means the step is the permanent bottleneck and lag grows linearly forever; a share far above it buys nothing because the arrivals are not there. When the two numbers disagree, that is the real finding - the demand does not fit the dependency - and the answer is a capacity conversation, a cheaper call, or fewer calls, not a larger bound. ## Where the limit lives | placement | what it isolates | utilisation | symptom when it trips | |---|---|---|---| | per flattening step | that one pipeline | lowest - each share sits idle when unused | that pipeline's lag grows, visible in its own metrics | | shared permit pool per process | all pipelines in the process from the outside world | higher - shares are pooled | whichever pipeline asks next waits, so blame is unclear | | at the dependency | nothing on the caller side | highest | failures appear across all callers at once | The useful principle is that you want the limit **you own** to trip first. A bound in your own step fails in a place you can observe, with one symptom you can attribute, and a number you can change. A limit enforced at the dependency arrives as failures spread across every caller and tells you least about which one caused them. Having both is fine, provided yours is the tighter of the two. ## Isolation against utilisation Separate per-pipeline bounds are the isolation choice: a batch backfill cannot consume the share that live traffic depends on, and each pipeline's behaviour is legible on its own. The cost is idle capacity, since one pipeline's unused share is not available to another. A single pooled bound is the utilisation choice and gives that back, at the price of a noisy neighbour being able to take the whole pool. For a shared classifier serving both interactive and batch work, the defensible default is separate bounds with the batch share deliberately small, because the failure you are protecting against is the backfill, and it is the one you can predict. ## What the bound does not do It does not reject work. When the bound is reached the step simply stops subscribing to further inner sources, which in a demand-driven pipeline means it stops requesting elements and the pressure propagates back upstream. Whether excess work should be refused, shed or delayed is a separate decision made elsewhere, and conflating the two leads to a bound set as though it were an admission policy. It also does not protect you from retries. A retry inside the inner call multiplies effective concurrency by the retry factor at exactly the moment the dependency is struggling, so the bound must be sized against the retried volume, not the nominal one. ## Making it a standard rather than a number What a lead owns here is the process, not the integer: 1. Publish the dependency's budget and who holds which share, so the number in any one repository is traceable to a decision. 2. Require every flattening step that calls a shared service to state its bound explicitly, so nothing inherits an implicit default. 3. Make the bound configurable without a deploy, since it must move when latency moves. 4. Alert on in-flight count pinned at the bound, which is the signal that the share is now the bottleneck, and re-derive from Little's Law when it fires. 5. Re-divide the budget whenever a caller is added, because the sum is the thing that must hold, not any individual share.

  • What symptom says the bound is too low?
    In-flight count pinned at the bound while the dependency sits well under its capacity, with lag growing linearly as arrivals continue. The step, not the service, is the bottleneck. The fix is a larger share of the budget, not more instances, which would multiply the load instead of relieving it.
  • Two pipelines were each sized to the dependency's full capacity - what happens?
    Together they exceed it. Latency rises, which lengthens the time each call stays in flight, which raises concurrency further at the same arrival rate; retries inside the calls compound it. The sizing has to come out of one shared budget, or the sum is nobody's number.
  • Where would you rather the limit trip: in your pipeline or at the dependency?
    In the pipeline. A bound you own fails in a place you can observe and change, with one attributable symptom. A limit enforced at the dependency surfaces as failures spread across every caller at once, which tells you least about which caller caused them.

saying these in an interview costs you the question

  • Picks a round number with no reference to the dependency's capacity
  • Sizes each pipeline to the dependency's full capacity
  • Forgets the bound multiplies by the number of running instances
  • Leaves the fan-out unbounded because the burst is rare
  • Thinks a higher bound always raises throughput
  • Confuses limiting in-flight calls with rejecting arriving work