A nightly job fans out tens of thousands of HTTP calls to one downstream service through a concurrency limiter. How do you decide what the limit should be, and why is a hardcoded constant like 10 usually the wrong long-term answer?
answer
- a property of the dependency
- throughput times latency
- per-process caps multiply by replicas
- watch for the throughput knee
- the results array is the other leak
basics
~20 sDerive the limit from the dependency's budget, not from your loop: target throughput times average latency gives the in-flight count you need. A per-process constant is wrong because the downstream sees every replica's limiter multiplied together, and because the right number drifts as latency and capacity change.
solid answer
~50 sI start from the downstream rather than from the code. If each call takes about 200 ms and my agreed share of the service is 50 requests per second, then by Little's law I need roughly 10 in flight — throughput times latency, not a guess. Then I check the other ceilings: any published rate limit, the socket and connection budget on my side, and the memory held per in-flight call plus whatever I am accumulating in the results array. The reason a constant rots is that it is per-process: eight replicas each capped at 10 present 80 to a service that only ever agreed to 10, so the meaningful bound is global and belongs to the client for that dependency, with each job drawing from it. It also drifts — when latency doubles, a fixed limit silently halves your throughput. I validate by raising the limit until throughput flattens and latency climbs; that knee is saturation, and past it extra concurrency only buys queue time.
go deeper
Know that a concurrency limit is chosen from what the downstream can take, not from what makes your loop finish fastest, and that raising it past the service's capacity makes everything slower rather than faster.
Be able to compute a starting point from throughput and latency, and to list the other ceilings — rate limits, connections and sockets, memory per in-flight response. Explain why a doubling of latency changes what a fixed limit delivers.
Show how you would validate the number empirically: raise it until throughput flattens and p99 climbs, and set the limit below that knee. Point out that the results accumulator, not the concurrency, is often the real memory bound.
Own the aggregate: per-process constants multiply by replica count, so argue for where enforcement truly belongs — configurable budgets, per-dependency bulkheads, gateway or server-side admission control — and say when an adaptive controller earns its complexity versus a tuned constant.
## The limit belongs to the dependency, not to the loop The most common mistake is treating the concurrency cap as a property of the code that happens to be looping. It is not — it is a statement about how much of a shared resource you are entitled to consume. That reframing produces most of the right answers on its own: the number lives with the client for that dependency, it is one number per dependency rather than one per job, and it has to be reasoned about across every process that uses it. ## Start with arithmetic, not a guess Little's law gives the starting point: in-flight items = throughput × latency. If you want 50 requests per second and each takes 200 ms, you need about 10 concurrent. If latency is 2 s, the same throughput needs 100. Two things follow immediately. First, the limit and the latency are coupled: when the downstream slows down, a fixed limit means your throughput falls proportionally — which is arguably correct behaviour, since it is backpressure, but you should know it is happening rather than discover it as a missed batch window. Second, if you have a throughput target and a latency measurement, you never have to guess; you have to *decide* a target and then compute. ## Then check the other ceilings - **Agreed rate.** A published quota or a negotiated share is a hard ceiling; your computed number must sit under it. - **Connections.** Each in-flight call holds a socket, and in a browser or behind a proxy the host may cap connections per origin and simply queue the rest. Concurrency above that ceiling buys you queued work and memory, not parallelism. - **Memory per item.** Response bodies, parsed objects and the closure state of each pending call are all live at once. 200 concurrent 5 MB responses is a gigabyte you did not budget. - **The accumulator.** This is the sneaky one and it is not bounded by your limit at all. A pool that returns one big results array holds every result for the entire run. For 200,000 items that array, not the concurrency, is your memory problem — write each result out inside the worker and return nothing. ## Why the per-process constant rots A constant in the source is a per-process bound. The downstream experiences the sum: eight replicas, or a job that got parallelised by shard, present 8× what the author intended, and nobody edited a line to cause it. Autoscaling makes this worse, because the multiplier is now dynamic — precisely when the system is busiest, the effective fan-out grows. The practical responses, in increasing order of investment: make the limit configurable rather than literal so it can be tuned without a deploy; divide a global budget by the known replica count; or move enforcement to a place that sees all traffic — a shared gateway or the downstream's own admission control. The last is the only one that is actually correct, since only the receiver knows the true aggregate; client-side limits are cooperative, and a cooperative limit is a good citizen rather than a guarantee. ## Separate limits per dependency One global limiter shared across unrelated downstreams couples them: a slow dependency occupies all the slots and starves calls to a fast, healthy one. Give each dependency its own limiter sized to its own budget. This is the same reasoning behind bulkheads — isolation of one dependency's degradation from another's — and it is why "we have a limit of 50" is an incomplete sentence without naming what it is 50 of. ## Measuring rather than asserting Run the fan-out at increasing limits and plot throughput and p99 latency together. Below saturation, throughput rises roughly linearly and latency is flat. At the knee, throughput flattens and latency starts climbing — that is queueing, and every unit of concurrency past it is added latency with no added work done. Set the limit at or slightly below the knee. If throughput actually *falls* past the knee, the downstream is in congestion collapse and you are the cause. ## Adaptive limits For long-lived services the mature answer is not a number at all but a controller: raise the limit slowly while latency and error rate stay healthy, cut it sharply when they degrade — additive increase, multiplicative decrease. It tracks the dependency's real, varying capacity instead of your estimate of it on the day you wrote the code, and it degrades gracefully during an incident. The cost is a feedback loop to tune and reason about, so it pays off for a hot path and is over-engineering for a nightly job. ## What to say Compute a starting number, name the ceilings you checked, explain that the bound is global rather than per-process, and say how you would validate it by finding the knee. Ending with "and I would make it configurable and per-dependency" is worth more than any specific number.
- How would you make the limit adapt instead of pinning it?Run a controller over the limiter's slot count: increase it gradually while latency and error rate stay within targets, and cut it by a multiplicative factor on degradation. It tracks the dependency's real capacity as that changes, and it sheds load automatically during an incident. The cost is a feedback loop with its own tuning and failure modes, so reserve it for hot paths.
- One job calls three different services. One limiter or three?Three. A single shared limiter couples them: if one dependency slows to seconds per call, its calls occupy every slot and starve the healthy ones. Separate limiters sized to each dependency's own budget keep one service's degradation from becoming an outage across all three — the bulkhead argument.
- The job holds steady at 10 concurrent yet still runs out of memory. What are you missing?Almost certainly the accumulator rather than the concurrency. A pool that resolves to one array of every result holds the entire output for the whole run, which grows with input length and ignores the limit completely. Write each result to its destination inside the worker and return nothing, so live memory tracks the concurrency instead of the input size.
saying these in an interview costs you the question
- Picks a round number with no reference to latency or quota
- Assumes a per-process cap is what the downstream experiences
- Raises the limit when latency climbs, mistaking queueing for slack
- Uses one global limiter across unrelated dependencies
- Believes the concurrency cap also bounds total memory