To cap concurrency at 10 over 1,000 jobs you can either slice the list into sequential batches of 10 and await `Promise.all` on each batch, or run 10 workers that pull from a shared queue. What does the batching version cost you?
answer
- Promise.all per batch is a barrier
- slowest job sets the batch cost
- idle slots waiting on stragglers
- pool refills a slot immediately
- batches buy clean checkpoints
basics
~20 sEach batch is a barrier: it takes as long as its slowest job while the other nine slots sit idle. With uneven durations that wastes most of the capacity, whereas a pool refills a slot the instant one frees. Batching is simpler and gives clean checkpoints.
solid answer
~50 sBatching enforces the same peak concurrency but not the same utilisation. `await Promise.all(batch.map(...))` is a barrier: nothing from batch two starts until every job in batch one is done, so a single 2-second straggler among nine 50-millisecond jobs leaves nine slots idle for almost 2 seconds, and the total run becomes batch-count times worst-case latency instead of total-work divided by 10. A pool has no barrier — a runner claims the next index the moment its own item settles — so it stays saturated and finishes far closer to the theoretical floor whenever latencies are heavy-tailed. Batching is still a reasonable choice: it is five lines, it gives a natural checkpoint boundary for committing or logging progress, and it is the right shape when the downstream genuinely takes a batch per call. I default to a pool for latency-variable I/O and reach for chunks when the batch boundary means something.
code
javascript · 23 linesconst sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function inChunks(items, size, worker) {
const out = [];
for (let i = 0; i < items.length; i += size) {
const chunk = items.slice(i, i + size);
// Note: (item) => worker(item), not chunk.map(worker) - that index is chunk-local.
const settled = await Promise.all(chunk.map((item) => worker(item)));
out.push(...settled);
}
return out;
}
(async () => {
const durations = [500, 20, 20, 20, 500, 20, 20, 20];
const started = Date.now();
await inChunks(durations, 4, async (ms) => {
await sleep(ms);
return ms;
});
// ~1000 ms: two batches, each paying its own 500 ms straggler.
console.log('batched ms:', Date.now() - started);
})();go deeper
Know that awaiting Promise.all on a batch waits for the slowest job in that batch, so the next batch cannot start until then. Recognise that a worker pool starts the next item as soon as any single one finishes.
Explain the barrier and estimate its cost: batch time equals worst-case latency, so idle slot-time is the batch count times the gap between the slowest and the average job. Note that the two shapes converge when durations are uniform.
Demonstrate judgment about latency distributions — argue from p99 versus median, quantify the throughput gap, and name the cases where the barrier is the feature: bulk endpoints, checkpointing, resumable prefixes, and memory flushes.
Frame it as a resumability and operability decision as much as a throughput one: prefix-shaped progress is cheap to restart, scattered progress is not. Be ready to defend a hybrid — pooled segments with a barrier at a coarse checkpoint boundary.
## Two shapes that look equivalent ```js // A: sequential batches for (let i = 0; i < items.length; i += size) { const chunk = items.slice(i, i + size); out.push(...(await Promise.all(chunk.map((item) => worker(item))))); } // B: N runners over a shared cursor // each runner: while (next < items.length) { const i = next++; out[i] = await worker(items[i]); } ``` Both cap in-flight work at `size` / N. Both preserve input order. They differ in one respect that matters enormously under real latency distributions: **A has a barrier and B does not.** ## The barrier cost, with numbers Suppose 1,000 jobs, batch size 10, and a latency distribution where nine jobs in ten take 50 ms and one in ten takes 2 s. Each batch of 10 therefore contains, on average, one straggler. - **Batched:** `Promise.all` resolves when the *last* member settles, so each batch costs about 2 s. One hundred batches ≈ **200 s**. For roughly 1.9 s of each batch, nine of your ten slots are doing nothing. - **Pooled:** total work is 900 × 50 ms + 100 × 2 s = 45 s + 200 s = 245 s of work spread over 10 always-busy slots ≈ **24.5 s**. An order of magnitude, from a change that does not touch the concurrency limit at all. The general rule: batching costs you `batches × (max latency − mean latency)` of idle slot-time. When durations are uniform the two converge, which is exactly why the problem hides in tests and appears in production, where p99 is many multiples of the median. ## Where batching is genuinely right Do not read this as "chunking is wrong". It wins when: - **The downstream takes batches.** If the API is `POST /users/bulk` with up to 100 ids per call, the chunk *is* the unit of work, and there is no straggler problem because there is one call per chunk. - **You need a checkpoint.** Writing a resume cursor, committing a transaction, or flushing to disk after every 500 items is far easier at a barrier, because at that instant nothing is in flight and the completed prefix is exactly known. - **You want to bound accumulated memory, not just concurrency.** If each result is large, a batch loop lets you write out and discard after every chunk. (A pool can too, but you have to arrange it deliberately.) - **You are pacing deliberately.** A burst-then-pause shape is sometimes what a fragile downstream tolerates best. ## Failure and progress semantics With batching, a rejection settles that batch's `Promise.all` and — if you do not catch it — breaks out of the loop, so everything before the current batch is done and everything after it is untouched. That prefix property is genuinely useful for resumability. With a pool, failure finds you mid-flight across an arbitrary set of indexes, so "what completed" is a scattered set rather than a prefix, and a resume story has to track completions individually. ## The hybrid You can usually have both. Keep the pool for utilisation and add checkpointing by counting completions inside the runner and flushing every N — accepting that the flush covers a set, not a prefix. Or run a pool per segment: pool the 500 items of segment one, checkpoint, pool the next 500. That amortises the barrier over 500 items instead of 10, which reduces the idle-slot penalty by the same factor while keeping a clean resume point. ## One implementation trap in the batched form `chunk.map(worker)` passes `(value, index, array)` to the worker, and that `index` is chunk-local — item 7 of chunk 3 arrives as index 7, not 37. If the worker uses its second parameter for anything (labelling output, computing an offset, deriving a key), the batched version silently corrupts it while the pooled version, which passes the true cursor value, does not. Write `chunk.map((item) => worker(item))` unless you are deliberately passing an index. ## What to say in the interview Lead with the barrier, quantify it with a straggler example, then name the cases where the barrier is the point rather than the cost. That progression — mechanism, cost, and the situations that invert the tradeoff — is what separates a memorised answer from judgment.
- How would you keep a pool but still checkpoint every 500 items?Count completions inside the runner and, at each multiple of 500, flush progress. The catch is that the completed set is not a contiguous prefix — a few earlier indexes may still be in flight — so record completions individually, or drain the pool at the boundary and accept one barrier per 500 items instead of one per 10.
- Does batching bound memory better than a pool?Not for in-flight work: both hold at most N operations at once. It helps only with what you accumulate, because a batch boundary is a natural place to write results out and drop them. A pool can do the same by writing each result inside the runner instead of collecting into one array.
- If durations are near-identical, is there any reason left to prefer the pool?Little on throughput — the two converge as variance goes to zero. The pool still avoids the chunk-local index trap and keeps behaviour stable if the latency distribution later develops a tail, which it usually does once retries, cold caches, or a slow shard enter the picture.
Batching is a supermarket that lets ten shoppers to the tills, then makes everyone wait at the exit until the slowest one is done before admitting the next ten. A pool is the normal arrangement: the moment a till frees, the next shopper steps up.
saying these in an interview costs you the question
- Says both shapes have identical throughput because the cap matches
- Thinks Promise.all resolves as soon as the first job finishes
- Believes chunking is required to preserve result order
- Assumes a straggler only delays its own result
- Uses chunk.map(worker) without noticing the index is chunk-local