skip to content

You run `await Promise.all(urls.map(url => fetch(url)))` over an array of 5,000 URLs. What actually happens, and why can Promise.all not limit how many requests are in flight?

level: juniorimportance: must knowfreq 62%

answer

  1. combinator, not a scheduler
  2. map runs to completion first
  3. promises start when created
  4. hold functions, not promises
  5. cap belongs where work begins

basics

~20 s

All 5,000 requests start immediately: map calls fetch for every URL before Promise.all ever runs. Promise.all only observes promises that are already in flight, so any concurrency cap must be applied where the work is started.

solid answer

~50 s

`map` is synchronous — it walks the whole array and calls `fetch` 5,000 times before it returns, so by the time `Promise.all` sees the array every request is already issued. A promise is a handle on work that has already begun, not a description of work to do later, and `Promise.all` is a combinator: it subscribes to an iterable of promises and resolves to their results in input order. None of the combinators takes a concurrency argument, because by the time they run there is nothing left to schedule. In practice you get thousands of queued sockets, memory held for every in-flight request and its response, and a downstream that starts rate-limiting or timing out all at once. The cap has to live in the code that *calls* `fetch`: keep the work lazy as thunks and start a new one only when a slot frees.

code

javascript · 18 lines
javascript
const load = (url) => new Promise((r) => setTimeout(() => r(url), 50));
const urls = ['/a', '/b', '/c', '/d', '/e'];

// Lazy: nothing is in flight yet, these are just functions.
const tasks = urls.map((url) => () => load(url));

let next = 0;
async function runner() {
  while (next < tasks.length) {
    const task = tasks[next++];
    console.log(await task());
  }
}

(async () => {
  await Promise.all([runner(), runner()]); // at most 2 in flight
  console.log('done');
})();

go deeper

for a junior

Be able to say plainly that map fires every call before Promise.all is reached, so all 5,000 requests are already in flight. Know that the fix is to start work in batches or through a limiter rather than all at once.

for a middle

Explain the mechanism: promises are eager, so a promise is work already begun, and a limiter therefore has to hold functions instead. Show the thunk-plus-runners shape and say why the concurrency number is independent of list length.

for a senior

Talk about what breaks concretely — sockets and descriptors, memory held per in-flight response, downstream rate limits, and latency numbers that become queue time. Frame the defect as concurrency being a function of input size rather than a chosen bound.

for a principal

Own it as a code-review and platform rule: unbounded fan-out is a class of defect, so shared limiters belong in the client layer every job uses rather than in each job. Be ready to argue where the bound is enforced and how a team keeps it from regressing.

## What that one line actually does Read it inside-out. `urls.map(url => fetch(url))` is a plain synchronous array method: it iterates all 5,000 entries, invokes the arrow function for each, and returns an array of 5,000 promises. `fetch` is called on iteration 1, then 2, then 3 — with no awaiting anywhere inside `map`, the whole loop runs to completion in one uninterrupted turn. Every request is dispatched before the array even exists as a value. Only then does `Promise.all` receive it, and only then does `await` suspend the surrounding async function. So the sequence is: 5,000 requests issued, then a single wait. Nothing about `Promise.all` or `await` was ever in a position to hold requests back. ## Promises are eager This is the mechanism worth naming in an interview. In JavaScript a promise does not represent "work I could do"; it represents "work that has started, whose result will arrive". The moment `fetch(url)` is evaluated, the request is on its way. Contrast the two shapes: ```js const p = fetch(url); // running right now const thunk = () => fetch(url); // nothing has happened yet ``` A limiter can only work with the second shape. If you hold an array of promises, the fan-out already happened and your "limiter" would just be choosing the order in which you *look* at results. If you hold an array of zero-argument functions, you can call as few or as many as you like, whenever you like. "Hold functions, not promises" is the whole trick behind every concurrency limiter you will ever write or install. ## Promise.all is a combinator, not a scheduler `Promise.all` takes one argument: an iterable. It attaches handlers to each element, keeps a counter, and resolves with an array whose positions match the input positions. It has no timing behaviour of its own and no knob for concurrency — and neither do the other standard combinators. When someone asks "how do I make Promise.all only do 5 at a time", the honest answer is that the question is aimed at the wrong function. ## What actually breaks at 5,000 - **Downstream.** The service sees a burst it never agreed to. You trip rate limits, or you fill its queue so that every one of your requests sits behind the others and they all approach the timeout together. - **Client resources.** Each in-flight request holds a socket (and in Node a file descriptor, plus DNS and TLS work), plus request state and eventually a response body. Hosts also queue connections per origin, so a large slice of your 5,000 may simply be sitting in a queue you cannot see — you paid the memory cost without buying any parallelism. - **Observability.** Everything is issued at t=0 and completes in one clump, so per-request latency stops meaning anything useful; you are measuring queue time. - **Blast radius.** A job that behaves fine on the 12-item test fixture behaves like a load test on the 200,000-item production list. The bug is that concurrency is a function of input length rather than a chosen number. ## Where the cap belongs Keep the work lazy, then drain it with a fixed number of concurrent runners: ```js const tasks = urls.map((url) => () => fetch(url)); // lazy let next = 0; async function runner() { while (next < tasks.length) { await tasks[next++](); } } await Promise.all([runner(), runner(), runner()]); // at most 3 in flight ``` Note that `Promise.all` is still here and still useful — it now waits on three long-lived runners instead of 5,000 requests. The concurrency number is `runner()` calls, chosen by you, independent of `urls.length`. The same laziness explains why `items.map(x => limiter(() => work(x)))` is fine even though it is still a `map`: the mapped function hands a *thunk* to a limiter that queues it, so mapping creates queue entries rather than requests. ## The review heuristic Any time you see `.map(asyncFn)` handed to a combinator over a list whose length you do not control, ask: what caps this? If the answer is "nothing", it is unbounded fan-out, and it will be found in production rather than in review. For a fixed, small list — four config endpoints at startup — unbounded is perfectly correct and idiomatic. The defect is fan-out proportional to data.

  • If you want to build the list of work up front anyway, what should that array hold?
    Zero-argument functions (thunks) rather than promises: `urls.map(url => () => fetch(url))`. Building the array then costs nothing, and a limiter can invoke as many thunks as its slot count allows and hold the rest back. An array of promises is already-running work, so nothing downstream of it can reduce concurrency.
  • Does awaiting the mapped array one element at a time in a for...of loop reduce concurrency?
    No. The requests were all issued by `map` before the loop started; iterating with `await` only changes the order in which you observe results, not when the work runs. It also makes error handling worse, since rejections that settle early sit unobserved until the loop reaches them.
  • Is the unbounded form ever acceptable?
    Yes, when the list is fixed and small and you know the downstream tolerates it — fetching four config endpoints at startup, or hydrating three widgets on a page. The problem is fan-out whose width is set by input length: the same code that issues 4 requests in a test issues 200,000 in production.

saying these in an interview costs you the question

  • Thinks Promise.all accepts a concurrency or limit option
  • Believes Promise.all runs the promises, or runs them one by one
  • Says awaiting later means the requests start later
  • Claims JavaScript runs the fetches on parallel threads
  • Assumes a limiter can throttle an array of already-created promises

context