skip to content

How does asyncio.Semaphore bound a fan-out of hundreds of coroutines?

level: middleimportance: must knowfreq 55%

answer

  1. A counter of permits, not a lock
  2. Guard the item, never the gather
  3. Bounds concurrency, not the number of tasks
  4. Waiters queue on futures and yield the loop
  5. async with sem: around the constrained call

basics

~20 s

asyncio.Semaphore(n) holds n permits. async with sem: takes one and suspends the task when none are left, so at most n coroutines are inside the guarded block at once; waiters are released in order as permits come back.

solid answer

~50 s

Create one `asyncio.Semaphore(n)` for the resource you are protecting and wrap **the per-item work**, not the `gather` call, in `async with sem:`. Each entry decrements the permit count; when it hits zero the next task awaits a waiter future and yields to the event loop, so the loop keeps servicing the n coroutines that hold permits. On exit the permit is returned and the first waiter is resumed. Two things it does *not* do: it does not limit how many task objects exist — 500 `create_task` calls still allocate 500 tasks and their coroutine frames — and it does not bound a producer feeding them. When you need both concurrency and memory bounded, use a fixed pool of worker tasks reading a bounded `asyncio.Queue`. `asyncio.BoundedSemaphore` raises ValueError if a permit is released more times than it was acquired.

code

python · 19 lines
python
import asyncio

sem = asyncio.Semaphore(5)
peak = live = 0

async def handle(item):
    global peak, live
    async with sem:            # inside the item, not around gather
        live += 1
        peak = max(peak, live)
        await asyncio.sleep(0.01)
        live -= 1
    return item * 2

async def main():
    results = await asyncio.gather(*(handle(i) for i in range(200)))
    print(len(results), peak)   # 200 5

asyncio.run(main())

go deeper

for a junior

Know the shape: create one asyncio.Semaphore(n) and write async with sem: inside the coroutine that does the work. Remember that the number reflects what the downstream resource can take, not how many items you have.

for a middle

Explain the counter-and-waiters mechanics and why the permit must be acquired per item rather than around the gather. Be able to state what the semaphore does not bound: task objects, coroutine frames and retained results.

for a senior

Show judgment about where the limit belongs — one semaphore per constrained resource, cancellation and timeout interactions, and knowing when to abandon fan-out-plus-semaphore for a bounded queue with a fixed worker pool.

for a principal

Own the throttling strategy: whether the limit belongs in the client, a shared gateway or the downstream service, how it composes with retries and pool sizes, and how the chosen number is derived, measured and revisited rather than guessed.

## The mechanism `asyncio.Semaphore(value)` is **a counter plus a waiter deque**. - `acquire()` is a coroutine: if the counter is above zero it decrements and returns immediately; otherwise the task appends a future to the deque and awaits it, which yields to the event loop. - `release()` increments the counter and wakes the first waiter. - `async with sem:` is the idiomatic form and is the only one that survives exceptions and cancellation cleanly — hand-rolled `acquire()` / `release()` pairs leak a permit the first time the body raises before the `release()`, and a leaked permit permanently shrinks your concurrency. ## Where the `async with` goes This is the single most common mistake and a favourite interview probe. Guarding the fan-out itself — `async with sem: await asyncio.gather(*coros)` — takes exactly one permit for the entire batch and bounds nothing. The permit must be taken *inside* the **per-item coroutine**, around the part that touches the constrained resource. Everything outside the guarded block still runs concurrently across all items, which is usually what you want: parse and validate freely, serialise only the expensive call. ## What it bounds, and what it does not A semaphore bounds *in-flight work*, not *scheduled work*. If you build 5,000 coroutines and hand them to `asyncio.gather` or an `asyncio.TaskGroup` under a semaphore of 10, the loop really does keep only 10 inside the guarded region — but 5,000 task objects, 5,000 coroutine frames and 5,000 result slots exist for the whole run, and `gather` holds every result in memory until the last one lands. On a large input that is a memory problem the semaphore cannot solve. The structural fix is the **producer/consumer shape**: a bounded `asyncio.Queue` plus a fixed number of long-lived worker tasks. There, concurrency is the worker count and memory is the queue's `maxsize` — two knobs instead of one, and the producer is throttled by backpressure rather than racing ahead. ## Choosing n The number is a property of the ***constrained resource***, not of the machine: - a remote endpoint's rate limit, - a connection pool's size, - a disk's useful queue depth, - or a downstream service's politeness budget. Use one semaphore per resource — a per-host semaphore held in a dict keyed by host is a common pattern — rather than one global number that conflates them. If the resource already has its own limit (a pooled client with a maximum connection count), adding a semaphore of a different size in front of it just moves the queueing and makes the real limit harder to reason about. ## Cancellation and timeouts A task cancelled while awaiting a permit removes its own waiter and never holds the permit, so the count stays correct. Combining `asyncio.timeout` (3.11+) with a semaphore is subtle: the timeout covers *both* the wait for a permit and the guarded work, so a saturated semaphore makes every item time out even though the work itself is fast. If you need to distinguish "we were queued too long" from "the call was too slow", time the two separately. ## Semaphore vs the alternatives - **`asyncio.Lock`** is a semaphore of one, but with an owner-ish contract and a clearer intent — use it for mutual exclusion, not for throttling. - **`asyncio.BoundedSemaphore`** behaves identically but raises `ValueError` on a release that would push the counter above its initial value, which turns a silent permit leak into a loud bug; prefer it when the acquire and release are not lexically paired. - **`asyncio.Condition`** is the right tool when tasks must wait for a *predicate* rather than a count. ## Concurrency is not a rate A semaphore of 10 says "at most ten of these at once", which is not the same promise as "at most ten per second". If each guarded call takes 50 ms, ten permits sustain roughly 200 calls per second; if the downstream slows to 500 ms per call, the same ten permits sustain twenty. A quota expressed per unit time needs a **token-bucket-style limiter** — the semaphore only shapes how many operations are simultaneously outstanding. The two are frequently conflated in code review, and the giveaway is a constant chosen to match a documented requests-per-second limit. `locked()` reports whether a permit is available right now, which is useful for instrumentation but is a snapshot, not a reservation: acting on it instead of awaiting `acquire()` reintroduces the race the semaphore exists to remove. ## Version notes The `loop` parameter was removed from `asyncio.Semaphore` in 3.10; on 3.14 the object binds to the running loop on first use. As with every asyncio primitive it is not thread-safe — a semaphore bounding work on a loop cannot bound work you have pushed onto threads with `asyncio.to_thread`, where the executor's own worker count is the limit.

  • You wrap the asyncio.gather call itself in `async with sem:`. What have you actually limited?
    Nothing useful. One permit is taken for the whole batch and released when the last coroutine finishes, so all of them run concurrently inside it. The only effect is that two separate fan-outs cannot overlap. The permit has to be acquired inside the per-item coroutine, around the constrained call, so that each item competes for it individually.
  • A semaphore of 10 guards 50,000 items, yet memory still climbs. Why?
    Because the semaphore bounds in-flight work, not scheduled work. Building 50,000 coroutines or tasks up front allocates 50,000 frames, and asyncio.gather retains every result until the last completes. Switch to a fixed pool of worker tasks draining a bounded asyncio.Queue: the queue's maxsize caps memory, the worker count caps concurrency, and the producer is throttled by backpressure instead of materialising the whole input.
  • When would you reach for asyncio.BoundedSemaphore over asyncio.Semaphore?
    When acquire and release are not lexically paired — for example when a permit is released by a callback or a different code path than the one that took it. BoundedSemaphore raises ValueError if a release would push the counter above its initial value, so a double-release surfaces immediately instead of quietly inflating your concurrency limit. With a plain `async with` on a single block, the two are equivalent.

A semaphore is a bowl of n parking permits at the door: you take one to go in, hang it back on the way out, and everyone else queues at the door instead of circling the block.

saying these in an interview costs you the question

  • Wrapping the gather call in the semaphore instead of each item
  • "A semaphore limits how many tasks exist"
  • Using acquire()/release() by hand with no try/finally, leaking permits on error
  • Picking n from CPU count rather than the constrained resource
  • Assuming an asyncio.Semaphore also throttles work sent to asyncio.to_thread
  • Confusing asyncio.Semaphore with threading.Semaphore in async code

context