How does asyncio's selector-based event loop use the selectors module underneath?
answer
- Same loop you would write by hand
- One selector, one thread, one queue
- Timeout comes from the nearest timer
- Ready events become queued callbacks
- add_reader is the raw registration seam
basics
~20 sIt owns a selectors.DefaultSelector, registers each managed socket for EVENT_READ or EVENT_WRITE, and calls select() with a timeout set by its nearest scheduled timer. Ready events become queued callbacks that the same single thread then runs.
solid answer
~50 s`asyncio.SelectorEventLoop` is the readiness loop you would write by hand, with a callback scheduler on top. It holds one `selectors.DefaultSelector` instance, so on Linux it is epoll-backed and on macOS and the BSDs kqueue-backed. Each iteration it computes a timeout from the earliest entry in its timer heap, calls `select(timeout)`, converts each returned `(SelectorKey, mask)` into the callbacks stored as that registration's `data`, appends them to a ready queue along with any timers that came due, and then runs every queued callback to completion. `add_reader()` and `add_writer()` on the loop are the public seam onto exactly that registration, and `remove_reader()` drops it. Two consequences follow directly: a callback that blocks - a CPU-heavy loop, a synchronous file or database call - stalls every registered connection, because the same thread that runs callbacks is the thread that calls `select()`; and the loop offers no parallelism, only overlapping of waiting.
code
python · 21 linesimport asyncio
import socket
async def main():
a, b = socket.socketpair()
loop = asyncio.get_running_loop()
ready = loop.create_future()
def on_readable():
loop.remove_reader(a.fileno())
ready.set_result(a.recv(64))
loop.add_reader(a.fileno(), on_readable)
b.sendall(b"hello")
print(await ready)
a.close()
b.close()
asyncio.run(main())go deeper
Recall that asyncio does not create threads for connections: one thread asks the kernel which sockets are ready and runs the resulting callbacks. That alone explains why blocking calls are forbidden inside a coroutine.
Be able to sketch the loop: selector plus timer heap plus ready queue, with the select() timeout derived from the nearest deadline. Explain what add_reader() registers and what constraints it inherits from the selectors layer.
Show that you can reason about tail latency from this structure - identifying an overrunning callback, using debug mode, and moving blocking or CPU-heavy work to an executor so the loop keeps polling.
Own the readiness-versus-completion distinction and its portability consequences, and decide when a single-loop-per-process design should become multiple processes, each with its own loop, rather than being tuned further.
### One selector, one thread, one queue Strip the coroutine machinery away and asyncio's selector-based event loop is a readiness loop with three data structures: - a `selectors.DefaultSelector` holding the registered sockets and, as each registration's `data`, the callbacks to invoke when it is readable or writable; - a **timer heap** of callbacks scheduled to run at a future time; - a **ready queue** of callbacks to run on this pass. One iteration does this: 1. Compute a timeout: `None` if there is nothing queued and no timer, otherwise the time until the earliest timer, otherwise `0` if callbacks are already queued. 2. Call `select(timeout)` on the selector - the single blocking point in the whole program. 3. For every `(SelectorKey, mask)` returned, push the read or write callback stored in `key.data` onto the ready queue. 4. Move every timer whose deadline has passed onto the ready queue. 5. Pop and run every callback currently queued, each to completion. Everything higher up is built on step 5. Awaiting a socket read ultimately registers a callback that resolves a future; resuming a coroutine is a callback that steps it to its next suspension point. The selector is where all the waiting actually happens. ### The public seam The loop exposes its own registration directly: `add_reader(fd, callback, *args)` and `add_writer(fd, callback, *args)` register interest in exactly the sense the `selectors` module means, and `remove_reader(fd)` / `remove_writer(fd)` withdraw it. These are the escape hatch for integrating a file descriptor the loop does not otherwise know about - a signal or self-pipe, a descriptor from a C library, a raw socket you want to drive yourself - without leaving the loop or spawning a thread. ```python loop = asyncio.get_running_loop() loop.add_reader(sock.fileno(), on_readable) ``` The same constraints apply as in a hand-written loop: the descriptor must be non-blocking, the callback must not block, and write interest should be held only while there is something to flush. ### Why the layering explains asyncio's rules **"Never block in a coroutine"** is not style advice; it is arithmetic. Steps 2 and 5 are the same thread. A callback that spends 200 ms in a synchronous call is 200 ms in which `select()` is not called, so no socket's readiness is noticed, no timer fires, and every connection registered with that loop is stalled - including ones with nothing to do with the blocking call. Debug mode (`asyncio.run(main(), debug=True)`) exists to log callbacks that overrun a threshold, precisely because the symptom - latency on unrelated connections - does not point at the culprit. **"asyncio is concurrency, not parallelism"** follows from the same picture: one thread runs step 5, so CPU-bound work is serialised no matter how many tasks exist. Offloading it to a thread or process executor is not a workaround, it is the only structure that keeps the loop calling `select()`. **Timer resolution comes from the timeout argument.** Scheduled callbacks are not run by a separate timer thread; they are run because the loop asked `select()` to return no later than the nearest deadline. A blocking callback therefore delays timers by exactly as much as it delays readiness. ### Where the model changes The selector-based loop is the default on Unix. On Windows, the default is a completion-based loop built on the operating system's I/O completion ports rather than on readiness, because Windows has no scalable readiness interface for sockets; that loop has been the platform default since Python 3.8. The distinction is not academic: `add_reader()` and `add_writer()` are the readiness API and are unsupported on the completion-based loop, which is why code that reaches for them is portable to Unix only. It is also the reason the two models are worth naming separately - **readiness** says "a call will not block, now do it", while **completion** says "the transfer you asked for has finished". ### What this is worth in an interview Being able to draw the loop above is the difference between using `async`/`await` and understanding it. It explains why one blocking call ruins a service's tail latency, why thousands of idle connections are cheap while one busy computation is not, why a `sleep` that is not the asyncio one freezes everything, and why the framework can only ever be as good as the selector underneath it. It is also the honest boundary of the abstraction: coroutines are an ergonomics layer over a callback loop over a readiness multiplexer, and every performance question eventually lands on one of those three floors.
- Where does the timeout passed to select() by an asyncio event loop come from?From the loop's own schedule. If callbacks are already queued it passes zero and polls; if a timer is pending it passes the time remaining until the earliest deadline; if there is neither, it blocks indefinitely. That is why scheduled callbacks need no timer thread - they run because the readiness call was told to return by then - and why a long-running callback delays timers by exactly the time it overran.
- Why does one synchronous blocking call inside a coroutine hurt connections that have nothing to do with it?Because the thread executing callbacks is the thread that calls `select()`. While the blocking call is in progress the loop is not asking the kernel which sockets are ready, so no readiness is noticed and no timer fires for any registration. The latency lands on every connection served by that loop, which is why the symptom rarely points at the offending call. Offloading it to a thread or process executor restores the loop's ability to poll.
- Why are the loop's add_reader() and add_writer() methods unavailable on Windows by default?Because the default event loop there is completion-based, built on I/O completion ports rather than on a readiness multiplexer, and has been since Python 3.8. Readiness registration has no meaning in that model: the operating system reports that a requested transfer has finished rather than that a call would not block. Code that registers descriptors through those methods is therefore Unix-only, and portable code uses the transport and stream layers instead.
The coroutine layer is a script; the callback queue is the stage manager reading it; the selector is the one doorman who tells the stage manager which door just opened. If the stage manager stops to do arithmetic, the doorman is never asked again.
saying these in an interview costs you the question
- Thinks the event loop runs callbacks on a thread pool
- Believes asyncio gives parallelism for CPU-bound work
- Says a separate timer thread fires scheduled callbacks
- Cannot connect blocking a callback to stalled connections
- Assumes every platform's default loop is readiness-based