Your selectors-based webhook receiver leaks memory across a 340-case replay pack. How do you find the cause?
answer
- Two structures can grow, not more
- Registration count against buffered bytes
- Everything is reachable, nothing is garbage
- Unregister before close, always in finally
- Cap the frame, expire the idle
basics
~20 sMeasure the selector's registration count first. If it never returns to its idle value, connections are being abandoned without unregister() and close(). If it stays flat, the growth is in the per-connection buffers you attached at registration.
solid answer
~50 sThere are only two places a single-threaded readiness loop can grow without bound, and one measurement separates them. Export `len(sel.get_map())` as a gauge. If it climbs through the replay run, registrations are leaking: a handler raised, or the peer closed and the `recv()` returning `b""` was not treated as end-of-file, so the socket was never unregistered - and the state you attached as `data` is retained with it. If the map is flat, the growth is inside that state: an input buffer appending every chunk while waiting for a frame end that a malformed case never sends, or an output queue growing because a slow consumer stopped being write-ready. Confirm with `tracemalloc` snapshots taken before and after the run and compared by line. Fix with hard caps on both buffers and `try`/`finally` teardown that always unregisters before closing.
code
python · 32 linesimport selectors
import socket
MAX_BUFFER = 64 * 1024
sel = selectors.DefaultSelector()
def drop(sock):
try:
sel.unregister(sock)
finally:
sock.close()
a, b = socket.socketpair()
a.setblocking(False)
sel.register(a, selectors.EVENT_READ, data=bytearray())
b.sendall(b"x" * 100)
b.close()
while sel.get_map():
for key, _mask in sel.select(timeout=1):
chunk = key.fileobj.recv(4096)
if not chunk:
drop(key.fileobj)
continue
key.data.extend(chunk)
if len(key.data) > MAX_BUFFER:
drop(key.fileobj)
print("still registered:", len(sel.get_map()))
sel.close()go deeper
Know that every registered socket must eventually be unregistered and closed, and that an empty result from recv() means the peer is gone. That single habit prevents the most common growth in this loop shape.
Explain the two growth sites - the registration table and the per-connection buffers - and how a registration-count gauge tells them apart. Be able to write the teardown helper that unregisters and then closes.
Demonstrate the diagnosis end to end: gauges first, tracemalloc snapshot diff second, then a structural fix with capped buffers, one teardown path and idle expiry, plus the regression check that memory returns to baseline between runs.
Own the limits as policy rather than as constants scattered in a handler: maximum request size, maximum queued bytes per connection, idle deadline, and what the service does when it hits them. Decide whether one process should be holding this many connections at all.
### Why this loop shape leaks A readiness loop holds long-lived state in exactly two structures, and both are easy to grow and easy to forget. The first is the **selector's own registration table**. Every `register()` puts an entry in a map keyed by file descriptor holding a `selectors.SelectorKey` - which holds `fileobj` and, crucially, whatever you passed as `data`. That entry lives until `unregister()`. A connection you stop caring about but never unregister keeps its socket object, its buffers and its parsed state alive for the life of the process, and it also keeps being reported ready forever, so the leak comes with a CPU cost. The second is the **per-connection buffers** inside that `data`. Because TCP has no message boundaries, every such loop accumulates input until a complete frame is parsed out, and queues output until the peer takes it. Both are unbounded by default. ### Diagnose the shape before reading code The cheapest discriminator is the registration count: ```python print("registered:", len(sel.get_map())) ``` Export that as a gauge, and record total buffered bytes alongside it - `sum(len(k.data.inbuf) + len(k.data.outbuf) for k in sel.get_map().values())`. Run the replay pack and watch which one moves. **Count climbing** means abandoned registrations. Look for every path that leaves the loop's handling of a connection: an exception from a handler that skipped teardown, a `recv()` of `b""` that was treated as "nothing yet" rather than end-of-file, an error path that called `close()` but not `unregister()`. Note that closing a socket without unregistering is the worse of the two orders - the descriptor number is freed and may be reused by the next accepted connection while the selector still holds a stale entry for it, which is how a loop starts routing one connection's readiness to another connection's state. **Count flat but bytes climbing** means the buffers. In a webhook receiver the usual cause is framing: a request whose declared body length never arrives, or a case in the pack that sends a header and then holds the connection open, so the input buffer grows for as long as the peer is willing to trickle. The other cause is outbound: a consumer that stopped being write-ready while the loop kept appending responses to its output queue. ### Confirm with the standard tools `tracemalloc` answers "which line allocated the memory that is still alive". Call `tracemalloc.start()` before the run, take a `tracemalloc.take_snapshot()` at a quiet point, run the pack, snapshot again, and compare the two snapshots by line; the top entries name the allocation site directly. For object-count shape rather than allocation site, `gc.get_objects()` bucketed by type tells you whether it is `bytearray`s, sockets, or your own connection objects that are multiplying, and `len(gc.garbage)` plus `gc.collect()` returning a large number points at reference cycles - a connection state object referring back to the socket that refers to it, with cleanup relying on refcounting alone. One warning specific to this leaf: because the loop is single-threaded and holds the only references, nothing is "leaked" in the C sense. Every byte is reachable from the selector's map. Tools that look for unreachable garbage will report nothing wrong. The leak is a bookkeeping bug, not an allocator bug. ### Fix it structurally, not case by case Four rules make the shape safe. 1. **One teardown path.** A single `drop(sock)` helper that unregisters inside a `try` and closes in the `finally`, called from every exit: end-of-file, parse error, timeout, handler exception. Never two call sites that each do half. 2. **Cap the input buffer.** Decide the largest request the service accepts, and when a connection's buffer exceeds it, drop the connection rather than keep reading. An unbounded frame buffer is not only a leak, it is a way for one peer to exhaust the process. 3. **Cap the output queue.** When a peer stops being write-ready and the queue passes its limit, stop producing for it - pause reading its requests, or close it. Buffering indefinitely for a consumer that is not consuming converts a slow client into an outage. 4. **Expire idle connections.** Track a last-activity timestamp in the connection state and use the `select(timeout)` argument to run a periodic sweep that drops connections past their deadline. Without a timer pass, a peer that connects and says nothing costs a registration and a buffer forever - and in a replay harness that finishes without closing its sockets, that is by itself enough to make memory climb run over run. The test that proves the fix is the one the symptom already gave you: run the pack twice and assert that the registration count and total buffered bytes return to their idle values between runs.
- Why is closing a socket without unregistering it from the selector worse than the reverse mistake?Closing frees the file descriptor number, and the operating system reuses the lowest free number for the next accepted connection. The selector's map is keyed by that number, so a stale entry now collides with a live connection: readiness for the new socket can be dispatched against the old connection's state, or the new registration can fail outright. Unregistering first and closing in a `finally` keeps the map and the descriptor space in step.
- What would you add to the loop so a peer that connects and then says nothing cannot accumulate?A deadline. Store a last-activity timestamp in each connection's state, pass a bounded `timeout` to `select()` so the loop wakes even with no traffic, and sweep the registration map each pass for connections past their idle limit, dropping them through the same teardown path. This is also what makes the `select(timeout)` argument matter: it is the loop's only source of time-based work.
- The memory is climbing but gc.collect() frees nothing. What does that tell you?That there is no cycle and no unreachable garbage - the objects are still reachable, which means your own data structures are holding them. In a readiness loop that points straight at the selector's registration map or the buffers hanging off it, not at a collector problem. It rules out the reference-cycle explanation and sends you to the bookkeeping instead.
saying these in an interview costs you the question
- Reaches for the garbage collector before measuring registrations
- Calls close() on a socket without unregistering it
- Reads until a delimiter with no size limit
- Keeps buffering for a peer that stopped being write-ready
- Treats an empty recv() result as nothing arrived yet
- Blames the allocator when everything is still reachable