skip to content

An image-thumbnail worker raises OSError Too many open files after hours of steady work; how do you find the leaking descriptors?

level: seniorimportance: should knowfreq 45%

answer

  1. The traceback blames the innocent
  2. Graph a number nobody was graphing
  3. Memory stays flat while it climbs
  4. Read what each entry points at
  5. Reproduce small, with tracebacks on

basics

~20 s

Confirm the descriptor count grows with work done by sampling /proc/<pid>/fd, then classify the entries as files, sockets or pipes to point at the guilty code, and reproduce under development mode with tracemalloc to get the exact opening line.

solid answer

~50 s

The traceback is misleading here: `errno.EMFILE` surfaces wherever the next open happens, not where the leak is. Start with a count. On Linux `len(os.listdir("/proc/self/fd"))` from inside the process, or the same directory for the pid from outside (`/dev/fd` on macOS and the BSDs), sampled against a throughput counter, tells you whether descriptors grow monotonically or merely plateau. Note that resident memory usually looks flat — a worker with a 2.4 GB working set of decoded images shows no memory signal at all — so the rising fd count is the only indicator. Next, read each symlink target to classify the leak: regular files point at an `open()` without `with`; sockets at per-item clients; pipes at a `subprocess.Popen` whose `communicate()` or `wait()` is never called, which leaks zombies too. Finally, reproduce under `-X dev -X tracemalloc=25` so the `ResourceWarning` carries the allocation traceback and names the line.

code

python · 14 lines
python
import os
import sys


def open_fd_count():
    d = "/proc/self/fd" if sys.platform == "linux" else "/dev/fd"
    return len(os.listdir(d))  # the listing itself holds one descriptor


before = open_fd_count()
keep = [open(os.devnull) for _ in range(5)]
print(before, open_fd_count())
for f in keep:
    f.close()

go deeper

for a junior

Recognise Too many open files as descriptor exhaustion rather than a disk or memory problem, and know that the fix direction is closing files rather than restarting the service.

for a middle

Explain how to measure it: count the entries under the process fd directory over time, and know that pipes from subprocess and sockets count the same as files. Be able to name the with-block and ExitStack fixes.

for a senior

Demonstrate the full path from symptom to line: confirm monotonic growth, classify the descriptors by what they point at, reproduce under development mode with tracemalloc, and fix ownership rather than the ceiling. Note that memory graphs stay flat throughout.

for a principal

Own prevention: a descriptor gauge with slope alerting as a standard service metric, development mode in the default CI command, and an ownership convention in shared libraries so resources belong to a scope or an object that closes them.

### First, confirm the shape of the failure `OSError` with `errno.EMFILE` — *Too many open files* — means the process hit the ceiling on simultaneously-open descriptors. Almost always the reported line is innocent: it is simply the next piece of code that tried to open anything. So do not start by reading the traceback's function. Start by asking whether the count of open descriptors grows monotonically with work done. That is a directly measurable number. On Linux, `/proc/<pid>/fd` contains one symlink per open descriptor, so its length is the answer; from inside the process, `len(os.listdir("/proc/self/fd"))` works, and on macOS and the BSDs the equivalent directory is `/dev/fd`. Sample it every N items processed and log it next to a throughput counter. If it climbs and never falls while the work loop runs, you have a leak; if it plateaus at a high number, you have a pool sized larger than the ceiling, which is a different bug with a different fix. A revealing detail in this kind of incident: resident memory can look completely healthy. A thumbnail worker holding a 2.4 GB working set of decoded pixel buffers shows a flat, boring memory graph while its descriptor count marches upward, because a descriptor costs a few bytes in userspace. Memory profiling will not find this. That mismatch — flat RSS, rising fd count, failures clustered at the open call — is the signature. ### Second, find out *what* the descriptors are Reading the target of each `/proc/<pid>/fd` symlink (`os.readlink`) classifies the leak immediately, and the class points at the code: * **Regular files** under one directory — a helper that opens without `with`, or an error path that returns before an explicit `close()`. * **Sockets** — HTTP or database connections created per work item and never closed, or a client object that owns a connection pool being constructed inside the loop instead of once. * **Pipes** — the classic `subprocess` mistake: launching a child with `stdout` set to a pipe and never calling `communicate()` or `wait()`. The parent keeps the read end open, and the child is never reaped, so you leak descriptors and zombie processes in step. * **`eventpoll` / `inotify` / `timerfd`** entries — an event loop, watcher or notifier being created per item rather than once. From the outside, the same census comes from the standard `lsof -p <pid>` view; inside the process the `/proc` or `/dev/fd` listing needs no extra tooling and can ship as a health metric. ### Third, get a line number Reproduce the workload — a few hundred items is usually enough — under `python -X dev -X tracemalloc=25 worker.py`. Development mode un-silences `ResourceWarning`, which CPython emits whenever a file, socket or pipe object is deallocated while still open, and `tracemalloc` attaches the object's allocation traceback to the warning. That turns "something is leaking" into "line 84 of `thumbnail/io.py` opened this and nobody closed it". Two caveats. Objects that are still referenced never warn, because the warning fires at deallocation — so a leak held alive by a cache is invisible to this technique and shows up only in the descriptor census. And an exception path can keep a frame (and its locals) alive through the traceback, delaying the close long enough to matter under load even when the code "looks closed". ### Fourth, fix ownership, not the symptom The fix is to tie every descriptor to a scope: `with open(...)`, `with socket.socket() as s`, `with subprocess.Popen(...) as proc`, and `contextlib.ExitStack` when the count is dynamic. Connection and client objects that own descriptors should be created once and closed on shutdown, not per item. Where a leak came from an error path, the `with` block fixes it for free, which is why the rewrite is usually smaller than the investigation. Raising the process's descriptor ceiling is worth doing when the *legitimate* concurrent demand is genuinely above the default, but it is not a fix for a leak — it only changes how many hours pass before the same outage. If the count grows without bound, no ceiling is high enough. ### Leave a tripwire Once fixed, keep the cheap descriptor gauge in the metrics and alert on its slope, and run the test suite with development mode on so a reintroduced leak fails in CI rather than in production at 3am. Descriptor exhaustion is one of the few production failures with a perfect leading indicator; the only reason it ever takes anyone by surprise is that nobody was plotting the number.

  • Why is raising the process descriptor limit not the fix here?
    Because an unbounded leak defeats any ceiling; a higher limit only changes how many hours pass before the same outage, while making each incident larger and slower to recover. Raising it is legitimate when the *legitimate* concurrent demand genuinely exceeds the default, which you can tell apart by the shape of the curve: a plateau means sizing, a monotonic climb means a leak.
  • The census shows hundreds of pipe entries. What code shape produces that?
    A loop that launches children with `subprocess.Popen` and pipes for their output but never calls `communicate()` or `wait()`. The parent keeps the read end of each pipe open and the exited children are never reaped, so descriptors and zombie processes accumulate together. Using `subprocess.run`, or `with subprocess.Popen(...) as proc:`, closes the pipes and waits for the child.
  • Your CI run under development mode is clean but production still leaks. What is the likely explanation?
    ResourceWarning fires only at deallocation, so anything still referenced never warns. A cache, module-level registry or long-lived worker that accumulates open objects is invisible to the warning and visible only in the descriptor census. Sampling the count during a soak run, rather than relying on warnings from a short test, is what catches that class.

saying these in an interview costs you the question

  • Blames the line in the traceback where the error surfaced
  • Reaches for the descriptor limit as the first fix
  • Looks for the leak in memory profiles or heap growth
  • Assumes a subprocess pipe closes itself when the child exits
  • Believes a clean short test run proves there is no leak
  • Restarts the process on a timer and calls it solved

context