skip to content

How does Python's development mode help you find a file-descriptor leak in a service that dies at a 1,200-request-per-minute peak?

level: seniorimportance: should knowfreq 28%

answer

  1. The interpreter already emits the signal
  2. The default filters are hiding it
  3. One switch shows it, another locates it
  4. Collection time is not leak time
  5. Reproduce under load, not in production

basics

~10 s

Development mode unhides ResourceWarning, which CPython raises when an unclosed file, socket or subprocess is garbage-collected. Add -X tracemalloc=25 and each warning carries the traceback of where the leaked object was allocated.

solid answer

~40 s

The leak already announces itself; the default filters hide the announcement. CPython emits a `ResourceWarning` when a file, socket or `subprocess.Popen` object is collected while still open, and `python -X dev` makes those visible. On its own the warning describes only the object being collected, so add `-X tracemalloc=25` (development mode does not enable memory tracing) and the warning gains the traceback of the allocation site — the exact line that opened the descriptor. Reproduce under load rather than in production: run the service at its peak profile in a staging process with both switches on, and consider `-W error::ResourceWarning` in a narrow reproduction so the first leak raises with a full traceback. Then fix the ownership — a `with` block or an explicit close — rather than raising the descriptor limit.

code

console · 1 line
console
python -X dev -X tracemalloc=5 -c "import socket; s = socket.socket(); del s"

go deeper

for a junior

Know that CPython warns about a file left unclosed, that the warning is hidden by default, and that a with block is the fix. Recognising Too many open files as a leak rather than a limit problem is the expectation here.

for a middle

Explain that ResourceWarning is emitted by the object's finalizer, that development mode unhides it, and that tracemalloc is what adds the allocation traceback. Be able to say why the two switches are separate.

for a senior

Demonstrate the investigation: reproduce under a matching load profile in staging, combine the switches, interpret bursty warnings as reference cycles, and fix ownership with context managers rather than raising limits. State plainly that the instrumented run is not a performance measurement.

for a principal

Own the prevention: which environments run instrumented, how a resource-leak reproduction is kept cheap and repeatable, and how the signal becomes a permanent CI check so the next leak is caught by a test rather than by a traffic peak.

## The shape of the problem A service — say a flight-schedule differ that compares published and actual departure times — runs fine for hours and then fails at its 1,200-request-per-minute peak with `OSError: [Errno 24] Too many open files`. The descriptor count climbs with traffic and never comes back down. Raising the limit buys hours and fixes nothing. ## CPython is already telling you File objects, sockets and `subprocess.Popen` objects all emit a `ResourceWarning` from their **finalizer** when they are collected while still open. In a normal run you never see it, because the default warning filters ignore `ResourceWarning` entirely. That is the first move: run the service under `python -X dev`, which applies the default filter as part of **development mode**, and the unclosed objects start reporting themselves as they are collected. ## Why the bare warning is often not enough The warning is emitted at *collection* time and can only describe the object being collected — `unclosed file <_io.TextIOWrapper ...>` — which tells you what leaked but not who opened it. In a codebase with one helper that opens everything, that is nearly useless. The fix is to **combine two switches**: development mode does **not** enable `tracemalloc`, so add `-X tracemalloc=25` (or `PYTHONTRACEMALLOC=25`). With memory tracing active, CPython attaches the traceback of the *allocation* site to the `ResourceWarning`, and you get the frames that ran when the descriptor was opened. The number is the **frame depth**: deeper is more informative and more expensive, and 10–25 is usually plenty to cross the helper boundary and reach real application code. ## Why collection time matters - In CPython, an object whose refcount drops to zero is finalized immediately, so a leaked file inside a function that returns normally warns almost at once. - But an object caught in a **reference cycle** — the common case in a service, where a handler holds a client that holds a connection that holds a callback back to the handler — is only reclaimed when the cyclic collector runs, so the warning arrives late, in a batch, on an unrelated request's stack. Do not read the timing of the warning as the timing of the leak; read the allocation traceback. If the warnings only appear when you force a collection, you have learned something important: the objects are in cycles, and a `with` block would have closed them deterministically regardless. ## Making the leak reproducible Do this under a load profile that matches the peak, **in staging, not in the production process**. Development mode costs real performance — the allocator debug hooks touch every allocation and free — and `tracemalloc` adds per-allocation bookkeeping on top, so throughput at 1,200 requests per minute will not match production and should not be expected to. You are hunting a functional bug, not measuring latency. In a narrow reproduction — one request replayed in a loop — `-W error::ResourceWarning` is even better than displaying: the first leaked descriptor raises, and you get a full traceback from the point of collection alongside the allocation traceback. ## Two more things development mode gives you here - It makes `io.IOBase` destructors *log* exceptions raised by `close()` instead of swallowing them, which is how you discover that a file was being closed but the close itself was failing. - And it enables `faulthandler`, so if descriptor exhaustion ends in a hard crash somewhere in a C extension rather than a clean `OSError`, you get a Python traceback out of it. ## What it will not find Development mode surfaces resource, deprecation, import and allocator problems. It does not check your logic. If the same service also computes a slightly wrong delay in minutes because of a floating-point rounding drift, no amount of `-X dev` will say a word about it — that is a test's job. Being clear about the boundary is part of a good answer: this is an **instrumentation switch with a specific catchment, not a correctness checker**. ## Then fix ownership, not symptoms The durable fix is **deterministic closing**: - a `with` block around each acquisition, - `contextlib.ExitStack` where the number of resources is dynamic, - an explicit `close()` in a service shutdown path for long-lived clients, - and never relying on the garbage collector to release an operating-system resource. Keep the `ResourceWarning` visible in CI afterwards so the next one is caught by a test rather than by a peak.

  • Would development mode have caught a floating-point rounding drift in the computed delay minutes?
    No. Development mode surfaces resource, deprecation and import warnings, checks encoding arguments, and instruments the C allocator; none of that inspects arithmetic. A rounding drift is a correctness bug and belongs to tests and to a deliberate choice of numeric type — `decimal.Decimal` or integer minutes — not to an interpreter switch. Knowing what a diagnostic tool does not cover is as useful as knowing what it does.
  • The warnings appear in bursts on unrelated stacks. What does that tell you?
    That the leaked objects are in reference cycles, so they are not finalized when the last name goes out of scope but only when the cyclic garbage collector runs. The collection stack is therefore meaningless — read the allocation traceback from `tracemalloc` instead. It also argues for the fix: deterministic closing with a `with` block releases the descriptor regardless of when the object itself is reclaimed.
  • Why not just raise the process file-descriptor limit?
    It converts a fast, obvious failure into a slow one. The leak still grows with traffic, so the service now dies later, under a heavier load, with more state in flight and a harder-to-reproduce trigger. A higher limit is legitimate as a temporary shield while you deploy the real fix, and only when you know the steady-state descriptor need genuinely exceeds the current limit.

saying these in an interview costs you the question

  • Assuming the garbage collector reliably closes descriptors
  • Reading the warning's stack as the leak's location
  • Enabling development mode in production to investigate
  • Expecting `-X dev` to enable tracemalloc as well
  • Raising the descriptor limit and calling it fixed
  • Believing development mode checks program logic

context