skip to content

A ctypes callback segfaults during a 6-hour nightly email-digest run, and the restart re-sends digests. How do you diagnose it?

level: seniorimportance: should knowfreq 22%

answer

  1. Two memory models, no shared bookkeeping
  2. C keeps an address, not a reference
  3. Collected object, address still readable
  4. faulthandler prints a traceback on the signal
  5. Debug allocator makes it fail immediately

basics

~20 s

Suspect object lifetime: C keeps a raw address and no Python reference, so a collected callback object or buffer leaves it dereferencing freed memory. Enable faulthandler for a traceback and re-run under -X dev to make it deterministic.

solid answer

~50 s

A crash at the `ctypes` boundary is almost always something on the C side outliving the Python object whose address it holds. The classic case is a `ctypes.CFUNCTYPE` callback object referenced only by a local in the function that registered it: it is collected on return, and the library calls into freed memory hours later, when the allocator finally reuses the block. Buffers behave the same way when the library stores the pointer instead of using it call-scoped. I would enable `faulthandler.enable()` to get a Python traceback from the fatal signal, re-run under `-X dev` so the debug allocator makes a use-after-free crash immediately and reproducibly, then hoist the callback and buffers to module scope and see if the crash disappears. Separately, the duplicate sends are their own defect: a segfault gives no `finally` and no flush, so the job needs a durable per-recipient sent record it can resume from.

code

python · 27 lines
python
import ctypes

SEND_CB = ctypes.CFUNCTYPE(ctypes.c_int, ctypes.c_uint32)
_REGISTERED = []                 # C stores only the address, never a reference


def deliver(recipient_id):
    if recipient_id == 0:
        raise ValueError("no such recipient")


def on_recipient(recipient_id):
    try:
        deliver(recipient_id)
    except Exception:            # an exception cannot cross the C frames
        return 1                 # so hand C an error code it understands
    return 0


def register(fn):
    cb = SEND_CB(fn)
    _REGISTERED.append(cb)       # without this the object dies at return
    return cb


handler = register(on_recipient)
print(handler(4200), handler(0))

go deeper

for a junior

Take away one rule: anything you hand to C — a callback object, a buffer, an array — must stay referenced from Python for as long as C might use it, because C keeps only the address.

for a middle

Be able to explain the mechanism: refcounts drop when the last Python name goes away, freed memory stays readable for a while, and that delay is why the crash looks random.

for a senior

An interviewer expects a diagnosis method, not a guess: faulthandler for the signal, the debug allocator to force determinism, hoisting objects to test the hypothesis, a native debugger for the C frames, and re-verifying the prototypes.

for a principal

Own the consequence: a native crash is an availability event with no traceback and no cleanup. Decide whether the library runs in-process at all, or behind a subprocess that can die alone, and require long jobs to be resumable and idempotent regardless.

### The failure class At the `ctypes` boundary two memory models meet. Python frees an object when its last reference goes away; C holds raw addresses and holds **no** Python reference at all. Nothing bridges the two, and almost every crash of this shape is the same bug: something on the C side outlived the Python object whose address it kept. The timing is what makes an overnight job hard to diagnose. Freed memory usually stays readable and plausible for a while, so the callback appears to work through most of the run and only dies once the allocator hands that block to something else — which depends on how much unrelated work has happened. That is exactly why a six-hour run crashes in hour four and a ten-minute reproduction never fails. ### The three lifetime traps **1. The callback object itself.** `ctypes.CFUNCTYPE(restype, *argtypes)(python_function)` builds a C function pointer backed by a Python object. Register it inside a helper and return, and if the only reference was that helper's local, the object is collected immediately and the library is left holding a pointer into freed memory. Keep the object referenced for at least as long as the library can call it: at module level, as an attribute of a long-lived object, or in an explicit registry you delete from when you unregister. **2. Buffers whose address is retained.** Passing `ctypes.byref(buf)` for a call-scoped out-parameter is safe, because `ctypes` keeps arguments alive for the duration of the call. It is fatal when the library stores the pointer for later use. The same trap in miniature: assigning a freshly built array to a struct's pointer field, `job.ids = (ctypes.c_uint32 * n)(...)`, frees the array on the very next line, because the array's only reference was the temporary. **3. Memory the library owns.** A pointer the library returned belongs to the library. Copying the bytes out with `ctypes.string_at` and then calling the library's own free function is correct; keeping the pointer past that free, or freeing it with a different allocator than the one that allocated it, is not. ### The silent-corruption trap next door A Python exception cannot propagate out through C frames. When a `ctypes` callback raises, `ctypes` prints the traceback to stderr and returns a **default result** — zero — to the C caller, which most libraries read as success. A digest run whose per-recipient callback has started raising will keep going and keep reporting success, and stderr is the only evidence. Every callback body should therefore be wrapped in `try`/`except`, log the failure, and return an explicit error code the C side understands. ### Diagnosis, in order **Get a Python traceback out of the fatal signal.** `faulthandler.enable()` at startup — or `-X faulthandler`, or `PYTHONFAULTHANDLER=1` — makes CPython dump the Python stack of every thread on a segmentation fault. That tells you which Python frame was live when C died. For a callback crash the telling detail is often that there is no Python frame above the foreign call at all: the library was calling into a pointer Python no longer owned. **Make it deterministic.** Run under `-X dev`, which enables the debug allocator: freed memory is overwritten with a fill pattern and blocks carry guard bytes, so a use-after-free or an overrun crashes immediately and reproducibly instead of hours later. `PYTHONMALLOC=debug` is the environment spelling of the same switch. Forcing a `gc.collect()` right after registration is the cheap version of the experiment. **Test the hypothesis by removing it.** Hoist the callback object and every buffer to module scope so nothing can be collected, and re-run. If the crash disappears, it was a lifetime bug, and you can put the objects back one at a time until it returns. **Read the C frames.** A native debugger on the core file gives the C side of the stack that `faulthandler` cannot show; the faulting frame names the library function that dereferenced the dead pointer. **Re-check the declarations while you are there.** A wrong `restype` or a missing `argtypes` entry produces the same symptom class — a fatal signal with no Python traceback — and re-verifying the prototypes against the C header is cheaper than chasing a refcount. ### The duplicated sends are a second, separate defect A segmentation fault kills the process outright: no `finally`, no `atexit` handler, no flush of anything buffered. Whatever the run had done but not durably recorded is lost, so the restart repeats it and some recipients receive the digest twice. Fixing the lifetime bug removes this crash, not the class of crash — a native library can still abort the process tomorrow. So treat the duplicate as its own finding: record each recipient as sent in durable storage as part of the same step that sends, resume from that record rather than from the top, and keep the run re-entrant. Any long job that calls into native code should be written on the assumption that it can vanish mid-iteration without warning, because at that boundary it genuinely can.

  • Why does running under -X dev turn an intermittent crash into a reproducible one?
    It enables CPython's debug allocator, which fills freed blocks with a pattern and puts guard bytes around live ones. A use-after-free then reads obvious garbage instead of stale-but-plausible data, and an overrun is detected at the next check rather than whenever the heap happens to notice. The failure moves from 'hour four, sometimes' to 'the first time the dead pointer is touched'. `PYTHONMALLOC=debug` is the environment-variable spelling.
  • Is passing ctypes.byref of a freshly created object ever safe?
    Yes, for a call-scoped out-parameter: ctypes keeps the arguments alive for the duration of the call, so the library can write into the buffer and you read it afterwards. It is unsafe the moment the library retains the pointer past the call — a registration function, an object it stores in its own state, or a pointer field copied into a struct it keeps. Then you need a named owner that outlives the library's use of it.
  • The library invokes your callback from a thread it created itself. What has to happen for that to be safe?
    ctypes attaches a thread state and acquires the GIL before running the Python callback, so the code itself is safe to execute. What is on you is that the callback runs concurrently with the rest of your program on a thread you did not create: it must be thread-safe, must not assume any thread-local or main-thread context, and should do very little, since it holds the GIL while it runs and blocks other Python threads.

You gave a contractor your address on a slip of paper and then moved out. The slip is still readable; the house behind it now belongs to someone else.

saying these in an interview costs you the question

  • Assumes the garbage collector knows C holds a pointer
  • Expects an exception in a callback to reach the Python caller
  • Says a segmentation fault prints a Python traceback by default
  • Registers a callback keeping only a local reference to it
  • Treats the restart duplicates as the same bug as the crash
  • Blames the GIL for a crash inside a foreign call

context