skip to content

A clinical-lab result loader dies with SIGSEGV inside a compiled extension - how do you capture a core file and read a Python stack from it?

level: seniorimportance: should knowfreq 26%

answer

  1. The process is gone; the image is not
  2. First check whether a core was even allowed
  3. The kernel decides where it lands
  4. Without symbols the stack is question marks
  5. A helper turns C frames into Python frames

basics

~20 s

Raise the core-size limit, find where the kernel writes cores, reproduce the crash, then open the core in gdb against the exact interpreter binary that produced it and use CPython's helper to print Python frames.

solid answer

~50 s

First make sure a core can be written at all: the soft `RLIMIT_CORE` is commonly 0, so raise it with `ulimit -c unlimited` in the launcher or `resource.setrlimit` early in the process, and find out where the kernel puts the file - on Linux that is `/proc/sys/kernel/core_pattern`, which may pipe cores to a collector rather than writing them next to the process. Then reproduce and open the core with gdb against the *same* interpreter binary and the same extension build; a different build gives you nonsense addresses. Debug symbols are what make it readable - install the interpreter's debuginfo and build or fetch an unstripped extension, otherwise the backtrace is a wall of `??`. CPython ships a gdb helper script that adds `py-bt`, `py-locals` and `py-up`, which walks the interpreter's C frames and prints the Python call stack that was live at the moment of the fault.

code

python · 8 lines
python
import resource

soft, hard = resource.getrlimit(resource.RLIMIT_CORE)
print("before:", soft, hard)

# Raising the soft limit up to the hard limit needs no privileges.
resource.setrlimit(resource.RLIMIT_CORE, (hard, hard))
print("after:", resource.getrlimit(resource.RLIMIT_CORE))

go deeper

for a junior

You are not expected to drive a debugger here. Recognise the shape - process gone, no traceback, status 139 - and know that the crash came from compiled code, so the next step is someone capturing a core file rather than reading application logs.

for a middle

Be able to enable cores and find them: raise the core-size limit in the launcher or with resource.setrlimit, check the kernel's core destination, and understand that the core must be opened against the exact interpreter and extension binaries that produced it.

for a senior

Demonstrate the full loop under production constraints - preserving binaries before a redeploy removes them, getting debug symbols in place, reading the C and Python stacks together, and recognising a corruption signature where the crash site is not the bug site.

for a principal

Own the readiness question: whether hosts are configured to keep cores at all, where they are stored and for how long given they contain process memory and therefore live data, and whether native extensions run in isolated worker processes so one crash is contained.

## The situation A loader that ingests result files runs for about 45 seconds of cold start - warming an in-memory index before it touches the first record - and then dies. No traceback, nothing in the application log, status 139. Everything Python-level has already told you all it can. The only remaining witness is the process image at the moment of the fault, which is what a core file is. ## Step 1: make a core possible The most common reason there is no core file is that there was never going to be one. The soft limit `RLIMIT_CORE` defaults to 0 on many systems, so the kernel writes nothing. Raise it in whatever starts the process (`ulimit -c unlimited`) or, if you cannot change the launcher, at the top of the program itself with `resource.setrlimit(resource.RLIMIT_CORE, ...)` - raising the soft limit up to the hard limit needs no privileges. Then find out where cores go. On Linux the destination is the kernel-wide `/proc/sys/kernel/core_pattern`, and on most modern distributions it is not a filename at all but a pipe into a collector daemon, so looking in the working directory finds nothing while the core is sitting in the collector's store. On macOS cores land in `/cores`, and the limit still has to be raised. This is a host setting, not a per-process one, so it is worth confirming rather than assuming. Other reasons a core goes missing: the filesystem is full or read-only, the process changed its user id, or the container the process runs in has no visibility of where the host kernel wrote the file. ## Step 2: keep the exact binaries A core file is only meaningful together with the binaries that produced it. Open it with a *different* build of the interpreter, or against an extension that has since been rebuilt, and the debugger will map addresses to the wrong symbols and quietly show you a fictional stack. Capture the interpreter path, its exact version and build, and the extension binaries alongside the core - in an environment that redeploys frequently, the binaries can be gone within the hour. ## Step 3: debug symbols Symbols are the difference between a usable backtrace and a column of `??`. You need them in two places: for the interpreter itself, which usually means installing the matching debuginfo package or running a build that was not stripped, and for the extension, which means a binary compiled with `-g` and not stripped. Wheels are routinely stripped for size, so the practical move is often to rebuild the extension locally with debug information and reproduce against that. Also worth knowing: the interpreter's own compiler optimisations can inline frames, so a release build's C backtrace is approximate even with symbols present. A debug build of the interpreter additionally turns on internal assertions, which frequently catch the real error - a bad reference count, an object used after free - earlier and much closer to its cause than the eventual segfault does. ## Step 4: get Python frames out of C frames The C backtrace shows the interpreter's evaluation machinery, which is true but not useful on its own: you want to know which Python function was executing. CPython ships a gdb helper script installed next to the interpreter for exactly this. When gdb loads it, you gain commands that walk the interpreter's C stack, recognise its evaluation frames, and print the Python-level call stack from them - a Python backtrace, the local variables of a chosen Python frame, and movement up and down the Python stack rather than the C one. If gdb refuses to load the helper, it is nearly always the auto-load safe-path setting, and the helper can be sourced explicitly instead. If the Python backtrace comes out empty while the C one is fine, the usual cause is missing interpreter symbols: the helper needs to read the interpreter's internal structures by name. ## Step 5: read the two stacks together The pair is what tells the story. The Python stack says which call your code made; the C stack says what the extension did with it. A fault in an extension entered directly from the Python frame you see is a straightforward argument or lifetime problem. A fault deep in allocation or garbage-collection machinery, with no obvious relationship to the Python frame, is the signature of memory corrupted earlier by someone else - the crash site is not the bug site, and you should be looking at what changed rather than at the frame you landed in. ## A cause worth suspecting in this scenario When a crash starts after a deployment that reused a cached build directory, suspect a stale compiled extension. Extension filenames carry an ABI tag - `importlib.machinery.EXTENSION_SUFFIXES` shows the shape the current interpreter accepts - and a binary compiled against a different interpreter minor version or a different C library will happily load in some setups and then fault on the first call into it. Rebuild from a clean tree before you spend a day in a debugger. ## What core files do not cover A core is a corpse. For a process that is still alive but misbehaving, 3.14 added a remote-debugging entry point (`sys.remote_exec`, PEP 768) that injects a script into a running interpreter at a safe point - a different tool for a different failure. Reach for a core when the process is already gone.

  • The native backtrace is mostly question marks and the Python backtrace is empty. What is missing?
    Debug symbols. The interpreter needs its matching debuginfo installed or an unstripped build, and the extension needs to have been compiled with debug information and not stripped - published wheels usually are. The Python-level backtrace fails specifically when the interpreter's symbols are absent, because the helper script reads the interpreter's internal structures by name. A second possibility is that gdb never loaded the helper at all, which is usually its auto-load safe-path setting.
  • You reproduce the crash but no core file appears anywhere. What are the usual reasons?
    The soft core-size limit is 0, which is the common default; the kernel's core destination is a pipe to a collector rather than a path, so the file is in the collector's store; the target directory is full, read-only, or not writable by the process user; or the process changed its user id, which disables dumping. In a container, the host kernel's setting decides the destination, so a core written on the host is invisible from inside.
  • The crash appeared right after a redeploy that reused a cached build directory. What would you suspect first?
    A stale compiled extension binary built against a different interpreter version or a different C library. Extension filenames carry an ABI tag, so a genuine mismatch is often rejected at import - but a binary that matches the tag while linking against a changed dependency will load and then fault on first use. Rebuild from a clean tree and reproduce before spending time in the debugger; it is far cheaper than a core-file investigation.
  • The C stack shows the fault deep inside allocation machinery, unrelated to your call. What does that tell you?
    That the memory was corrupted earlier and the fault is only where the damage was discovered. The classic cause is a reference-count error in an extension: an object is freed while still referenced, and the crash happens later when that block is reused. Chase what changed rather than the frame you landed in, and reproduce under a debug build of the interpreter, whose internal assertions usually fire much closer to the real bug.

saying these in an interview costs you the question

  • Hunts for the bug in Python code with no native evidence
  • Opens the core against a different interpreter build
  • Accepts a backtrace full of question marks as unreadable and stops
  • Assumes the core-size limit is already unlimited
  • Thinks a core file contains the Python traceback as text
  • Adds a supervisor restart loop and calls it fixed

context