skip to content

Aborts From Native Code

When the process dies with no Python-level story at all: a segfault or abort inside a compiled extension, what a core file still gives you, and how to get a Python stack out of one.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

3

Why does a segfault inside a compiled Python extension module leave no traceback?

level: middleimportance: must knowfreq 30%

answer

  1. The interpreter never gets a turn
  2. Death happens below Python, not in it
  3. A signal kills; an exception unwinds
  4. Shell status is 128 plus signal
  5. 139 is SIGSEGV, 134 is SIGABRT

basics

~20 s

A SIGSEGV kills the process at the operating-system level before CPython can run any Python code, so there is no unwinding, no except clause and no traceback - only a shell status of 139 and possibly a core file.

solid answer

~50 s

An exception is interpreter machinery: CPython builds an exception object, unwinds frames, runs `finally` blocks and prints a traceback from those frames. A segmentation fault is not an exception. The kernel delivers `SIGSEGV` because the process dereferenced an address it may not touch, and the default disposition of that signal is immediate termination plus an optional core file - the interpreter never regains control, so nothing Python-level runs. A compiled extension module is machine code loaded into the interpreter's own address space with no isolation, so a bad pointer or a refcount mistake in it corrupts the same heap the interpreter is using. What you get instead of a traceback is the exit status: a shell reports 128 plus the signal number, so 139 for SIGSEGV and 134 for SIGABRT, and `subprocess.run(...).returncode` reports the negative signal number directly.

code

python · 10 lines
python
import signal
import subprocess
import sys

crash = "import ctypes; ctypes.string_at(0)"
done = subprocess.run([sys.executable, "-c", crash], capture_output=True)

print("returncode:", done.returncode)          # -11
print("stderr empty:", done.stderr == b"")     # True: no traceback
print("signal:", signal.Signals(-done.returncode).name)  # SIGSEGV

go deeper

for a junior

Recall the shape of the failure: no traceback at all, the process simply gone, and a strange exit status instead of an error message. Knowing that this means a crash below Python rather than a bug in your own function is enough here.

for a middle

Be ready to explain the mechanism: an exception is interpreter machinery, a signal is not, and an extension module runs in the interpreter's own address space. Name the status convention of 128 plus signal number and what 139 and 134 each indicate.

for a senior

Show the diagnostic reflex. Isolate the crash in a subprocess, read the exit status rather than the logs, treat the last visible output as stale because of buffering, and rule out a partially written output file before restarting anything.

for a principal

Own the policy question: which workloads may load native extensions at all, how you contain a crash-prone one in a separate process so the blast radius is a worker rather than a service, and what your restart path assumes about half-written state.

## Two different kinds of death When Python code goes wrong, CPython raises an *exception object*, unwinds the frame stack, runs `finally` blocks and context-manager `__exit__` methods on the way out, and finally prints a traceback assembled from the frame objects it walked. Every part of that is interpreter machinery, and it needs a running, self-consistent interpreter to happen at all. A segmentation fault is not that. `SIGSEGV` is delivered by the kernel because the process touched an address it is not allowed to touch - a null pointer, a block that was already freed, memory past the end of a buffer. The default disposition of that signal is to terminate the process immediately and, if the core-size limit allows it, write a core file. No Python bytecode executes after the faulting instruction, so there is no unwinding, no `finally`, no `atexit` handler, and no traceback. The same is true of `SIGABRT`, which is what a native library raises when an assertion fails or a C++ exception escapes. ## Why an extension module can do this at all A compiled extension module is a shared object (a `.so` on Unix, a `.pyd` on Windows) that the import system loads into the *same address space* as the interpreter. There is no sandbox between them. That is exactly why extensions are fast, and exactly why a mistake in one is fatal to the whole process rather than raisable as a Python error. The classic causes are worth recognising by shape: * **Reference-counting errors.** One `Py_DECREF` too many frees an object that Python still names. The crash then happens later, in unrelated code, when that memory is reused - the crash site is not the bug site. * **C API calls without holding the GIL**, or state shared between threads without a lock. Free-threaded builds (experimental in 3.13, officially supported in 3.14 under PEP 779) remove the incidental serialisation that used to hide these bugs, so extensions that were quietly thread-unsafe start crashing. * **Raw pointer arithmetic** over a buffer, including via `ctypes`, where a wrong length or a wrong type reads off the end of an allocation. ## What you get instead of a traceback The exit status. A Unix shell reports a signal death as 128 plus the signal number: 139 for `SIGSEGV` (11), 134 for `SIGABRT` (6), 137 for `SIGKILL` (9). From Python, `subprocess.run(...).returncode` is the *negative* signal number, and `signal.Signals(-returncode).name` turns it into a name. Distinguishing these matters: 139 points at a bad pointer inside native code, while 137 usually means something outside the process decided to kill it. `SIGABRT` often comes with a parting message on stderr - a library's assertion text, or a line beginning `Fatal Python error:` - and that message is worth reading before anything else, because it names the component that gave up. ## The output you see is not the last thing that happened When stdout is redirected to a file or a pipe it is block-buffered, so several kilobytes of the most recent output die in the buffer with the process. Engineers routinely conclude that the crash happened right after the last line they can see; it did not. When you are chasing a native crash, log to stderr or run unbuffered, or you will be reading a stale picture of where the process got to. ## Why `except` cannot help `except Exception` catches exception objects, and no exception object exists here. Installing a Python-level handler for `SIGSEGV` with `signal.signal` does not rescue you either: CPython only runs Python signal handlers between bytecode instructions, and a faulting process typically cannot get back to the eval loop to reach one. Treating a segfault as something to retry around is a category error - the process image is already suspect, and continuing risks writing corrupted data. ## What the first diagnostic moves are Reproduce with the suspect work isolated in a subprocess so you can read the exit status cleanly. Note which extension modules the workload had imported, and whether the crash correlates with a version bump of one of them. Then move to the tooling that *can* see below Python: a native backtrace from a core file, and the interpreter's own gdb helper to translate the C frames back into Python frames. Everything after that step is native debugging, and it starts from accepting that this failure never had a Python-level story to tell.

  • How do you turn a process exit status into the signal that killed it?
    A shell reports 128 plus the signal number, so 139 means signal 11. From Python, `subprocess.run(...).returncode` is already the negative signal number, and `signal.Signals(-returncode).name` gives you the name. Beware the ambiguity: a program that calls `sys.exit(139)` produces the same shell status as a segfault, which is one reason the negative return code from `subprocess` is the more trustworthy signal.
  • Do context managers and buffered writes get cleaned up when a process dies from SIGSEGV?
    No. `__exit__` methods never run, `finally` blocks never run, `atexit` callbacks never run, and anything sitting in Python's or C's I/O buffers is lost. Temporary files are not removed and half-written output files are left as they were. This is why a native crash in a writer process is a data-integrity event, not just an availability one - assume the last records are missing or truncated.
  • The crash points at one extension module, but that module has been stable for months. Where do you look?
    At refcount errors, which crash far from their cause: a decref too many frees an object that is still named, and the fault appears later when that memory is reused by whatever code happens to allocate next. So the module in the backtrace is often the victim, not the culprit. Look at what changed - an interpreter upgrade, a rebuilt binary, a newly added extension sharing the process - rather than at the frame you landed in.

saying these in an interview costs you the question

  • Says the crash can be caught with except Exception
  • Expects finally blocks or atexit handlers to still run
  • Assumes any crash without a traceback means out of memory
  • Reads exit status 139 as an application-defined error code
  • Trusts the last logged line as the crash location
  • Proposes retrying the operation inside the same process

context

open as a page

A clinical-lab result loader dies with SIGSEGV inside a compiled extension - how do you capture a core file and read a Python stack from it?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Raise the core-size limit, find where the kernel writes cores, reproduce the crash, then open the core in gdb against the exact interpreter binary that produced it and use CPython's helper to print Python frames.

open as a page

What does Python's os.abort() do, and how does it differ from sys.exit()?

level: juniorimportance: nice to knowfreq 14%

basics

~20 s

os.abort() raises SIGABRT and ends the process in the hardest way available: no exception, no finally blocks, no atexit callbacks, exit status 134, and a core file if the limit allows. sys.exit() merely raises SystemExit, which unwinds normally.

open as a page