skip to content

What does a negative multiprocessing.Process.exitcode mean?

level: juniorimportance: should knowfreq 40%

answer

  1. Three shapes, plus a fourth before termination
  2. The sign carries the meaning
  3. Positive means the child chose
  4. Negative means the OS chose
  5. -9 and -11 are the usual suspects

basics

~20 s

A negative exitcode of -N means the child was killed by signal N and ran no cleanup. None means it has not finished, 0 means a clean exit, and a positive number is the status the child chose.

solid answer

~50 s

`multiprocessing.Process.exitcode` is `None` until the child has been started and has terminated. After `join()` it holds one of three shapes. Zero means the child ran to completion normally. A positive integer is the status the child chose: whatever it passed to `sys.exit()`, with an uncaught exception in the target function becoming `1` after multiprocessing prints the child's traceback. A negative integer `-N` means the child did not choose at all — the operating system killed it with signal `N`, so nothing in the child ran afterwards: no `finally` blocks, no `atexit` handlers, no flushing of buffered writes. The classic negative codes you see in production are `-9` (the process was hard-killed, typically by an out-of-memory killer or a `Process.kill()` call) and `-11` (a segmentation fault inside a native extension). So negative always means "died", never "failed".

code

python · 14 lines
python
import multiprocessing as mp
import os
import signal


def crash():
    os.kill(os.getpid(), signal.SIGKILL)


if __name__ == "__main__":
    p = mp.Process(target=crash)
    p.start()
    p.join()
    print(p.exitcode)

go deeper

for a junior

Recall the four states: None before it finishes, 0 for a clean exit, a positive number the child chose, and a negative number meaning a signal killed it. Always join() before you read the attribute.

for a middle

Explain why negative means no cleanup ran — no finally, no atexit, no flushed output — and name the two you actually see: -9 for a hard kill such as an out-of-memory kill, -11 for a segfault in a native extension.

for a senior

Show how you branch on the sign in a supervisor: positive sends you to the child's own traceback, negative sends you to memory limits, native code or whoever sent the signal. Own the cleanup the dead child could not do.

for a principal

Argue for the invariant that makes crashes survivable: tasks idempotent enough to retry after a kill at an unknown point, and cleanup owned by the parent or the platform rather than by code inside the process that dies.

### The three shapes of an exit status When you start a `multiprocessing.Process`, the parent keeps a handle on a real operating-system child process. `Process.exitcode` is the parent's view of how that child ended, and it is only meaningful once the child has actually terminated — before `start()`, and while the child is still running, the attribute reads `None`. Calling `join()` (or polling `is_alive()`) is what makes the value settle. Once the child is gone, the value follows the underlying wait-status convention, flattened into a single signed integer: * **`0`** — the child ran its target function to the end and exited normally. * **a positive integer** — the child terminated itself with that status. `sys.exit(3)` inside the child yields `3`. An uncaught exception in the target function yields `1`: multiprocessing catches it in the child's bootstrap, prints the child's traceback to that child's stderr, and exits with status 1. Note that the exception object itself does not travel back through `Process` — only the number does. `Process` gives you no result channel, so any diagnosis has to come from the child's own output or from a queue you set up yourself. * **a negative integer `-N`** — the child was killed by signal number `N` and never got to run any Python code afterwards. ### Why negative is the interesting case The negative case is the one interviewers are actually probing, because it is the only one that means the child had no say. Nothing in the child unwound: `try/finally` blocks did not run, context managers did not exit, `atexit` callbacks did not fire, buffered `print` output was never flushed, temporary files were never removed, and any lock the child held in shared state stayed held. That is why a crashed worker so often leaves the rest of the system in a strange state rather than simply producing a missing result. The two negatives you meet most often in real systems are: * **`-9`** — the child was hard-killed. On Linux this is very frequently the kernel's out-of-memory killer choosing the fattest process on the box, which under memory pressure is usually a worker. It is also what `Process.kill()` does, and what a container runtime does when a memory limit is exceeded. * **`-11`** — a segmentation fault, which in Python effectively always means a bug in a compiled extension module or in the C library underneath it, since pure Python cannot segfault by itself. A `-15` normally means somebody asked politely first: that is what `Process.terminate()` sends, and it is also the signal an orchestrator sends before escalating to a hard kill. ### Reading it correctly Two mistakes are common. The first is treating `exitcode` as a *result*: it is a disposition, not a return value, and a child that computed the wrong answer perfectly happily still exits `0`. The second is checking it too early. `exitcode` is `None` while the child lives, so `if p.exitcode:` before a `join()` silently reads as false and hides the failure. Join first (optionally with a timeout), then inspect. A useful pattern in a supervisor loop is to branch on the sign rather than the value: zero means done, positive means the child raised or exited deliberately and its own traceback is in the log, negative means the child was killed from outside and you should be looking at memory limits, native extensions, or whoever sent the signal — not at your Python traceback, because there will not be one. ### Cleanup is a separate concern Because a negative exit means no in-child cleanup ran, the parent has to own anything that must survive a kill. That means the parent removing scratch files it can name, the parent releasing shared resources, and the parent deciding whether the task is safe to retry. It also means designing tasks to be idempotent, since "killed at an unknown point" is exactly the state where a half-applied side effect is possible.

  • What exitcode does a child get when its target function raises an uncaught exception?
    `1`. The child's bootstrap catches the exception, prints that child's traceback to its own stderr, and exits with status 1. The exception object does not come back to the parent through `multiprocessing.Process` — there is no result channel on a bare `Process`, so if the parent needs the error it has to travel over a queue or a pipe you set up, or you should be using a pool, whose result path carries exceptions for you.
  • If the parent process is hard-killed, what happens to its children?
    Non-daemonic children keep running and are reparented to the init process — they become orphans that nothing joins. Setting `daemon=True` before `start()` makes multiprocessing terminate the child when the parent exits *normally*, via an exit handler, but that handler cannot run if the parent itself was killed by a signal. Surviving that reliably needs something outside Python: a process group the supervisor signals as a unit, or a container that reaps everything inside it.
  • How do you check whether any children are still alive from the parent?
    `multiprocessing.active_children()` returns the list of the parent's live children, and calling it has the side effect of joining any that have already finished, which clears them out. It only knows about children this process started through multiprocessing, so it is a hygiene check inside one parent, not a system-wide process listing.

A positive exit code is an employee handing in a resignation letter with a reason on it; a negative one is finding their desk empty and the building security log showing they were escorted out.

saying these in an interview costs you the question

  • Thinking exitcode carries the child's return value
  • Reading exitcode before join(), while it is still None
  • Assuming a negative code means the child raised an exception
  • Believing finally blocks run when a child is signal-killed
  • Treating -9 as a Python bug rather than an external kill

context