skip to content

Process and Pool

Spawning workers with multiprocessing.Process or fanning work out through Pool — the standard answer to using every core. Interviewers check that you can size it and collect results without hanging.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What do multiprocessing.Process.start() and join() do?

level: juniorimportance: must knowfreq 65%

answer

  1. Two calls, two different jobs
  2. One creates, one waits
  3. There is a method that does not spawn
  4. Status attribute is None until it ends
  5. join(timeout) tells you nothing by itself

basics

~20 s

start() launches a new operating-system process that runs the target callable; join() blocks the caller until that child has exited and reaps it. Calling run() instead executes the work in the current process and spawns nothing.

solid answer

~40 s

Constructing a `multiprocessing.Process` only records the target and its arguments — no process exists yet, and `pid`, `is_alive()` and `exitcode` reflect that. `start()` creates the child and returns immediately in the parent; the child calls `run()`, which invokes the target. `join()` blocks until the child exits and reaps it; `join(timeout)` waits at most that long and returns `None` regardless, so you check `is_alive()` or `exitcode` afterwards to learn what happened. `exitcode` is `None` until exit, then `0` for a clean return, the value given to `sys.exit(n)`, `1` for an uncaught exception in the child, or a negative signal number if it was killed. `terminate()` kills rather than shuts down: the child's `finally` blocks never run, and you still have to `join()` it.

code

python · 16 lines
python
import multiprocessing as mp
import os


def work(n):
    print("child", os.getpid(), "n =", n)


if __name__ == "__main__":
    p = mp.Process(target=work, args=(3,))
    print("before start:", p.pid, p.exitcode, p.is_alive())
    p.start()
    print("after start :", p.pid, p.exitcode, p.is_alive())
    p.join()
    print("after join  :", p.pid, p.exitcode, p.is_alive())
    print("parent", os.getpid())

go deeper

for a junior

Be ready to say plainly that one call creates the process and the other waits for it, and to spot the trap of invoking the worker method directly instead of starting a process.

for a middle

Explain the state machine: pid and exitcode before and after each call, why join(timeout) returns nothing useful, and why the parent must read exitcode instead of assuming a finished child succeeded.

for a senior

Show the production habits: bounded joins with a follow-up liveness check, escalation from a cooperative stop to terminate, and knowing that a killed child skips its finally blocks and can leave shared resources mid-update.

for a principal

Own the choice of granularity. Hand-rolled processes suit a few long-lived workers with distinct roles; per-item work belongs in a pool. Argue the process-creation cost and the supervision burden your team has to carry.

`multiprocessing.Process` is Python's wrapper around an operating-system process, and the interview turns on four observable states: **constructed, started, running, exited**. ## The lifecycle, call by call ### Construction binds; it does not launch `p = multiprocessing.Process(target=work, args=(3,))` records a callable and its arguments on an ordinary Python object living in the parent. Nothing has happened at the OS level: - `p.pid` is `None`, - `p.is_alive()` is `False`, - `p.exitcode` is `None`. ### `start()` creates the child It asks the interpreter to create a new OS process and arranges for that child to call `p.run()`, which is what actually invokes `target(*args, **kwargs)`. `start()` returns in the parent as soon as the child exists — it does **not** wait for the work. From that instant there are **two independent interpreters with two independent heaps**: the child's assignments to module globals are invisible to the parent, and anything the child computes has to travel back over an explicit channel. `start()` may be called only once per object; a second call raises. The single most common beginner error is calling `p.run()` instead of `p.start()`. `run()` is a plain method — it executes the target synchronously in whichever process calls it. The program prints the right answer, finishes, and achieves exactly zero parallelism. ### `join()` waits and reaps `p.join()` blocks the calling process until the child has exited, then collects its exit status so the child does not linger as a **zombie**. - `p.join(timeout)` waits at most that many seconds and **returns `None` either way** — it is not a boolean and it does not raise on timeout, so the only way to tell whether the wait succeeded is to inspect `p.is_alive()` or `p.exitcode` afterwards. - Joining a process that was never started raises `AssertionError`. You are not strictly obliged to join: non-daemonic children are joined automatically when the parent interpreter shuts down. But leaning on that means you cannot observe success or failure at the point where you need it, and it postpones every error to process exit. ### `exitcode` reports how it ended It is: - `None` while the child has not exited; - `0` after a clean return; - the integer passed to `sys.exit(n)`; - `1` when the target raised an uncaught exception — note that the traceback is printed by the child on its own stderr and is **not** re-raised in the parent; - and a negative number when the OS killed the child with a signal. Because it is only meaningful after exit, the idiomatic sequence is `p.join()` then `if p.exitcode: ...`. ## `terminate()` is a kill, not a shutdown It signals the child to die immediately (SIGTERM on Unix, the equivalent on Windows). The child does not unwind: `finally` blocks, `atexit` handlers and context-manager exits inside the child never run, and anything it held half-updated stays half-updated. After `terminate()` you must still `join()` to reap it, and its `exitcode` will be negative. `kill()` is the harsher variant. Treat both as the escalation path after a cooperative stop has failed, never as ordinary shutdown. ## Daemonic children Setting `p.daemon = True` *before* `start()` marks the child as **expendable**: the parent kills it on exit instead of joining it, and the child is forbidden to create children of its own. That suits a background helper whose partial work is worthless, and suits nothing whose result you intend to read. ## The shape interviewers want to see 1. Build the `Process`, 2. `start()` it, 3. do something useful in the parent or start siblings, 4. `join()` each one, 5. then check each `exitcode` before trusting the outcome. If you need many short-lived workers rather than a handful of long-lived ones, a pool is the right tool — hand-rolling a `Process` per work item pays the process-creation cost every time. One last thing worth saying out loud in an interview: a `Process` is a *process*, not a thread. There is no shared address space, so "just append to a global list in the worker" silently does nothing useful. Results have to come back through a channel designed for it, and that is a deliberate cost you accept in exchange for real parallelism across cores.

  • What is the difference between calling p.start() and calling p.run()?
    `start()` creates a new operating-system process and arranges for it to invoke `run()`. `run()` is just a method: calling it directly executes the target synchronously in the current process, so the code appears correct and delivers no parallelism at all. `run()` exists so subclasses can override it as the body of the child; you override `run`, you call `start`.
  • If you never call join(), does the child still finish?
    Usually yes. A non-daemonic child is joined automatically when the parent interpreter shuts down, so its work completes before the program exits. But you learn nothing at the point where you needed the answer: `exitcode` stays `None` while it runs, failures surface only at interpreter exit, and a daemonic child is killed outright rather than awaited. Join when the result or the status matters.
  • How do you tell whether a child failed rather than succeeded?
    Join it, then read `exitcode`: `0` is a clean return, `1` means the target raised an uncaught exception (the traceback went to the child's own stderr and was never re-raised in the parent), a positive value is whatever `sys.exit(n)` was given, and a negative value means a signal killed it. Treating a finished process as a successful one is the bug.

start() is posting a letter and walking away; join() is standing at the mailbox until the reply arrives. Calling run() is writing the reply yourself and pretending it came back.

saying these in an interview costs you the question

  • Calling p.run() to start a worker, which runs it in the caller
  • Expecting exitcode to be 0 right after start()
  • Believing join(timeout) returns True or False
  • Thinking join() kills a child that overruns its timeout
  • Using terminate() as normal shutdown and expecting finally blocks to run
  • Assuming the child's writes to globals are visible in the parent

context

open as a page

How do multiprocessing.Pool's map, imap and imap_unordered differ?

level: middleimportance: must knowfreq 55%

basics

~20 s

map() blocks and returns a list of every result in input order. imap() returns a lazy iterator that yields results in input order as they arrive. imap_unordered() yields each result the moment it is ready, in completion order.

open as a page

What does the chunksize argument to multiprocessing.Pool.map control?

level: middleimportance: should knowfreq 40%

basics

~20 s

chunksize is how many items of the iterable are batched into a single task handed to one worker. Larger chunks mean fewer dispatches and less per-item overhead; smaller chunks spread work more evenly and shorten the idle tail.

open as a page

Why can a with-block around multiprocessing.Pool lose pending results?

level: seniorimportance: nice to knowfreq 25%

basics

~10 s

Pool.exit calls terminate(), not close() then join(). Leaving the with-block while apply_async or map_async submissions are still outstanding kills the workers immediately, so those tasks never run and their AsyncResult objects are never fulfilled.

open as a page