skip to content

Why does an exception inside a ThreadPoolExecutor job stay hidden until Future.result()?

level: juniorimportance: must knowfreq 72%

answer

  1. the failure has to cross threads
  2. the worker stores its outcome somewhere
  3. one accessor re-raises, one returns
  4. done() covers failed and cancelled too
  5. an unread Future is an unread failure

basics

~10 s

The executor catches whatever the submitted callable raises and stores it on that job's Future instead of letting it propagate. Nothing surfaces until you call Future.result(), which re-raises it, or Future.exception(), which returns it.

solid answer

~40 s

`Executor.submit` returns a `Future` immediately and the callable runs later on a worker thread, so a raise there cannot unwind your stack. `concurrent.futures` therefore wraps every call: on success it stores the value on the `Future`, on failure it stores the exception. `Future.result()` blocks until the job is finished and then re-raises the stored exception in the calling thread; `Future.exception()` blocks the same way but returns the exception object instead of raising. `Future.done()` only means "finished" — it is true for succeeded, failed and cancelled jobs alike. The classic production bug is fire-and-forget: code submits jobs, never keeps the returned `Future`, and every failure disappears, because exiting the `with` block calls `Executor.shutdown(wait=True)`, which waits for the work but never inspects a single result.

code

python · 15 lines
python
from concurrent.futures import ThreadPoolExecutor

def charge(account_id):
    if account_id == 3:
        raise ValueError("card declined")
    return account_id * 10

with ThreadPoolExecutor(max_workers=4) as pool:
    futures = [pool.submit(charge, i) for i in range(5)]
    print("submitted; no exception yet")
    for fut in futures:
        if fut.exception() is None:
            print("ok", fut.result())
        else:
            print("failed", type(fut.exception()).__name__, fut.exception())

go deeper

for a junior

Recall the two-step shape: submit returns a Future right away, and the outcome — value or exception — is read back later with result(). Be ready to say out loud that ignoring the returned Future means ignoring the failure.

for a middle

Explain the mechanics: the work item catches BaseException and stores it, result() re-raises in the calling thread while exception() returns the object, and done() covers success, failure and cancellation alike.

for a senior

Show how you make failures visible in a real batch: collect futures, drain them, count and log per-item failures, and set an exit status. Interviewers expect you to spot the fire-and-forget submit loop in a code review immediately.

for a principal

Own the policy question: does a batch that loses 2% of its items report success? Decide where partial failure is aggregated, what the process exit code and alerting mean, and make that contract the same across every pool in the codebase.

`Executor.submit(fn, *args)` does two separate things: it puts a work item on the executor's internal queue, and it hands you back a `concurrent.futures.Future` — a thread-safe handle to a result that does not exist yet. The call returns in microseconds. `fn` runs later, on one of the pool's worker threads in a `ThreadPoolExecutor`, or in a worker process in a `ProcessPoolExecutor`. ### Why the exception cannot simply propagate A `raise` unwinds the stack of the thread that raised it. Your calling thread is somewhere else entirely by then — it has already returned from `submit` and moved on. There is no stack frame of yours for the worker's exception to travel through. If the pool let the exception escape the worker's run loop, the worker thread would die and the pool would silently shrink until nothing ran at all. So the executor wraps each call. It invokes the callable inside a `try`, and either records the return value on the `Future` or records the exception on it. It catches `BaseException`, not just `Exception`, so even something like a `KeyboardInterrupt` raised inside the callable lands on the `Future` rather than killing the worker. ### Where the exception comes back `Future.result()` blocks until the job reaches a final state, then returns the value or re-raises the stored exception in the *calling* thread. The exception object still carries the worker's traceback, and CPython appends the frames of the `result()` call site, so a printed traceback shows both halves: your loop at the top, the failing worker frames at the bottom. `Future.exception()` is the non-raising counterpart: it blocks the same way and returns the exception object, or `None` when the job succeeded. Both take a `timeout=` argument, and both raise `TimeoutError` if the job has not finished in time. Since Python 3.11, `concurrent.futures.TimeoutError` is an alias of the built-in `TimeoutError`, so a plain `except TimeoutError` catches it; on older versions it was a distinct class. ### `done()` is not `succeeded()` This is the trap juniors fall into. `Future.done()` returns `True` for a job that finished, a job that raised, and a job that was cancelled. `Future.cancelled()` singles out the cancelled case, and `result()` on a cancelled `Future` raises `CancelledError` rather than returning anything. The only way to learn whether a job *succeeded* is to ask for its outcome, with `result()` or `exception()`. ### The bug this design creates in real code ```python with ThreadPoolExecutor(max_workers=8) as pool: for record in records: pool.submit(handle, record) print("done") ``` This prints `done` and exits cleanly whether every job succeeded or every single one raised. `Executor.__exit__` calls `shutdown(wait=True)`, which waits for the queue to drain — it never looks at a result, and there is nowhere for a worker error to go. A batch job written this way reports success while quietly dropping a fraction of its work. The fix is to keep the futures and drain them: ```python futures = [pool.submit(handle, r) for r in records] failures = [] for fut in futures: try: fut.result() except Exception as exc: failures.append(exc) ``` Or, when you only want to log rather than re-raise, use `Future.exception()` and skip the `try` entirely. A third option is `Future.add_done_callback(fn)`: the callback fires when the job finishes, fails or is cancelled, receives the `Future` itself, and must call `result()` or `exception()` to see the outcome. Note that an exception raised *inside* a done-callback is logged by `concurrent.futures` and otherwise swallowed, so callbacks are not a place to put fragile code. ### One extra twist with process pools With a `ProcessPoolExecutor` the exception has to travel back from a child process, which means it is pickled. An exception class that cannot be pickled, or one whose custom `__init__` signature does not survive a round trip, comes back degraded. And if the worker dies outright — a segfault in a C extension, or the OS killing it for memory — there is no exception to store at all: the pool is declared broken and every pending `Future` is completed with a `BrokenProcessPool` error from `concurrent.futures.process`. ### How to say it in an interview One sentence carries it: *the worker cannot raise into my thread, so the pool stores the outcome on the Future and `result()` replays it at the call site — which means an unread Future is an unread failure.* Then name the fire-and-forget bug, because that is what the interviewer is actually probing for.

  • How do you check whether a job failed without blocking on it?
    Ask `Future.done()` first; only when it is true will `Future.exception()` return immediately rather than blocking. You can also pass `timeout=0` to `exception()` and catch `TimeoutError` for the not-finished case. For a push model, register `Future.add_done_callback` — it fires once the job finishes, fails or is cancelled, and the callback reads the outcome off the `Future` it is handed.
  • What happens if the submitted callable raises something that is not an Exception, such as KeyboardInterrupt?
    It is still captured. The executor's work item catches `BaseException`, so a `KeyboardInterrupt` or `SystemExit` raised inside the callable is stored on the `Future` and re-raised by `result()` like anything else. That keeps the worker thread alive, but it also means such a signal raised in a worker will not interrupt your main thread — you only see it when you read the result.
  • Does a cancelled Future look any different from a failed one?
    `done()` is true for both, but `cancelled()` is true only for the cancelled one, and `result()` on it raises `CancelledError` rather than the job's own exception. `exception()` on a cancelled `Future` also raises `CancelledError` instead of returning it, so cancellation is a state you test for, not an exception you read out.

saying these in an interview costs you the question

  • Claims submit() raises the worker's exception at the call site
  • Reads Future.done() as True meaning the job succeeded
  • Submits jobs, keeps no Future, and calls that error handling
  • Thinks exiting the with block re-raises worker exceptions
  • Believes Future.exception() re-raises rather than returns
  • Expects worker tracebacks to print themselves automatically

context