skip to content

A Python worker pool renders no more invoices — how do you read an all-threads stack dump to find the stuck frame?

level: seniorimportance: must knowfreq 52%

answer

  1. Sample twice, never once
  2. Compare stacks against process CPU
  3. Frames are printed innermost first
  4. Waiting workers are rarely the culprit
  5. Find the thread that is not waiting

basics

~10 s

Take two dumps a few seconds apart. Identical stacks at near-zero CPU means blocked; a pinned core means spinning. Then read each worker's innermost frames and find the one thread that is not waiting.

solid answer

~50 s

One dump is a photograph; two dumps are a diagnosis. I take an all-threads dump, wait five to ten seconds, take a second, and diff them against the process's CPU usage. Identical stacks at near-zero CPU means the threads are genuinely blocked, not slow. Then I read the innermost frames per worker: `threading.py` in `wait` beneath `queue.py` in `get` means the worker is idle and the producer is the suspect, not the pool; `threading.py` in a lock acquisition beneath application code means contention, and the interesting thread is the one that is *not* waiting — usually still inside a loop it should have left. That last thread is where a boundary bug lives: a page-index loop whose exit condition is off by one never terminates, so it holds the lock and every other worker stacks up behind it. Remember the dump shows Python frames only.

code

python · 20 lines
python
import faulthandler
import os
import queue
import signal
import threading
import time

jobs = queue.Queue()


def worker():
    jobs.get()


faulthandler.register(signal.SIGUSR1, all_threads=True)
for n in range(2):
    threading.Thread(target=worker, name=f"render-{n}", daemon=True).start()
time.sleep(0.2)
os.kill(os.getpid(), signal.SIGUSR1)
time.sleep(0.2)

go deeper

for a junior

Know that a stack dump lists each thread's calls innermost-first, and that a thread sitting in a queue read is waiting for work rather than causing the problem. Reading one dump top to bottom is the skill being checked.

for a middle

Explain how to separate blocked from spinning: two dumps compared against the process's CPU usage, and what the frames from the standard library beneath your own code are telling you about what the thread is waiting on.

for a senior

Show the full triage — sample twice, cross-check the thread inventory, find the thread that is not waiting, and state the limits of the evidence. Being explicit that the dump narrows the search rather than naming the bug is what separates this level.

for a principal

Own the practice, not the incident: a standing rule that two dumps are captured before any restart, dumps landing somewhere durable, worker threads named by role, and per-thread CPU available so nobody has to guess which thread is the spinner.

## The method, before the details A wedged process has three plausible states and you can separate them with two dumps and one CPU reading: | CPU | Two dumps compared | Diagnosis | |---|---|---| | ~0% | identical | Threads are **blocked** — waiting on a lock, a condition, a socket, a child process | | one core pinned | identical *or* moving | A thread is **spinning** — a loop that never exits, or a tight C call | | busy, spread | moving | The process is **working**, just slower than you expected | The single most common mistake is taking one dump, seeing everything parked in `wait`, and announcing a deadlock. One sample cannot distinguish stopped from slow. Take the second. ## Reading a worker's frames A `faulthandler` all-threads dump prints, per thread, the id and (on 3.14) the `Thread.name`, then frames **innermost first**. So the top line is where the thread is *now*, and the lines below are how it got there. Common innermost signatures: - **`threading.py` in `wait`, under `queue.py` in `get`.** The worker is idle, parked on an empty work queue. This is a healthy worker in an unhealthy process: if all eight look like this, the pool is not stuck at all — whatever feeds it has stopped. Go and look for the feeder thread; if it is missing from `threading.enumerate()`, it died with an exception you never logged. - **`threading.py` in `wait` or in a lock acquisition, under your own code.** Genuine contention. Now find the thread that is *not* waiting, because it is the holder. - **A socket read, a subprocess wait, a file read.** Blocked on the outside world. The dump tells you *which* external dependency; the fix is a timeout, not a lock change. - **The same application frame on every worker, with the process burning CPU.** Not a lock at all: a hot loop, or a data-dependent path that got quadratic. ## The subtle case: identical stacks at 100% CPU This looks like a contradiction and it is the case worth being able to explain. The dump prints a **line number**, not a bytecode offset. A thread grinding through a tight loop whose body spans two lines can produce byte-identical dumps seconds apart while consuming a full core. So never conclude "blocked" from identical stacks alone — pair the dumps with the CPU reading. Equally, a thread inside a single long-running C call will show one unchanging Python frame no matter how much work it is doing, because the dump cannot see into the extension. ## Working the invoice-renderer example A renderer's pool of eight workers stops completing jobs. The dumps show seven workers parked in a lock acquisition under `render_document`, and one worker whose innermost frame is a `while` line in `paginate` — the same line in both dumps, with one core pinned. That combination reads cleanly: the eighth worker entered a pagination loop whose exit condition is off by one at the final page boundary, so it never leaves; it holds the lock the other seven want; the pool's throughput drops to zero without a single error being logged. The 2.4 GB working set is a corroborating detail rather than the cause — it stops growing when the pool stops, which tells you the process is not still allocating, it is stalled. Notice what the dump did **not** tell you: it did not say the loop is wrong, only that the process is inside it and staying there. The dump narrows a whole service to one function; you still read that function. ## Cross-checks that make the reading solid - **`threading.enumerate()` alongside the dump.** Confirms the expected worker count and gives you `ident`-to-name mapping. A worker missing from both is a worker that died. - **Per-thread CPU from the operating system**, joined on `Thread.native_id`, tells you exactly which thread is the spinner rather than making you infer it. - **A third dump.** If a stack is unchanged across three samples spread over a minute, "it is just slow" stops being a credible story. - **Resident memory over the same interval.** Flat memory with flat stacks is a stall; climbing memory with flat stacks means something is still accumulating on your behalf. ## Limits to state out loud The dump covers Python frames in this interpreter only. Native frames are invisible — a thread deep inside a compiled extension shows only the Python frame that called in, so if every dump points at the same call into an extension, the next tool is an operating-system-level debugger, not another dump. And a thread that never touched Python at all will not appear with frames worth reading.

  • Two dumps thirty seconds apart are identical, yet the process is pinning a full core. How is that possible?
    Two ways. The dump prints line numbers, not bytecode offsets, so a thread grinding round a tight loop can show the same line in every sample while burning CPU. Or the thread is inside a single long call in a compiled extension, which the dump cannot see into and therefore renders as one unchanging Python frame. Identical stacks mean *not moving between frames*, not *not running* — always pair them with a CPU reading.
  • All eight workers are parked in queue.Queue.get. What is your next step?
    Stop looking at the pool: those workers are idle and healthy. The failure is upstream, in whatever feeds the queue. Check whether the producer thread is still in `threading.enumerate()` — a thread that raised out of its target vanishes silently while the process keeps running — and whether its exception output was swallowed. If the producer is alive, read *its* frames: it is probably blocked on the source it reads from.
  • The innermost visible frame of the stuck worker is a call into a compiled extension. What does the dump still owe you, and what do you reach for next?
    The dump has done all it can: it shows Python frames, so it names the call into the extension but nothing beyond it. That is still valuable — it identifies the dependency and the arguments' origin. Beyond that you need an operating-system-level debugger attached to the process, or per-thread CPU accounting to say whether the thread is computing or blocked inside that call.

A single stack dump is a group photograph: it shows who was standing where. Two photographs a few seconds apart show who was actually moving.

saying these in an interview costs you the question

  • Calls it a deadlock from a single dump
  • Assumes a worker parked in queue.Queue.get is the bug
  • Reads the innermost line as where the bug was written
  • Expects C-extension frames to appear in the dump
  • Ignores process CPU when separating a spin from a block
  • Restarts the process before taking any dump at all

context