skip to content

How do you profile a live Python image-thumbnail worker without restarting it?

level: seniorimportance: should knowfreq 38%

answer

  1. Never restart the evidence
  2. Attach from outside the process
  3. Read memory, walk interpreter frames
  4. Privileges are a pre-incident decision
  5. Wall-clock first for a waiting worker

basics

~10 s

Attach a sampling profiler from outside the process. An out-of-process sampler reads the target's memory, walks its interpreter frames and aggregates stacks without redeploying, restarting, or adding per-call instrumentation to the hot path.

solid answer

~50 s

Restarting destroys the evidence: the warm caches, the leaked handles and the queue backlog that made the worker slow at its 1,200-request-per-minute peak are all gone the moment it comes back up. So sample it in place. An out-of-process sampler attaches with operating-system facilities, reads the target's memory, walks the interpreter's frame structures and reconstructs Python stacks without the target cooperating — no import, no code change, and overhead borne mostly by the sampler rather than the worker. Take a **wall-clock** profile first, since a leak or a stalled resource shows as waiting rather than as CPU. Interpret it knowing the blind spots: native frames collapse into their calling Python frame, and threads you did not ask for do not appear. Python 3.14 also adds `sys.remote_exec`, a supported way to have a running interpreter execute a script at its next safe point, which gives attach-style tooling a sanctioned entry point.

code

python · 11 lines
python
import faulthandler
import sys

faulthandler.dump_traceback_later(0.2, repeat=True, file=sys.stderr)
try:
    total = 0
    for i in range(20_000_000):
        total += i
finally:
    faulthandler.cancel_dump_traceback_later()
print(total)

go deeper

for a junior

Take away the headline: you can look at a running Python process from outside without stopping it, and restarting a sick process usually destroys the evidence you needed.

for a middle

Explain how attaching works at a mechanical level — a separate process reading memory and reconstructing interpreter frames — and why that costs the target far less than instrumenting every call.

for a senior

Show operational judgement: pick the wall-clock profile for a waiting worker, know the privilege and container prerequisites in advance, collect enough samples, and state the blind spots when you present the finding.

for a principal

Own it as a capability rather than a trick: decide whether remote attach is enabled in production at all, who may use it, how it is audited, and what the standing alternative is where policy forbids it.

### Why not just restart with profiling on The instinct is to redeploy the worker with a profiler enabled and reproduce the problem. On a live system that is usually the wrong move, and the reasons are worth being able to state: - **The state is the bug.** A worker degrading at a 1,200-request-per-minute peak because a resource is left unclosed has been accumulating that state for hours. A restart resets it and the symptom disappears for a while — you have destroyed the only reproduction you had. - **Peak is not reproducible on demand.** The interesting behaviour exists only under the traffic you cannot conjure in staging. - **Per-call instrumentation changes the answer.** Instrumenting every call inflates exactly the call-heavy code you are trying to judge, so even a successful reproduction gives distorted proportions. ### What attaching actually does An out-of-process sampler runs as a separate program. It attaches to the target by process id using operating-system facilities for reading another process's address space, locates the interpreter's thread and frame structures in that memory, and walks them to reconstruct a Python-level call stack — code object names, filenames, line numbers. It repeats that at its sample rate and aggregates the stacks. The target does not import anything, does not run profiler code, and in the common design is not even paused for long: the sampler pays most of the cost. This is why it is the only technique that is genuinely safe to point at production. It requires privileges — reading another process's memory is exactly what a debugger does, so the same permission model applies, and in a container it usually means the sampler must share the process namespace and hold the relevant capability. That is a deployment decision to make *before* the incident, not during it. ### The 3.14 addition Python 3.14 shipped a supported remote-attach mechanism: `sys.remote_exec(pid, script_path)` asks a running interpreter to execute the given script at its next safe point. It exists for debuggers and diagnostic tooling, and it makes it possible to inject a stack-dumping or sampling routine into a process that was started with no such affordance. It is cooperative in that the target executes the code itself, so it depends on the target reaching a bytecode boundary — a process wedged inside a long native call will not service it promptly. It can also be disabled at build time or by configuration in environments where allowing arbitrary code injection into a running interpreter is unacceptable, which is a control worth knowing exists. ### Choosing the clock, and reading the result Take wall-clock samples first. The failure being investigated — a resource left unclosed under sustained load — consumes very little CPU; it manifests as time parked in opens, closes, retries and finalization, plus growing waits as the pool of available handles shrinks. A CPU-time profile of the same worker would look almost empty and would be read, wrongly, as "the code is fine". Then interpret with the blind spots in mind: - **Native frames collapse.** Time inside a C extension is attributed to the Python frame that called it; you see the boundary, not what is below it, unless the tool also walks native stacks. - **Thread selection matters.** A profile of one thread in a multi-threaded worker is a profile of one thread. Sample all of them, or know why you did not. - **Sample count still rules.** Thirty seconds at a low rate on a busy worker is not enough to trust a 5% figure; let it run. - **Sub-interval calls are individually invisible**, so a fast function called constantly shows up while a fast function called rarely does not appear at all. ### The in-process fallbacks If attaching is impossible — no privileges, a locked-down runtime — there are weaker options that must be armed in advance. `faulthandler.dump_traceback_later(timeout, repeat=True)` prints every thread's stack on a timer, which is a crude, human-readable sampler at very low frequency and is genuinely useful for finding a hang. A sampler thread using `sys._current_frames()` can be shipped disabled and switched on by a signal or an admin endpoint. Both live inside the process, so both compete for the GIL and both need to have been thought about before the incident. ### The answer an interviewer wants Lead with "do not restart, attach" and say why the state matters. Name the mechanism (a separate process reading the target's memory and walking interpreter frames), the privilege requirement, and the choice of a wall-clock profile for this symptom. Close with a limitation you would state in the postmortem — collapsed native frames, or the number of samples you actually collected. That is the difference between having read about this and having done it at three in the morning.

  • What does an out-of-process sampler need from the operating system, and what breaks in a container?
    It needs permission to read another process's address space — the same permission a debugger needs — plus visibility of the target's process id. In containers that usually means sharing the process namespace with the target and granting the debugging capability, and it means the sampler must see the same interpreter binary to interpret the memory layout. Discover and grant this before an incident; scrambling for it during one is how outages get longer.
  • Why does time spent inside a C extension not show up as its own frames in a Python stack sample?
    The sampler reconstructs Python frames from interpreter structures, and native code does not create them. Its cost is attributed to the Python frame that called across the boundary, so you see a wide leaf at the call site with nothing underneath. Confirming what is below it requires a tool that also walks the native stack, or timing the boundary directly.
  • What would make you take a CPU-time profile of this worker instead?
    Evidence that the box is saturated rather than waiting: cores pinned near 100%, throughput flat while queue depth grows, and a wall-clock profile already explained by compute rather than by blocking calls. At that point the waits are noise and a CPU-charged profile concentrates the samples on the code actually consuming the cores.
  • How would you keep this capability available without leaving the process open to code injection?
    Treat remote attach as a privileged, auditable operation: restrict who can reach the host and the process namespace, prefer read-only memory sampling over injecting code, and where policy demands it disable the interpreter's remote-execution support entirely and rely on an in-process sampler that you ship deliberately and can enable through an authenticated control path.

saying these in an interview costs you the question

  • Restarts the process with profiling enabled first
  • Adds per-call instrumentation to a hot production path
  • Assumes attaching needs no special privileges
  • Takes a CPU-time profile of a worker that is waiting
  • Draws conclusions from a few seconds of samples
  • Expects C extension internals to appear as Python frames

context