skip to content

Why might a script sent with sys.remote_exec never run in a process that has stopped making progress?

level: seniorimportance: should knowfreq 18%

answer

  1. Attach is a request, not an interrupt
  2. Only between complete bytecode instructions
  3. Blocked in C means never noticed
  4. Delivered like a signal, when safe
  5. Silence means no bytecode is running

basics

~20 s

Python 3.14 runs an injected script only at a safe point, between complete bytecode instructions on the target's main thread. A main thread parked in a blocking system call or a long C call never reaches one, so the queued request simply waits.

solid answer

~50 s

Attaching is a request, not an interrupt. `sys.remote_exec` writes a path and raises a flag; the target's main thread acts on it the way it acts on a signal — at the next **safe point** in the evaluation loop, between complete bytecode instructions, where the interpreter's invariants hold. Executing arbitrary Python halfway through an instruction, or inside a C call that holds internal locks, would corrupt interpreter state, so CPython will not do it. The operational consequence is that a main thread blocked forever on a socket read, waiting on a lock held by someone else, or grinding inside a C extension that released the GIL, never returns to the evaluation loop and never runs your script. Silence after an attach is therefore data, not a bug: it says the main thread is not executing bytecode, and you need a technique that does not require the target's cooperation.

code

python · 6 lines
python
import sys
import traceback

for thread_id, frame in sys._current_frames().items():
    print(f"--- thread {thread_id} ---", flush=True)
    traceback.print_stack(frame)

go deeper

for a junior

Remember that attaching to a live Python 3.14 process asks it to run your file when it can, rather than stopping it on the spot. If the process is stuck in something that never returns, nothing happens.

for a middle

Explain the two phases: the caller writes a request, and the target's main thread acts on it at a safe point between bytecode instructions. Be able to name states that never reach one, such as a blocking socket read or a long call into C code.

for a senior

Demonstrate that you read the outcome diagnostically: script runs means the main thread is executing Python; total silence narrows the fault to a blocked call. Have a plan for the silent case, and make your first injected script cheap enough to prove liveness.

for a principal

The judgement to own is when cooperative inspection is the wrong tool for your fleet. If your typical incident is a wedged process, standardise on inspection that needs no cooperation from the target and treat attaching as the follow-up, not the first move.

### What a safe point is CPython's evaluation loop executes one bytecode instruction after another. Between instructions there is a moment where the interpreter's own bookkeeping is consistent: no half-built frame, no object mid-mutation, no internal lock held for a partially finished operation. The loop checks a small set of pending requests at exactly those moments — that is where signal handlers run, where pending calls are dispatched, and, since PEP 768 in Python 3.14, where a remote debugging request is noticed. So `sys.remote_exec` has two phases with a gap between them that is entirely outside the caller's control. Phase one: the caller writes the script path into the target's control area, sets the pending flag, and returns. Phase two: the target's **main thread** reaches a safe point, sees the flag, reads the file, and executes it. The gap can be microseconds or forever. ### Why CPython insists on this The alternative — vectoring the target into your code wherever its instruction pointer happens to be — is how you corrupt a process. The thread might be inside a C function that is midway through resizing a container, holding an internal lock, or working with a borrowed reference that is only valid until the function returns. Re-entering the interpreter there could deadlock on a lock the same thread already holds, or leave reference counts and container invariants broken in a way that crashes minutes later somewhere unrelated. A safe point is a correctness requirement, not a performance nicety: it is the only place where 'run this arbitrary Python now' has a defined meaning. ### The states that never reach one * **Blocked in a system call.** A main thread sitting in a socket read with no timeout, waiting on a pipe, or joining a child, is not executing bytecode. It will notice the request only when the call returns — which, for a hung dependency, may be never. * **Inside a long C call.** A C extension performing a long computation, typically with the GIL released, has left the evaluation loop; the flag is checked after it returns. * **Deadlocked in native code.** A lock cycle inside a C library never unwinds, so no safe point ever arrives. * **Stopped by the operating system.** A process that is suspended executes nothing at all. A subtlety worth stating explicitly: other threads reaching safe points do not help, because the injected script is delivered to the main thread. A busy worker pool proves nothing about whether the attach will land. ### Reading the silence: a worked case A four-person team owns a nightly museum-catalogue importer. Tonight's run stops advancing part way through the batch, and the standing suspicion is a locale-dependent date format that drives the record parser down a pathological path. Nothing has crashed; the process is simply not finishing. Attach and inject a stack dumper. Two outcomes, and both are informative: 1. **Output appears immediately.** The main thread is executing bytecode, so it is looping or grinding in Python. The dumped frames name the function — and if they land in the date-parsing code, the locale hypothesis is now evidence rather than a hunch. 2. **Nothing appears at all.** The main thread is not running bytecode. It is blocked in a call that has not returned — a database socket, a lock, a native library. Attaching has told you where the problem is *not*, which is a real narrowing. In case two, escalate to a technique that requires no cooperation from the target: read the interpreter's state from outside the process, or look at the native stack. The distinction is exactly the tradeoff between the two families of tool — an attach runs your code inside the process and needs it healthy enough to reach a safe point; an out-of-process reader does not need the target to do anything, and correspondingly cannot run your logic inside it. ### Practical implications Never treat a queued request as an executed one: there is no completion signal, so 'I attached' is not the same claim as 'my script ran'. When you attach during an incident, make the first injected script prove liveness cheaply — one line to a file or to stderr — before you inject anything elaborate. And be aware of the timing skew: if the target *does* reach a safe point later, your script runs then, possibly long after you have moved on, which is a good reason for injected scripts to be idempotent and to identify themselves in whatever they write.

  • Does releasing the GIL count as reaching a safe point?
    No. The check lives in the evaluation loop, and a thread that dropped the GIL inside a C call is not executing bytecode; it sees the pending flag only after that call returns. Releasing the GIL lets *other* threads run Python, but the injected script is delivered to the main thread, so a busy worker thread does not get you an attach.
  • Why not simply interrupt the target wherever it happens to be?
    Because the interpreter's invariants only hold between instructions. Mid-instruction the thread may hold an internal lock, be resizing a container, or be working with a borrowed reference; re-entering the interpreter there can deadlock or corrupt state that crashes much later. Safe points are the same discipline CPython already uses to deliver signal handlers.
  • Your attach produces no output at all. What is the next move?
    Treat the silence as the finding: the main thread is not running bytecode, so it is blocked in a call or suspended. Switch to a method that needs no cooperation — read the interpreter's state from outside the process, or inspect the native stack — and check at the operating-system level whether the process is even runnable.

It is like leaving a message for a surgeon: they will read it when they step out of the operating theatre, and if the operation never ends, the message is never read.

saying these in an interview costs you the question

  • Thinks attaching stops the target immediately
  • Believes a deadlocked process still runs the injected script
  • Assumes the script preempts a running C call
  • Treats safe points as an optimisation rather than correctness
  • Confuses a queued request with an executed one
  • Expects a busy worker thread to run the injected script

context