skip to content

Syscall Retry and EINTR

Since 3.5 the interpreter retries most system calls a signal interrupted instead of raising, so the old retry loops are dead code. What still surfaces the error, and what a recomputed deadline costs.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What did PEP 475 change in Python 3.5 about system calls interrupted by a signal?

level: middleimportance: must knowfreq 30%

answer

  1. It used to be your loop
  2. Now the loop lives in the interpreter
  3. Handler runs first, then reissue
  4. Timeout becomes a monotonic deadline
  5. A raising handler still wins

basics

~20 s

Since Python 3.5, the interpreter reissues a system call that a signal interrupted instead of raising InterruptedError, after running the Python-level handler and recomputing any remaining timeout. Hand-written EINTR retry loops around stdlib calls became dead code.

solid answer

~50 s

Before 3.5, any blocking stdlib call could fail with `InterruptedError` — an `OSError` subclass carrying `errno.EINTR` — whenever a signal arrived, so correct code wrapped every wait in a retry loop and almost nobody did. PEP 475 moved that loop into the interpreter: the Python-level handler runs, and if it returns normally the call is reissued automatically. For calls that take a timeout the interpreter recomputes the remaining time from a monotonic deadline, so repeated interruptions cannot stretch a bounded wait. Two things still escape: a handler that raises aborts the retry and its own exception propagates, and operations where reissuing would be unsafe — closing a descriptor is the classic case — ignore the interruption rather than repeat the call. `InterruptedError` still exists as a class, and you can still meet it from code that reaches libc directly through a C extension or `ctypes`.

code

python · 22 lines
python
import os
import signal
import socket
import threading
import time

signal.signal(signal.SIGUSR1, lambda signum, frame: print("handler ran"))
sock_a, sock_b = socket.socketpair()
sock_a.settimeout(1.0)


def interrupt() -> None:
    time.sleep(0.2)
    os.kill(os.getpid(), signal.SIGUSR1)


threading.Thread(target=interrupt).start()
start = time.monotonic()
try:
    sock_a.recv(1)
except TimeoutError:
    print(f"timed out after {time.monotonic() - start:.2f}s, not 1.2s")

go deeper

for a junior

Know the headline: on modern Python you do not write retry loops for signal-interrupted calls, because the interpreter reissues them for you. Recognise InterruptedError as the old symptom rather than something you should be handling.

for a middle

Explain the ordering — kernel returns the error, the Python handler runs, then the call is reissued — and that a timeout is recomputed from a deadline rather than restarted. Name the version: 3.5, PEP 475.

for a senior

Show judgement about the exceptions to the rule: a raising handler aborts the retry, unsafe-to-repeat operations are not retried, and native code outside CPython's wrappers still needs its own handling. Flag inherited except-OSError loops as over-broad.

for a principal

Frame it as a class of intermittent, environment-dependent bug that the platform absorbed. Decide when your own native or ctypes-facing code owes callers the same guarantee, and how signal-driven aborts should surface across a codebase.

## The problem PEP 475 solved On Unix, delivering a signal to a thread blocked in an interruptible system call makes that call return early with `EINTR`. That is a kernel-level fact and it predates Python by decades. Before Python 3.5 the interpreter passed it straight through: `os.read`, `time.sleep`, a socket receive, `select.select`, waiting on a child — any of them could raise `InterruptedError`, an `OSError` subclass whose `errno` is `errno.EINTR`. The practical result was that *correct* Python code had to wrap every blocking call in a retry loop, and essentially nobody did. Programs worked in development, where no signals arrive, and failed in production, where a process reaping children, rotating logs or running under a supervisor gets signals routinely. The failure was intermittent, load-dependent and looked like a random `OSError` in an unrelated place. ## What the change actually does PEP 475, shipped in Python 3.5, moved the retry loop into the interpreter. The sequence for an interrupted call is now: 1. The kernel returns `EINTR` from the call. 2. The interpreter runs any pending Python-level signal handlers, in the main thread, at the next check. 3. If none of them raised, the call is reissued. 4. If the call takes a timeout, the timeout passed to the reissued call is the *remaining* time, computed from a deadline taken on a monotonic clock before the first attempt. Step 4 is the part people forget, and it is what makes the change safe rather than merely convenient. Naive retrying re-passes the original timeout, so a call being interrupted every second with a ten-second timeout waits forever. The interpreter instead treats your timeout as a deadline, so total wall-clock time stays bounded no matter how often signals arrive. ## Where the retry does not apply **A handler that raises wins.** If a Python-level handler raises — the default `SIGINT` handler raising `KeyboardInterrupt`, or your own handler raising a shutdown exception — the interpreter does not reissue the call. The handler's exception propagates out of the blocked call. This is the documented escape hatch that makes Ctrl-C and signal-driven aborts work at all. **Operations where reissuing is unsafe.** Repeating a call is only correct when the first attempt did nothing. Closing a descriptor is the standard counter-example: on Linux the descriptor is released even when the call reports `EINTR`, so reissuing could act on a number that has already been handed out to something else. Those cases ignore the interruption rather than repeat the call. **Code that bypasses the interpreter's wrappers.** The retry lives in CPython's own I/O and syscall wrappers. A C extension that calls libc itself, or your own `ctypes` call into a blocking function, gets raw `EINTR` and must handle it — which is why `InterruptedError` is still a live built-in exception in 3.14 rather than a historical curiosity. ## What this means for code you write today Delete the retry loops. A pattern like `while True: try: data = sock.recv(n); break; except InterruptedError: continue` is pure noise on any supported Python, and worse than noise if it also re-passes a timeout. If you inherit such a loop, check whether it swallows more than it means to — `except OSError: continue` around a wait quietly retries permission errors, connection resets and bad descriptors as well. Do not confuse this with the interpreter *ignoring* signals. Handlers still run, and they run promptly: CPython registers Python-level handlers in interruptible mode specifically so that the kernel unblocks the call and gives the eval loop a chance to dispatch, rather than restarting the call itself and delaying your handler until the read finally returns. Also remember what a retry cannot fix. The interpreter can only reissue calls that made no progress. A partially completed transfer reports the bytes it moved; that is a normal short read or short write, not an interruption, and code that assumes all-or-nothing was broken before PEP 475 too. ## Saying it in an interview One sentence: since 3.5 the interpreter runs your handler and then reissues the interrupted call, recomputing the remaining timeout, so `InterruptedError` only reaches you from outside CPython's own wrappers — and a handler that raises still beats the retry.

  • What does the interpreter do about a timeout when it reissues the call?
    It converts the timeout into a deadline on a monotonic clock before the first attempt, and passes the remaining time to each reissue. So a one-second wait interrupted five times still ends at about one second, not six. Re-passing the original timeout — which is what most hand-rolled retry loops do — is exactly the bug this avoids: with signals arriving faster than the timeout, the wait never expires.
  • Where can you still meet InterruptedError on Python 3.14?
    From code that reaches the operating system without going through CPython's own wrappers: a C extension calling libc directly, or your own `ctypes` call into a blocking function. Those return raw `EINTR`, and the class is still a live `OSError` subclass with `errno` set to `errno.EINTR`. It is also what you construct in tests when you want to simulate an interruption.
  • Why is reissuing not always the right answer?
    Because a retry is only safe when the first attempt had no effect. Closing a descriptor is the canonical counter-example: the descriptor can already be released when the call reports interruption, so repeating it risks acting on a number the kernel has since reused. Cases like that ignore the interruption instead of repeating the call, which is a different remedy from retrying.

saying these in an interview costs you the question

  • Says InterruptedError was removed from Python 3
  • Still writes except InterruptedError retry loops around stdlib calls
  • Thinks the retry happens before the signal handler runs
  • Claims each retry restarts the full timeout
  • Believes PEP 475 prevents handlers from raising
  • Confuses the automatic retry with ignoring signals entirely

context

open as a page

Why does Ctrl-C during a blocking socket.socket.recv raise KeyboardInterrupt instead of resuming the read?

level: juniorimportance: should knowfreq 35%

basics

~20 s

Ctrl-C delivers SIGINT, and Python's default SIGINT handler raises KeyboardInterrupt. Since Python 3.5 the interpreter reissues an interrupted call only when the handler returns normally; a handler that raises aborts the retry, so the exception leaves recv.

open as a page

A retry loop re-passes its 30-second timeout and catches Exception on every interruption; why does the wait never expire?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Re-passing the original timeout restarts the clock on every interruption, so a wait interrupted more often than every 30 seconds never expires. Catching Exception also swallows the abort a signal handler raised, so the shutdown request disappears with it.

open as a page

What does signal.siginterrupt(sig, False) change, and why does CPython leave calls interruptible by default?

level: seniorimportance: nice to knowfreq 12%

basics

~20 s

It asks the kernel to restart a system call that signal interrupts, instead of failing it. Python's own handler then cannot run until the blocked call returns, so an abort may be delayed for as long as the call blocks.

open as a page