skip to content

A worker is stuck in `re.search()` on one email digest for 27 minutes and ignores Ctrl-C. How do you confirm the cause and bound it?

level: seniorimportance: should knowfreq 40%

answer

  1. Confirm from the live process first
  2. Signals wait for the eval loop
  3. No time budget exists in the module
  4. Threads share the wedged interpreter
  5. Only the OS can stop it

basics

~20 s

Dump the live process stack with faulthandler to confirm it is inside the match. re has no timeout and the C matching loop defers signals, so the only hard bound is a separate process you can kill.

solid answer

~50 s

First confirm, do not guess: `faulthandler.register()` on a signal, or `faulthandler.dump_traceback_later()` armed at startup, will print the stack of a wedged process even though it is stuck in C — and 3.14's remote debugging hook (`sys.remote_exec`, PEP 768) lets you attach to a live pid. Ctrl-C does nothing because the matcher runs a C loop that never returns to the eval loop, so pending signal handlers are not run and the GIL is never released; a watchdog **thread** cannot help either. Then find the input: an off-by-one in a subject line, one missing closing bracket, turns a valid digest header into a near-match and the ambiguous pattern enumerates every split. Fix the pattern — remove the nested quantifier, or fence it with an atomic group — cap input length, and for anything you do not control, run the match in a separate process you can `terminate()`.

code

python · 29 lines
python
import multiprocessing
import re

PATTERN = r"(a+)+$"


def _worker(text, out):
    out.value = bool(re.search(PATTERN, text))


def match_or_give_up(text, seconds=0.5):
    ctx = multiprocessing.get_context("spawn")
    out = ctx.Value("b", False)
    proc = ctx.Process(target=_worker, args=(text, out))
    proc.start()
    proc.join(seconds)
    if proc.is_alive():
        proc.terminate()
        proc.join()
        raise TimeoutError("regex budget exceeded")
    return bool(out.value)


if __name__ == "__main__":
    print(match_or_give_up("aaa"))
    try:
        match_or_give_up("a" * 40 + "!")
    except TimeoutError as exc:
        print("killed:", exc)

go deeper

for a junior

Recall that a pegged CPU with no response to Ctrl-C looks like a hang but is a computation, and that re has no timeout option to pass.

for a middle

Explain the deferred-signal mechanism that makes Ctrl-C and SIGALRM useless here, and reproduce the blow-up with a short timing ladder over increasing input lengths.

for a senior

Show the whole path: dump the stack of the live process, identify the near-match input, fix the ambiguous pattern, then add the input cap and the killable-process budget so the class of failure is contained.

for a principal

Decide the policy — whether patterns from config or users are accepted at all, what per-item wall-clock budget the pipeline enforces, and how a hang is detected and alerted long before it is noticed by a customer.

## Step 1 — confirm it is the regex, from the live process A pegged CPU and an unresponsive worker have several causes; do not assume. The useful property of `faulthandler` is that it works when nothing else does, because its handlers dump the interpreter's stacks from C without needing the GIL: - `faulthandler.register(signal.SIGUSR1)` at startup, then send that signal to the pid: the traceback of every thread is written to stderr, and you will see the frame that called `re.search()`. - `faulthandler.dump_traceback_later(60, repeat=True)` arms a watchdog that dumps automatically if the process ever goes 60 seconds without resetting it — good as a standing defence in a batch worker. - On 3.14, PEP 768's remote debugging hook (`sys.remote_exec`) lets you attach to a running process by pid without having pre-armed anything. Then get the input. Log the digest id and the length of the text before each match, or reconstruct it from the queue; the offending item is the one still in flight. Reproduce offline with a timing ladder — the same pattern against the same text truncated to 18, 20 and 22 characters — and if each step roughly doubles, you have exponential backtracking rather than a merely large job. ## Step 2 — understand why nothing stopped it **Ctrl-C is ignored.** Python's signal handling is cooperative: the C-level handler sets a flag, and the *Python* handler runs when the eval loop next checks it. The matcher is a tight C loop that does not return to the eval loop until the match finishes, so `KeyboardInterrupt` is only raised after the match completes — which may be never in practice. **`signal.SIGALRM` does not save you either,** for the same reason: `signal.setitimer()` fires, the flag is set, and the handler still waits for the eval loop. **A watchdog thread cannot interrupt it.** The match holds the GIL for its whole duration, so other Python threads in that interpreter do not run. There is also no API anywhere in Python to interrupt one thread's computation from another. **There is no timeout parameter.** `re.search()`, `re.match()`, `re.fullmatch()`, `re.compile()` — none of them accepts a time budget, and there is no module-level setting or environment variable that imposes one. This is a deliberate, long-standing gap, and it is the single fact that shapes every mitigation. **A subinterpreter does not fix it.** 3.14 shipped multiple interpreters in the stdlib (`concurrent.interpreters`), and a separate interpreter gets its own GIL — but there is still no way to stop code already running inside one. Isolation is not cancellation. ## Step 3 — fix the pattern The digest header pattern was almost certainly something like `(\s*\w+)*` or a group with a nested quantifier, plus a trailing anchor or literal. One malformed subject line — a bracket short of correct — turns it into a near-match, and the engine must prove that no arrangement works. - Rewrite to remove the ambiguity: `\w+(\s+\w+)*` in place of `(\s*\w+)*`; collapse `(a+)+` to `a+`; make alternation branches disjoint. - Anchor with `re.fullmatch()` or a leading `^` so the attempt is not retried at every offset. - Since 3.11, fence a subpattern you cannot easily rewrite with an atomic group `(?>...)` or a possessive quantifier, which discards the saved alternatives — checking first that it does not change what matches. - Consider not using a regex: header parsing in particular is usually better served by a real parser or by `str.partition()`. ## Step 4 — bound the cost structurally Fixing this one pattern does not stop the next one. What actually holds: - **Cap input length** before matching. Most fields have a defensible maximum, and an exponential curve is harmless if `n` cannot exceed a small bound. - **Never accept a pattern from outside the codebase.** A user-supplied pattern is an unbounded compute primitive; if a product requires it, treat it the way you would treat running submitted code. - **Run untrusted matching in a killable process.** Because SIGTERM's default disposition terminates the process at the OS level rather than going through Python's deferred handling, `Process.terminate()` really does stop a wedged match — a thread or a subinterpreter does not. - **Timebox the work item, not the regex.** A per-item wall-clock budget in the worker, enforced by the process that owns it, converts a hang into a failed item plus an alert. - **Alert on it.** A batch stage whose per-item time is normally milliseconds should page long before 27 minutes. The senior point to make is the ordering: confirm from the live process, fix the pattern, then put the structural bound in place so the class of bug cannot take the worker down again.

  • Why can't you enforce the budget with `signal.setitimer()` and a `SIGALRM` handler in the same process?
    Because Python defers signal handlers to the eval loop. The C-level handler sets a flag when the timer fires, but the Python-level handler only runs when the interpreter next executes bytecode — and the matcher is a C loop that does not return until the match ends. The alarm fires and nothing happens. The technique works for pure-Python hot loops and for interruptible syscalls, not for this.
  • Would running the match in a thread pool with a result timeout contain it?
    No. A future's timeout only stops you waiting; the thread keeps running and there is no way to cancel it. Worse, the match holds the GIL for its whole duration, so the other Python threads in that interpreter make no progress either — you lose the whole worker, not just one item. Only a separate process can be killed.
  • The pattern is loaded from a config file that operators edit. What changes?
    It becomes an untrusted pattern, and no rewrite of today's pattern protects you. Validate on load, run matching inside the killable process with a budget, cap input length, and treat a rejected pattern as a config error surfaced at deploy time rather than a hang discovered in production. If operators genuinely need arbitrary patterns, that is a compute-sandbox decision, not a regex decision.

saying these in an interview costs you the question

  • Suggests passing a timeout argument to `re.search()`
  • Proposes a watchdog thread to cancel the match
  • Expects `SIGALRM` to interrupt the C matching loop
  • Says Ctrl-C works and the process must be blocked on I/O
  • Fixes the one pattern and adds no structural bound
  • Accepts operator-supplied patterns with no budget

context