skip to content

An email-digest sender gets SIGTERM with a 30-second window before SIGKILL — how do you drain without truncating a digest?

level: seniorimportance: should knowfreq 42%

answer

  1. Two phases against a clock
  2. The handler does almost nothing
  3. One signal cannot be caught at all
  4. The budget is smaller than the window
  5. Acknowledge before you record sent

basics

~20 s

Catch signal.SIGTERM with a handler that only sets a flag, stop pulling new digests, finish the ones in flight against a deadline shorter than the 30-second window, and exit. signal.SIGKILL cannot be caught, so anything unfinished must be safely retryable.

solid answer

~50 s

Split shutdown into two phases against a clock. The handler for `signal.SIGTERM` does one thing — set a `threading.Event` — because it runs at an arbitrary bytecode boundary in the main thread and must not do I/O. The main loop sees the flag, stops claiming new digests, and drains what is already in flight against a deadline computed with `time.monotonic()` that is comfortably inside the 30-second window, leaving room to flush and close. Truncation is prevented not by draining but by ordering: render the whole digest, hand it to the transport, and only mark it sent once the transport has acknowledged it — so a digest interrupted mid-render is never half-delivered and is simply re-sent on the next run. When the deadline passes, abandon the rest and exit; `signal.SIGKILL` at the end of the window cannot be caught, blocked or ignored.

code

python · 14 lines
python
import signal
import threading
import time

stopping = threading.Event()
signal.signal(signal.SIGTERM, lambda signum, frame: stopping.set())

pending = list(range(5))
sent = 0
deadline = time.monotonic() + 25.0
while pending and not (stopping.is_set() and time.monotonic() >= deadline):
    pending.pop(0)
    sent += 1
print("sent", sent, "abandoned", len(pending))

go deeper

for a junior

Know the shape: the terminate signal is catchable, the kill that follows is not, so a process gets a limited window to finish what it started and must not begin anything new.

for a middle

Explain why the handler only sets a flag, and how a monotonic deadline bounds the drain so the process exits on its own terms rather than being killed mid-write.

for a senior

Show that safety comes from ordering, not from draining: acknowledge before recording, keep work retryable, forward signals to children, and test shutdown with the signal production actually sends.

for a principal

Frame it as a property of the work pipeline. If a killed process can lose or half-deliver anything, the drain code is not the defect; the at-least-once contract is missing.

## The contract you are being asked to implement A supervisor stopping a process sends a catchable request first and an uncatchable one later. The process gets a **grace window** — here 30 seconds — between the two. `signal.SIGKILL` and `signal.SIGSTOP` cannot be caught, handled, blocked or ignored; `signal.signal` raises `OSError` if you try. So the design question is never "how do I survive the kill" but "what state is the world in when the kill lands". ## Phase one: stop taking work The handler's only job is to record intent: ```python import signal import threading stopping = threading.Event() signal.signal(signal.SIGTERM, lambda signum, frame: stopping.set()) ``` Handlers run at a bytecode boundary in the main thread while the rest of the program is mid-flight. Logging, locking or network calls inside one can deadlock against whatever the interrupted code was holding. Setting an already-set `threading.Event` is idempotent, which matters because a second terminate signal is a normal occurrence. The dispatcher then stops claiming new digests. This is the single highest-value step: the set of work you must finish stops growing the instant the signal lands, so the drain has a bounded amount to do. ## Phase two: drain against a deadline, not against hope ```python import time deadline = time.monotonic() + 25.0 while in_flight and time.monotonic() < deadline: finish_one(in_flight.pop()) ``` Use `time.monotonic()`, never `time.time()`: a wall-clock adjustment mid-drain would move your deadline. Budget strictly less than the window — 25 seconds inside 30 — because after the drain you still have to flush buffers, close connections and let the interpreter finalize, and because the supervisor's clock started before yours did. A thread pool needs the same discipline. `ThreadPoolExecutor.shutdown(wait=True, cancel_futures=True)` stops the queue from being drained into workers and waits for the running tasks, but it has no timeout parameter, so a task that never returns holds you past the window. Either keep the per-item work short enough to be bounded, or track the deadline yourself and stop waiting when it passes. ## Where the truncation actually comes from The silent truncation in a digest sender is almost never caused by the drain being too short. It is caused by **ordering**: a process that streams sections into an outgoing message as it renders them, or that marks a digest as sent before the transport acknowledged it, will happily deliver a message containing the first four sections of twelve and record the job as complete. Nobody sees an error; the recipient just gets less than they should have. The ordering that survives being killed anywhere: 1. Render the digest completely, in memory or to a staging location. 2. Hand the finished artefact to the transport in a single operation. 3. Only after the transport acknowledges it, record "sent" and advance the cursor. Interrupted before step 3, the digest is re-sent next run — at-least-once rather than half-delivered. That is a property of the pipeline, not of the shutdown code, and it is what makes abandoning work safe. Deduplicate on a per-digest idempotency key if a duplicate would be worse than a delay. ## asyncio wears the same shape ```python import asyncio import signal async def drain_in_flight(): await asyncio.sleep(0.1) async def main(): loop = asyncio.get_running_loop() stopping = asyncio.Event() loop.add_signal_handler(signal.SIGTERM, stopping.set) loop.call_later(0.1, stopping.set) await stopping.wait() async with asyncio.timeout(5): await drain_in_flight() print("drained before the deadline") asyncio.run(main()) ``` `loop.add_signal_handler` is the asyncio-correct registration: the callback is scheduled on the event loop rather than running at an arbitrary bytecode boundary, so it may safely touch loop objects. `asyncio.timeout` (3.11+) gives the drain its deadline. Cancelling the remaining tasks and awaiting them with `return_exceptions=True` keeps one failing cleanup from hiding the others. ## Children and connections If the sender spawns children, forward the signal to them and give them their own, smaller budget: `Popen.terminate()`, then `Popen.wait(timeout=...)`, then `Popen.kill()` if the timeout expires. Otherwise your children outlive you or get killed mid-write. Connections deserve an explicit close inside the budget. A transport that is abandoned rather than closed can leave the peer holding a half-open connection until its own timeout fires, which is a slow way to turn a clean deploy into elevated latency elsewhere. ## Prove it Test it the way production stops you, not the way your terminal does: start the service, send `signal.SIGTERM` from a test harness with `os.kill`, and assert on the exit status and on the state of the work store — nothing marked sent that was not acknowledged, nothing lost that was not retryable. Record how long shutdown took and how often the process was still alive when the window expired; an exit reported as killed by signal 9 is the signal that your budget is wrong.

  • Why set a flag in the handler instead of raising SystemExit there?
    Because the handler runs at whatever bytecode boundary the main thread has reached, which may be between handing a digest to the transport and recording that it was sent. Raising there unwinds correctly but at a point you did not choose. A flag lets the main loop finish the current item and begin shutdown at a boundary the program defines, which is what makes the outcome predictable.
  • The drain still has work left when the deadline passes. What should the process do?
    Stop, leave the remaining items unclaimed, close connections and exit — deliberately, with a status that says shutdown was truncated, rather than gambling on the kill arriving late. That is only acceptable because the pipeline is at-least-once: an unacknowledged digest is re-sent next run. If abandoning work were unsafe, the fix is in the work model, not in a longer wait.
  • How would you verify the drain works without waiting for a real deploy?
    Start the service in a test, send `signal.SIGTERM` with `os.kill`, and assert on both the exit status and the work store: nothing marked sent that the transport did not acknowledge, and everything unfinished still claimable. In production, record shutdown duration as a metric and alert on processes reported as killed by signal 9, which means the budget is wrong.

It is a shop closing: lock the door so no new customers enter, serve the ones already inside, and be out before the landlord changes the locks — which he will do whether or not you are ready.

saying these in an interview costs you the question

  • Trying to catch or block SIGKILL
  • Doing logging or network I/O inside the signal handler
  • Draining with no deadline and hoping it finishes
  • Using time.time instead of a monotonic clock
  • Marking work complete before the transport acknowledges
  • Continuing to accept new work after the signal

context