skip to content

How would you design an asyncio service's shutdown to fit an orchestrator's grace period?

level: principalimportance: should knowfreq 40%

answer

  1. the default teardown is not a policy
  2. stop accepting before you stop working
  3. the killer does not negotiate
  4. budget must fit inside the grace period
  5. idempotent retries beat a longer drain

basics

~20 s

Own the sequence rather than inheriting it: catch SIGTERM with loop.add_signal_handler, stop accepting new work, drain what is in flight under a budget well inside the grace period, cancel the rest, then let the runner finalize.

solid answer

~40 s

Treat shutdown as four explicit phases. **Quiesce**: a `SIGTERM` handler sets an `asyncio.Event`; you fail readiness so traffic stops arriving and close the listener. **Drain**: `asyncio.wait()` on the in-flight tasks with a timeout you choose — comfortably under the supervisor's grace period, because whatever is unfinished when it expires is killed with `SIGKILL` and gets no cleanup at all. **Cancel**: `Task.cancel()` on the stragglers, then `asyncio.gather(..., return_exceptions=True)`. **Finalize**: return from the main coroutine and let `asyncio.run()` finalize async generators and the executor. The judgement is in the budget and in what you promise upstream: the honest answer for most services is at-least-once delivery with idempotent handling, not a promise that everything in memory is flushed on the way out.

code

python · 20 lines
python
import asyncio

async def drain(budget: float) -> None:
    pending = asyncio.all_tasks() - {asyncio.current_task()}
    if not pending:
        return
    _, still_running = await asyncio.wait(pending, timeout=budget)
    for task in still_running:
        task.cancel()
    await asyncio.gather(*still_running, return_exceptions=True)

async def work(seconds):
    await asyncio.sleep(seconds)

async def main():
    running = {asyncio.create_task(work(s)) for s in (0.1, 30)}
    await drain(budget=0.5)
    print("drained;", sum(t.cancelled() for t in running), "task cancelled at the deadline")

asyncio.run(main())

go deeper

for a junior

Know that a supervisor sends SIGTERM first and SIGKILL later, and that the first one is your only chance to clean up. Be able to describe stopping new work before finishing existing work.

for a middle

Explain the mechanics you would write: a signal handler that sets an event, a bounded wait on in-flight tasks, cancel-then-gather for the stragglers, and why the handler itself should do almost nothing.

for a senior

Show the operational judgement — fail readiness before draining, size the budget from measured tail latency, keep uncancellable thread work out of the exit path, and prove the whole thing with a test that stops a loaded instance.

for a principal

Own the tradeoff and the promise: how long a drain the deploy cadence can afford, whether correctness rests on idempotent retries rather than on clean exits, and how much unflushed state the design is allowed to hold at any moment.

## The default is a policy, and it is rarely your policy Left alone, `asyncio.run()` cancels every pending task the instant your main coroutine returns, with no budget and no ordering. That is a fine default for a script and a poor one for a service, where "in flight" means a half-answered request or a batch that has been accepted but not yet durably stored. A production shutdown is something you design; the runner's teardown is the safety net underneath it. ## Four phases **1. Quiesce.** Install the handler on the loop, not with a raw signal handler: ```python loop = asyncio.get_running_loop() stop = asyncio.Event() for sig in (signal.SIGINT, signal.SIGTERM): loop.add_signal_handler(sig, stop.set) ``` The callback does nothing but set an event — no cancelling, no closing. Everything that follows is ordinary async code you can read and test. On waking: fail the readiness probe first so the load balancer removes you, then stop accepting new work (close the listener, stop pulling from the queue). Keep liveness green, or the supervisor will kill you for the crime of shutting down. **2. Drain with a deadline.** Wait for accepted work only, and bound the wait: ```python _, still_running = await asyncio.wait(in_flight, timeout=budget) ``` **3. Cancel the stragglers** and await them so their `finally` blocks actually run: `Task.cancel()` on each, then `asyncio.gather(*still_running, return_exceptions=True)`. **4. Finalize.** Return from `main()`. The runner cancels whatever you missed, finalizes async generators, joins the default executor, closes the loop. ## The budget arithmetic A supervisor typically sends `SIGTERM`, waits a grace period — 30 seconds is a common default — and then sends `SIGKILL`, which cannot be caught, handled or ignored. Everything you want to happen has to fit inside that window, minus a margin for the final flush and process exit. If your drain budget equals the grace period you will be killed mid-cleanup, which is worse than cancelling early, because a task that is cancelled at least gets to run its `finally`. One trap hides in the teardown itself: `asyncio.run()` joins the default executor with a **five-minute** budget on 3.14. Work handed to `asyncio.to_thread()` cannot be cancelled, so a thread still running at exit can hold the process well past any sane grace period — you will be `SIGKILL`ed in the middle of the join and never reach `loop.close()`. Either keep executor work short, or track it and refuse to enter shutdown while a long one is outstanding. ## State is the real constraint The decisive question is how much unflushed state you hold. A webhook receiver that accumulates a 2.4 GB working set of decoded events before writing cannot flush that inside a 30-second window on any realistic storage path, and no amount of shutdown engineering changes the arithmetic. The design fix is upstream of shutdown: checkpoint continuously so the amount at risk is always small, or do not accept the event until it is durable. That choice determines what you can promise. If handlers are idempotent and the sender retries, dropping in-flight work at the deadline is safe — at-least-once delivery with deduplication on a stable key. If they are not idempotent, every hard stop is a correctness event, and the honest response is to fix the handler rather than to lengthen the drain. ## The second signal Operators press Ctrl-C twice and supervisors escalate. The second signal should exit promptly rather than politely: stop draining, cancel immediately, exit. Because cleanup can be interrupted between two awaits, every shutdown path should be crash-safe anyway — write-then-rename, commit after durability, no half-updated structures across an `await`. If your process is only correct when it exits cleanly, it is not correct, because the machine can lose power. ## Verify it, don't assume it Shutdown is the least-tested path in most services and the one most likely to be exercised on every deploy. Make it a test: start the service under load, send `SIGTERM`, assert the exit code, assert the drain finished inside the budget, assert no partial records were written and that everything accepted was either completed or safely retryable. Export the drain duration as a metric — when it creeps toward the grace period you learn it from a dashboard rather than from a `SIGKILL` at 3am. ## The tradeoff to name out loud A longer drain means fewer interrupted requests and slower deploys, rollbacks and autoscaling; a shorter drain means fast, predictable restarts and more reliance on retries. Which one you pick is a product decision about the promise made to callers, not a tuning constant — and it belongs written down beside the grace period it must fit inside.

  • How can outstanding asyncio.to_thread() work wreck a drain that otherwise fits the grace period?
    Thread work cannot be cancelled. At teardown the runner joins the default executor with a five-minute budget on 3.14, so one long-running thread keeps the process alive far past a 30-second grace period and you are `SIGKILL`ed mid-join. Keep executor work short, or track outstanding thread work and wait for it inside your own budget instead.
  • What should the service tell callers during the drain window?
    Fail readiness immediately so the load balancer stops routing to you, keep liveness healthy so the supervisor does not kill you early, finish requests already accepted, and return a retryable status for anything you cannot serve. The window should be invisible to callers that retry and obvious in your metrics.
  • How do you decide the drain budget rather than guessing it?
    Measure the tail latency of the unit of work you promise to finish, take a high percentile, and check it fits inside the grace period with margin for the final flush and exit. If it does not fit, the fix is smaller units of work or durable checkpoints, not a longer window — the supervisor's `SIGKILL` is not negotiable.
  • Why is testing shutdown under load different from testing it idle?
    Idle shutdown always looks clean: nothing is in flight, so no path is exercised. Under load you find the real behaviour — tasks cancelled mid-write, cleanup that itself awaits a dead connection, drain durations that exceed the budget. Send `SIGTERM` to a loaded instance and assert on exit code, drain duration and data completeness.

Closing a kitchen: you stop seating new tables first, finish the orders already fired, bin what cannot be plated in time, and only then turn off the gas — all before the landlord locks the door on a fixed schedule.

saying these in an interview costs you the question

  • Relies on asyncio.run()'s implicit cancel as the drain policy
  • Sets a drain budget equal to or longer than the grace period
  • Assumes SIGKILL can be caught and cleaned up after
  • Holds large unflushed state and hopes for a clean exit
  • Blocks the event loop with synchronous flush work while shutting down
  • Never tests shutdown while the service is under load

context