skip to content

During a rolling deploy, a Celery 5.6 worker gets SIGTERM mid-way through a long PDF export and is killed after the grace period; what do warm, soft and cold shutdown do, and how do you keep the task?

level: seniorimportance: should knowfreq 30%

answer

  1. SIGTERM waits without a limit
  2. SIGQUIT cancels running work
  3. a timed wait before cancelling
  4. REMAP_SIGTERM
  5. only unacknowledged messages come back

basics

~20 s

SIGTERM starts a warm shutdown that waits, unbounded, for running tasks, so the later SIGKILL loses them. Cold shutdown (SIGQUIT) cancels them; soft shutdown (5.5+) first waits worker_soft_shutdown_timeout. Use REMAP_SIGTERM=SIGQUIT, a timeout under the grace period, and acks_late.

solid answer

~40 s

On `SIGTERM` a Celery worker does a **warm shutdown**: it stops taking work and waits for running tasks with no time limit. The platform's `SIGKILL` at the end of its grace period then kills the export mid-way, and with the default early acknowledgement its message is already gone. **Cold shutdown** (`SIGQUIT`) cancels running tasks and exits. Celery 5.5 added **soft shutdown**: with `worker_soft_shutdown_timeout` above 0 (default `0.0`, off), a cold shutdown first waits that many seconds, then cancels what is left, and messages still unacknowledged, those of `acks_late` tasks, go back to the broker for redelivery. To get that path from a deploy's `SIGTERM`, set the environment variable `REMAP_SIGTERM=SIGQUIT`, keep the timeout well under the grace period, and make the export `acks_late` and safe to run twice.

code

python · 12 lines
python
from celery import Celery

app = Celery("proj", broker="amqp://guest@rabbitmq//")
app.conf.update(
    worker_soft_shutdown_timeout=20.0,  # below the platform's 30 s grace period
    worker_enable_soft_shutdown_on_idle=True,  # also protects reserved ETA tasks
)


@app.task(acks_late=True, soft_time_limit=1500, time_limit=1560)
def render_pdf(export_id):
    ...  # idempotent: writes to a temp path, publishes atomically

go deeper

for a junior

Recall that SIGTERM asks a Celery worker to finish its running tasks and stop, and that SIGQUIT stops it at once, cancelling what is running.

for a middle

Explain the stages: warm waits without a limit, cold cancels, soft inserts a timed wait before the cancel, and only unacknowledged messages return to the broker.

for a senior

Design deploys around the grace period: REMAP_SIGTERM=SIGQUIT, a soft timeout below the grace window, acks_late on idempotent long tasks, and time limits as the backstop.

for a principal

Weigh long grace periods against re-running work on redelivery, and decide whether long exports should be split so deploys stop being the event that loses them.

## The four shutdown stages Celery 5.5 and later names four stages, and each signal moves the worker to one of them: | Stage | Triggered by | Running tasks | Available | |---|---|---|---| | **Warm** | `SIGTERM`, or the first Ctrl-C (`SIGINT`) | Allowed to finish, with no time limit | Every 5.x release | | **Soft** | The start of a cold shutdown, when `worker_soft_shutdown_timeout` is above 0 | Given that many seconds, then cancelled | 5.5 and later | | **Cold** | `SIGQUIT`, or the next Ctrl-C during a warm shutdown | Cancelled, then the worker exits | Every 5.x release | | **Hard** | Repeated Ctrl-C | Worker terminates immediately by force | 5.5 and later | ## Why the rolling deploy loses the export A typical rolling deploy stops the old worker with `SIGTERM` and follows up with `SIGKILL` if it has not exited within a grace period, say 30 seconds. 1. `SIGTERM` starts a **warm shutdown**: the worker stops consuming and waits for the 20-minute export. Further `SIGTERM` signals are ignored. 2. Thirty seconds later `SIGKILL` arrives. No process can catch it; the main process dies, and on Linux its pool children are killed with it. 3. With the default early acknowledgement, the message was acknowledged when the task **started**, so the broker has nothing to redeliver: the export is lost. 4. Messages the worker had prefetched but not started were never acknowledged, so the broker redelivers those. Warm shutdown is right when tasks are short. It is wrong when a task can outlive the grace period, because then the ending is decided by `SIGKILL`, not by Celery. ## Soft shutdown and REMAP_SIGTERM Soft shutdown is a **time-limited warm shutdown placed just before the cold one**. When a cold shutdown begins and `worker_soft_shutdown_timeout` is positive (the default `0.0` disables it), the worker waits that many seconds for running tasks, then cancels whatever is still running and exits. Cancelled tasks whose messages are still unacknowledged, those declared with `acks_late=True`, are left for the broker to redeliver to another worker; the log line `Restoring N unacknowledged message(s)` is that hand-back. The catch is that soft shutdown hangs off the **cold** path, while a deploy sends `SIGTERM`, which means warm. Setting the environment variable `REMAP_SIGTERM=SIGQUIT` on the worker makes `SIGTERM` start a cold shutdown instead, and so, with the timeout set, a soft one. The setting on its own does not change what `SIGTERM` does. ## What comes back and what does not - **`acks_late` task, not yet acknowledged**: cancelled, message returned for redelivery. The export runs again from the start elsewhere, so it must be safe to run twice. - **Early-acknowledged task, the default**: cancelled, and its message is already gone. It is lost exactly as with `SIGKILL`. - **Reserved but not started**: never acknowledged, so it goes back to the broker. - **ETA or countdown tasks on an idle worker**: soft shutdown is skipped when no task is running, and the worker guide warns this can lose reserved ETA tasks during the cold shutdown. `worker_enable_soft_shutdown_on_idle=True` forces the wait. ## A deploy recipe for long exports 1. Set `REMAP_SIGTERM=SIGQUIT` in the worker's environment. 2. Set `worker_soft_shutdown_timeout` comfortably **below** the platform's grace period, for example 20 seconds against 30, so cancellation and hand-back finish before `SIGKILL`. 3. Declare the export `acks_late=True` and make it idempotent: write to a temporary path and publish the finished file atomically. 4. Enable `worker_enable_soft_shutdown_on_idle` if workers hold ETA or countdown tasks. 5. Keep a `time_limit` on the export anyway, so a hung export cannot hold a worker between deploys either. ## Details that bite - In the 5.6.3 source the soft wait is a plain sleep of the **full** timeout whenever a task is active when it begins; it does not end early when the last task finishes. Budget it as a fixed cost per worker stop. - Before 5.6, the prefork pool stopped sending broker heartbeats during a warm shutdown, so the broker could close the connection and running tasks could not complete. 5.6 keeps the heartbeats going. - A second `SIGTERM` during a warm shutdown does nothing; `SIGQUIT` or Ctrl-C moves the worker to the next stage. - Redelivery of an `acks_late` export is a restart, not a resume. For exports that run for many minutes, splitting the work into smaller tasks shrinks what a deploy can throw away.

  • Why not raise the platform's grace period to an hour and keep plain warm shutdown?
    It works when every task is bounded, but each deploy then waits for the slowest export on every worker, and a hung export with no time limit never finishes at all, so the stop still ends in `SIGKILL`. If you choose this route, the grace period must exceed the task's `time_limit`, which makes the hard limit the real ceiling on how long a deploy can stall.
  • Why can an idle worker holding ETA tasks lose them in a cold shutdown, and what fixes it?
    Soft shutdown is skipped when no task is actively running, so the worker goes straight to cancelling and exiting. The worker guide warns that reserved ETA tasks can be lost in that path. Setting `worker_enable_soft_shutdown_on_idle=True` forces the soft wait even when idle, giving the worker time to hand those messages back.
  • What does a second SIGTERM do while a Celery worker is in a warm shutdown?
    Nothing: additional `SIGTERM` signals are ignored during a warm shutdown. `SIGQUIT`, or another Ctrl-C, moves the worker to the next stage, a cold shutdown, preceded by a soft wait when `worker_soft_shutdown_timeout` is set.

saying these in an interview costs you the question

  • SIGTERM makes a Celery worker cancel running tasks and requeue them
  • Soft shutdown is on by default in Celery 5.5 and later
  • Setting worker_soft_shutdown_timeout alone changes what SIGTERM does
  • A cancelled task that was acknowledged early goes back to the queue
  • Cold shutdown loses acks_late tasks just like early-acknowledged ones