skip to content

Monitoring & Shutdown

Inspecting workers with celery inspect and control, Flower and task events, soft and hard time limits, and warm, soft and cold shutdown. Asked because a task with no ceiling pins a worker.

on this pageshow

explore

questions

5

A Celery PDF-export task sometimes hangs on a remote font server; how do soft_time_limit, time_limit and SoftTimeLimitExceeded stop it pinning a worker?

level: middleimportance: must knowfreq 40%

answer

  1. no ceiling by default
  2. an exception first, a kill second
  3. catchable vs uncatchable
  4. the pool process is replaced
  5. prefork enforces both

basics

~20 s

Celery sets no time limit by default, so a hung task holds its worker slot forever. soft_time_limit raises SoftTimeLimitExceeded inside the task so it can clean up; time_limit kills and replaces the pool process and fails the task with TimeLimitExceeded.

solid answer

~40 s

By default `task_time_limit` and `task_soft_time_limit` are unset, so an export blocked on a dead font server occupies one concurrency slot until someone intervenes. Set limits per task (`@app.task(soft_time_limit=120, time_limit=150)`), per call in `apply_async`, globally, or with the worker's `--soft-time-limit` and `--time-limit`. When the soft limit passes, Celery raises `SoftTimeLimitExceeded` inside the task: catch it, delete the half-written file and re-raise. If the task is still running at the hard limit, the prefork pool kills that child process, starts a new one and records the task as failed with `TimeLimitExceeded`; no task code runs, and with `acks_late` the message is still acknowledged because `task_acks_on_failure_or_timeout` defaults to True. Keep the soft limit below the hard one, and give the font request its own client timeout.

code

python · 15 lines
python
from celery import Celery
from celery.exceptions import SoftTimeLimitExceeded

app = Celery("proj", broker="redis://localhost:6379/0")


@app.task(soft_time_limit=120, time_limit=150)
def render_pdf(export_id):
    path = f"/tmp/export-{export_id}.pdf"
    try:
        fonts = fetch_fonts(timeout=10)  # client timeout: the first defence
        build_pdf(export_id, fonts, path)
    except SoftTimeLimitExceeded:
        remove_partial(path)  # 30 s remain before the hard kill
        raise  # task ends as FAILURE, the slot frees at once

go deeper

for a junior

Recall that Celery tasks have no time limit by default, and that soft_time_limit raises an exception while time_limit kills the task outright.

for a middle

Explain the sequence: SoftTimeLimitExceeded inside the task for cleanup, then the hard kill that replaces the pool process and records TimeLimitExceeded, and why the soft limit must sit below the hard one.

for a senior

Size limits from real runtimes, pair them with client timeouts, and know that a hard-limit kill with acks_late is still acknowledged by default, so the task is not retried by redelivery.

for a principal

Set a house rule that every task declares limits and every outbound call a timeout, and decide which failures should be retried rather than simply failing at the ceiling.

## Why a hung task pins a worker A Celery worker runs a fixed number of tasks at once, its **concurrency**. Each running task occupies one slot until its function returns. If `render_pdf` opens a connection to a remote font server that accepts the socket and then never answers, the function never returns and the slot is gone. Celery sets **no time limit by default**: `task_time_limit` and `task_soft_time_limit` are both unset. Four stuck exports on a four-slot worker and that worker processes nothing else, while the messages it prefetched wait behind the hung ones. `celery -A proj inspect active` shows such tasks with a `time_start` far in the past, which is often how a hang is first noticed. ## Soft and hard limits | | `soft_time_limit` | `time_limit` | |---|---|---| | What happens | `SoftTimeLimitExceeded` is raised inside the task | The pool process is killed and replaced | | Can the task react? | Yes: catch it and clean up | No: no task code runs afterwards | | Worker log | `Soft time limit (...s) exceeded` warning | `Hard time limit (...s) exceeded` error | | Outcome if not handled | `FAILURE` with `SoftTimeLimitExceeded` | `FAILURE` with `TimeLimitExceeded` | Both exceptions are importable from `celery.exceptions`. The soft limit is a warning with a deadline behind it; the hard limit is the deadline. ## Where to set them - On the task: `@app.task(soft_time_limit=120, time_limit=150)`. - Per call: `render_pdf.apply_async(args=[42], soft_time_limit=300, time_limit=330)` for an unusually large export. - As app-wide defaults: `task_soft_time_limit` and `task_time_limit` in the configuration. - On the worker command line: `--soft-time-limit` and `--time-limit`. - At run time: `app.control.time_limit('exports.render_pdf', soft=60, hard=90, reply=True)` changes the limits on live workers, but only tasks that **start** afterwards are affected, and it needs remote control, which the RabbitMQ and Redis transports provide. ## What happens, step by step With `soft_time_limit=120, time_limit=150` on the default prefork pool: 1. At 120 seconds the pool signals the child process running the task, and `SoftTimeLimitExceeded` is raised inside the running task. 2. The task's `except SoftTimeLimitExceeded:` block deletes the half-written PDF and re-raises, so the task ends as `FAILURE` and the slot frees at once. 3. If the task swallows the exception, or its cleanup hangs too, at 150 seconds the pool kills that child process, starts a fresh one, and the worker records the task as failed with `TimeLimitExceeded`. 4. With `acks_late=True`, the message is still acknowledged on that timeout, because `task_acks_on_failure_or_timeout` defaults to `True`. The export is **not** redelivered. ## Pool support and caveats - The **prefork** pool enforces both limits; it is the default pool and the case the documentation describes. - The **gevent** pool does not implement soft limits, and does not enforce the hard limit while a task blocks without yielding. - The worker guide lists time-limit support for prefork and gevent only, so do not assume the other pools enforce them. - Time limits rely on the `SIGUSR1` signal and do not work on platforms that lack it. - `AsyncResult.get(timeout=30)` is **client side**: the caller stops waiting and gets a timeout error, while the task keeps running on the worker. ## Choosing the numbers - Put the soft limit above the slowest **legitimate** export, judged from real runtimes rather than the average, and the hard limit far enough above it for cleanup to finish. - Keep the soft limit **below** the hard one. Without a gap between them there is no window for cleanup, and the kill arrives before the handler can finish. - Treat limits as the backstop, not the fix. The font fetch should carry its own client timeout, so the task fails fast with a clear error instead of waiting two minutes for Celery to step in. - Never write a broad exception handler that swallows `SoftTimeLimitExceeded` and carries on: the task then runs straight into the hard kill with no cleanup at all. - Remember the hard limit is a process kill. Anything the task held in that process, such as an open temporary file or a half-sent upload, is abandoned rather than closed, which is exactly why the soft limit exists. - Watch the worker log for `Soft time limit` warnings. A steady trickle means either the limit is too tight for real exports or the font server is degrading, and both are worth knowing before hard kills start.

  • What happens if the task catches SoftTimeLimitExceeded and simply carries on?
    The soft limit fires once. The task keeps running until the hard limit, when the pool kills its process and the worker records `TimeLimitExceeded`. Nothing in the task gets to clean up at that point, so swallowing the soft exception trades a clean failure for a messier one. Catch it only to clean up, then re-raise or return.
  • Does AsyncResult.get(timeout=30) in the caller put a ceiling on the task?
    No. The timeout only limits how long the caller waits; when it expires the caller gets a timeout error and the task keeps running on the worker, still holding its slot. Only the task's own `soft_time_limit` and `time_limit`, enforced by the worker's pool, bound how long it runs.
  • Can you tighten the export's limits without restarting the workers?
    Yes: `app.control.time_limit('exports.render_pdf', soft=60, hard=90, reply=True)` broadcasts new limits to live workers. Only tasks that start after the change are affected, and it relies on remote control, available on the RabbitMQ and Redis transports. Put the same values in code too, or the next deploy restores the old ones.

A library at closing time: first the announcement that it closes in ten minutes, so readers can pack up and return books; then the lights go off, and whatever was still open stays open. The soft limit is the announcement, the hard limit the lights.

saying these in an interview costs you the question

  • Celery gives every task a default time limit, so hangs end by themselves
  • time_limit raises an exception the task can catch to clean up
  • AsyncResult.get(timeout=...) stops the task on the worker when it expires
  • A task killed by the hard limit is redelivered because acks_late is set
  • Every pool, gevent included, enforces the soft time limit
open as a page

In Celery, how do you list what each worker is running right now, and what do inspect active, reserved and scheduled each show?

level: juniorimportance: should knowfreq 35%

basics

~20 s

celery -A proj inspect active asks every worker, over the broker, which tasks it is executing. inspect reserved lists tasks it prefetched but has not started, inspect scheduled the ETA or countdown tasks it holds; messages still queued appear in none.

open as a page

In Celery, what are task events and the worker's -E flag, and what does Flower need from the workers to show a task's history?

level: middleimportance: should knowfreq 30%

basics

~20 s

Task events are messages a Celery worker publishes as each task is received, started, succeeds or fails. They are off by default; celery worker -E turns them on. Flower builds its task list from that stream, so without events it shows no task history.

open as a page

In Celery, you revoke a hung PDF-export task by id, yet it keeps running, and a revoked queued export later runs anyway; why, and what are terminate's risks?

level: seniorimportance: should knowfreq 25%

basics

~20 s

revoke only broadcasts the id to workers, which skip that task when they reach it; a running task continues unless terminate=True signals its pool process. Revoked ids live in worker memory, so a full restart forgets them unless workers use --statedb.

open as a page

During a rolling deploy, a Celery 5.6 worker gets SIGTERM mid-way through a long PDF export and is killed after the grace period; what do warm, soft and cold shutdown do, and how do you keep the task?

level: seniorimportance: should knowfreq 30%

basics

~20 s

SIGTERM starts a warm shutdown that waits, unbounded, for running tasks, so the later SIGKILL loses them. Cold shutdown (SIGQUIT) cancels them; soft shutdown (5.5+) first waits worker_soft_shutdown_timeout. Use REMAP_SIGTERM=SIGQUIT, a timeout under the grace period, and acks_late.

open as a page