skip to content

In Celery, you revoke a hung PDF-export task by id, yet it keeps running, and a revoked queued export later runs anyway; why, and what are terminate's risks?

level: seniorimportance: should knowfreq 25%

answer

  1. a broadcast flag, not a kill
  2. skipped when it arrives
  3. signals a process, not a task
  4. memory that restarts forget
  5. --statedb

basics

~20 s

revoke only broadcasts the id to workers, which skip that task when they reach it; a running task continues unless terminate=True signals its pool process. Revoked ids live in worker memory, so a full restart forgets them unless workers use --statedb.

solid answer

~50 s

`app.control.revoke(task_id)`, or `AsyncResult.revoke()`, broadcasts the id; each worker adds it to an in-memory revoked set and discards the task if it later receives it. A task already executing keeps going. `terminate=True`, or `celery -A proj control terminate SIGTERM <id>`, signals the pool **process** running it, `TERM` by default, and only prefork, eventlet and gevent support it; that process may already have moved on to another task, so the docs call it a last resort, never to be called from code. The revoked set is capped at 50,000 ids with a 3-hour expiry, and a restart of every worker empties it unless `--statedb` persists it. Revoke rides on remote control, so it works on RabbitMQ and Redis, not SQS. Since 5.6, workers also mark the task `REVOKED` in the result backend as soon as they get the command.

code

python · 9 lines
python
from proj.celery import app

task_id = "d9078da5-9915-40a0-bfa1-392c7bde42ed"

# not started yet: workers discard it when they reach it
app.control.revoke(task_id)

# already running and stuck: signal the pool process running it
app.control.revoke(task_id, terminate=True, signal="SIGTERM")

go deeper

for a junior

Recall that revoke tells workers to skip a task that has not started, and that stopping a running one needs terminate=True.

for a middle

Explain that revoke is a broadcast held in each worker's memory, that terminate signals the pool process, and which pools support it.

for a senior

Know why revoked tasks resurface after full restarts or past the three-hour expiry, use --statedb, target terminate at one node after inspect active, and replace manual kills with time limits.

for a principal

Decide how the system cancels work by design, through expiry, limits and application flags, rather than relying on revoke, which some transports do not support at all.

## What revoke actually does `app.control.revoke(task_id)`, or `AsyncResult(task_id).revoke()`, is a **remote-control broadcast**, not an operation on the broker's queue. Step by step: 1. The id is broadcast to every worker, or to the nodes named in `destination`. 2. Each worker adds it to its in-memory **revoked set**. 3. When a worker later receives that task, whether from the queue, from its prefetched reserve or from its ETA timer, it finds the id in the set and discards the task instead of running it. 4. Since Celery 5.6, a worker handling the command also marks the task `REVOKED` in the result backend straight away. Earlier releases wrote that state only when a worker reached the task, so ETA tasks could sit in `PENDING` indefinitely if workers stopped first. Nothing in those steps touches a task that is **already executing**. That is the first half of the puzzle: the hung export keeps running because a plain revoke only stops tasks that have not started. ## Revoke versus terminate | | `revoke(id)` | `revoke(id, terminate=True)` | |---|---|---| | Queued, reserved or ETA task | Skipped when reached | Skipped when reached | | Task already running | Keeps running | Its pool process receives a signal, `TERM` by default | | Pool support | All pools | Terminate only on prefork, eventlet and gevent | | CLI | `celery -A proj control revoke <id>` | `celery -A proj control terminate SIGTERM <id>` | `terminate` does not target a task; it targets the **process** executing it. By the time the signal lands, that process may have finished the export and started another task, which is then killed instead. That is why the worker guide calls it a last resort for an administrator with a stuck task and says never to call it programmatically. `signal='SIGKILL'` makes it unconditional, but leaves no chance for cleanup. ## Why a revoked task can still run The revoked set is **worker memory**, and it has limits: - **A full restart forgets it.** A starting worker copies revoked ids from running peers during its startup sync, but if every worker restarts together, or they start with `--without-mingle`, the set begins empty and a revoked export still in the queue runs when received. - **It is bounded.** It holds at most 50,000 ids (`CELERY_WORKER_REVOKES_MAX`), and ids expire after 10,800 seconds, three hours (`CELERY_WORKER_REVOKE_EXPIRES`). An export revoked with a countdown longer than that can outlive its revocation. - **It needs remote control.** Revoke works on the RabbitMQ and Redis transports; on SQS, which has no remote control, the broadcast reaches nobody. Starting workers with `--statedb=/var/run/celery/%n.state` persists the set to a file per node (`%n` expands to the node name), so it survives restarts. ## Stopping the hung export safely 1. Run `celery -A proj inspect active` and find the `exports.render_pdf` entry: its `id`, its node and its `worker_pid`. 2. Revoke it first, so no redelivered or duplicate copy of that id runs later. 3. If it must stop now, terminate it on that node only: `celery -A proj control -d celery@worker-pdf-1 terminate SIGTERM <id>`. 4. Check with `inspect active` that it is gone, and with `inspect revoked` that the id is recorded. 5. Then fix the reason it hung: a `time_limit` on the task makes the next occurrence end by itself, with no human racing the pool. ## Revoking many at once - `revoke` accepts a list of ids, and `GroupResult.revoke()` uses that to revoke every task in a group. - `revoke_by_stamped_headers` revokes tasks carrying a stamped header value instead of naming ids; its mapping is not persistent across restarts at all. - Revoking is not purging. The messages stay in the queue until a worker receives and discards them, so a large revoked backlog still has to be read off the broker. ## When revoke is the wrong tool - For a task that hangs regularly, revoke and terminate are manual repairs; time limits are the lasting fix. - For tasks that must never run after a cutoff, `expires` on the call is enforced by the worker from the message itself and needs no worker memory. - For SQS deployments, plan without remote control: expiry, time limits and application-level cancellation flags checked by the task.

  • Why does the worker guide say never to call terminate programmatically?
    Because `terminate` signals a pool process, not a task. Between reading the task's state and the signal arriving, the process may finish the export and start another task, which is killed instead. A human at a console checking `inspect active` first narrows that race; code firing on every timeout would hit it regularly. Time limits are the automated tool.
  • How would you cancel exports that must never run after a cutoff, on SQS where revoke does not work?
    Send them with `expires`, which the worker checks from the message itself when it receives the task, so no remote control or worker memory is involved. Add a `time_limit` for runs that start and hang, and, for business-level cancellation, have the task check a flag in the database before doing expensive work.

saying these in an interview costs you the question

  • revoke stops a task that is already running on a worker
  • terminate=True raises an exception inside the task so it can clean up
  • Revoked ids are stored in the broker, so they survive every restart
  • Revoke works the same on every transport, including SQS
  • terminate is safe to call from application code on every timeout