skip to content

In Celery, where does a task sent with apply_async(countdown=...) wait before it runs, and why is a countdown of hours risky?

level: seniorimportance: should knowfreq 35%

answer

  1. countdown becomes an eta
  2. published now, not later
  3. held unacknowledged by a worker
  4. redelivery and memory
  5. quorum queues change it

basics

~20 s

apply_async turns countdown into an eta and publishes at once. On Redis or a classic RabbitMQ queue a worker takes the message immediately and holds it unacknowledged in memory until due, so hours-long waits pile up and can run twice.

solid answer

~40 s

The producer converts `countdown=7200` into an `eta` timestamp in the message and publishes right away. A worker consumes it at once, keeps it in its timer and only starts it after the eta: "not before", never "exactly at". Until then the message is unacknowledged. On Redis, a message unacknowledged longer than the transport's `visibility_timeout` (3600 s by default) is restored to the queue, so a two-hour countdown can run twice. On RabbitMQ, a delivery left unacknowledged past the broker's consumer timeout closes the channel with `PRECONDITION_FAILED`. Thousands of scheduled invoices also sit in worker RAM. Since Celery 5.5, quorum queues switch on native delayed delivery, so the broker holds the wait. I keep countdowns to minutes, add `expires` so a late message is revoked rather than run, and use a scheduler for anything longer.

code

python · 16 lines
python
from datetime import datetime, timedelta, timezone

from celery import Celery

app = Celery('shop', broker='redis://localhost:6379/0')

@app.task
def render_invoice_pdf(order_id):
    ...

# Start no earlier than 2 minutes from now; drop it if an hour has passed
render_invoice_pdf.apply_async((42,), countdown=120, expires=3600)

# The same start time as an aware datetime
start = datetime.now(timezone.utc) + timedelta(minutes=2)
render_invoice_pdf.apply_async((42,), eta=start)

go deeper

for a junior

Recall that countdown and eta set the earliest start time and that expires stops a late task from running.

for a middle

Explain that countdown becomes an eta, the message is published and consumed immediately, and the worker holds it unacknowledged until due.

for a senior

Show the production failures of long countdowns: worker memory, Redis visibility-timeout duplicates, RabbitMQ consumer timeouts, and how quorum queues change the picture.

for a principal

Decide where delayed work belongs: short operational delays in the queue, business schedules in a database with a periodic enqueuer, and idempotent tasks either way.

## countdown, eta and expires `apply_async` accepts three timing options that belong to one call: | Option | Meaning | Type | |---|---|---| | `countdown` | earliest start, in seconds from now | int or float | | `eta` | earliest start as an absolute time | `datetime`, preferably timezone-aware | | `expires` | after this, the task must not run | seconds from publish or a `datetime` | `countdown` is a shortcut: the producer computes `now + countdown` and stores the result as the message's `eta` (see `celery/app/amqp.py`). Pass one of the two, not both. The guarantee is **"not before"**: the task runs at some point after the eta, later if workers are busy. `expires` is the opposite bound. When a worker receives a message whose expiry has passed, or reaches its start time after the expiry, it marks the task `REVOKED` and does not run the body. For the invoice pipeline, `render_invoice_pdf.apply_async((order_id,), countdown=120, expires=3600)` means "after the payment settles, but never an hour late". ## What happens after the call On Redis, SQS or a classic RabbitMQ queue: 1. The producer publishes the message **immediately**, with the eta in its headers. 2. A worker consumes it **immediately**, like any other message, and registers it on an internal timer. 3. The message stays **unacknowledged** while it waits; the worker raises its prefetch allowance so eta messages do not block ordinary ones. 4. At the eta, the task is handed to the pool and executes; with default early acknowledgement, it is acknowledged when it starts. The broker is not a scheduler here. The wait happens inside a worker process. ## Why long countdowns hurt - **Memory.** Every waiting task is an object in a worker's memory. Scheduling tomorrow's reminder for every order puts a day of messages on the workers. - **Duplicate runs on Redis.** Kombu's Redis transport emulates acknowledgement with a **visibility timeout**, 3600 seconds by default. A message unacknowledged for longer is restored to the queue and delivered again, so a task with a two-hour countdown can be executed by two workers. Raising `visibility_timeout` in `broker_transport_options` trades this for slower recovery of genuinely lost messages. - **Channel closures on RabbitMQ.** RabbitMQ enforces a consumer acknowledgement timeout; Celery's docs cite 30 minutes in current RabbitMQ versions. A delivery held longer makes the broker close the channel with `PRECONDITION_FAILED`, disrupting the worker. - **Restarts.** A worker that stops returns its unacknowledged messages to the broker; they are not lost, but they are redelivered and scheduled again by another worker. The Celery docs therefore recommend eta and countdown only for delays of a few minutes, and a database-backed periodic scheduler for the distant future. ## Quorum queues and native delayed delivery (5.5+) Celery 5.5 added support for RabbitMQ **quorum queues**. Quorum queues do not allow the per-worker prefetch adjustment that eta handling relies on, so a waiting eta message would block the worker. When quorum queues are detected (`worker_detect_quorum_queues`, default `True`), Celery automatically enables **native delayed delivery**: a message with `countdown` or `eta` is published into a set of broker-side delay queues and only reaches the work queue when due. The wait then lives in the broker, not in worker memory, which removes the memory and unacknowledged-message problems for that setup. Redis and classic queues still behave as above. ## Reading the schedule inside the task A running task can see its own schedule through `self.request.eta` and `self.request.expires` when it is declared with `bind=True`. Both are `None` for a task sent without timing options. This is useful for logging how late an invoice actually started compared with its planned time, which is the first number to look at when customers report slow invoices: a large gap points at busy workers or a backlog rather than at the countdown itself. ## Recommendations for the invoice pipeline - Use `countdown` for short, operational delays, such as waiting two minutes for a payment webhook. - Always pair it with `expires` when a late run is worse than no run. - Pass `eta` as a timezone-aware `datetime`, as the docs do, rather than a naive local time. - For "send a reminder in three days", store the due time in the database and let a periodic job enqueue what is due, instead of a three-day countdown. - Make the task safe to run twice for the same order, because redelivery is a real outcome of long waits.

  • How does a Celery task published with expires behave if it is still queued when the time passes?
    The worker checks the expiry when it receives the message and again before execution. If the time has passed, it marks the task `REVOKED` and never runs the body. Anyone reading the result sees a revocation rather than a success or failure, which is the signal that the invoice was skipped for being late.
  • What makes a Celery countdown task safe on Redis if a delay longer than an hour is unavoidable?
    Either raise `visibility_timeout` in `broker_transport_options` above the longest countdown, accepting that genuinely lost messages take that long to return, or move the wait out of the message: store the due time and let a periodic job enqueue due work. Either way, make the task idempotent per order.

saying these in an interview costs you the question

  • The broker holds a countdown message and delivers it only when it is due, whatever the transport.
  • A countdown guarantees the task starts exactly at that moment.
  • Countdowns of several days are fine because the message is safely acknowledged.
  • An expired task still runs once and is then marked as expired.
  • Redis redelivery cannot duplicate a task with a countdown.