skip to content

A payouts team on Celery's Redis broker raises visibility_timeout to 12 hours so delayed payouts stop running twice; what does that fix, and what does it cost?

level: seniorimportance: must knowfreq 40%

answer

  1. unacknowledged is not the same as lost
  2. countdown messages stay unacked until due
  3. restore happens only after the window
  4. a clean shutdown requeues, a kill does not

basics

~20 s

On Celery's Redis or SQS broker, a message left unacknowledged past visibility_timeout is redelivered, so long countdowns run twice. Twelve hours stops that, but a killed worker's tasks then wait up to 12 hours for redelivery.

solid answer

~40 s

Redis has no acknowledgements, so kombu's Redis transport keeps each delivered message in an unacked set and restores it once `visibility_timeout` passes, 3600 seconds by default (kombu's SQS transport uses 1800). A task sent with `countdown` or `eta` is fetched at once and held unacknowledged in the worker until it is due, so a six-hour countdown is restored and delivered again, and again, while still waiting. Setting `broker_transport_options = {"visibility_timeout": 43200}` above the longest countdown stops those duplicates. The cost is recovery time: a clean shutdown puts reserved messages back, but a worker that is killed or loses power leaves them stranded for the full 12 hours. The better fix keeps countdowns short and stores far-future payouts in the database for a periodic task to pick up, with the payout task idempotent either way.

code

python · 14 lines
python
from celery import Celery

app = Celery("payouts", broker="redis://redis:6379/0")
# default is 3600 s; must exceed the longest countdown or eta in use
app.conf.broker_transport_options = {"visibility_timeout": 43200}  # 12 hours


@app.task
def send_payout(payout_id):
    ...


# held unacknowledged in a worker for six hours before it runs
send_payout.apply_async((42,), countdown=6 * 3600)

go deeper

for a junior

Recall that Celery's Redis broker redelivers a message nobody acknowledged within visibility_timeout, one hour by default.

for a middle

Explain why countdown and eta messages stay unacknowledged in the worker and therefore get redelivered when they wait longer than the timeout.

for a senior

Show the trade: a larger timeout stops duplicates but delays recovery after killed workers, and prefer short countdowns, due-time rows and idempotent tasks.

for a principal

Treat long delays as a scheduling problem the broker is not built for, and decide where due work should live so a crash cannot hide it for hours.

## What visibility_timeout means on Celery's Redis transport A real broker such as RabbitMQ tracks which consumer holds which message and takes it back when that consumer's connection drops. Redis has no such protocol, so **kombu**, Celery's messaging library, emulates it: 1. A worker pops the message off the queue's Redis list. 2. kombu stores a copy in an **unacked** hash, indexed by the time it was delivered. 3. When the task is acknowledged, the copy is deleted. 4. Workers periodically scan the index and **restore** every copy older than `visibility_timeout` back onto the queue. The default is **3600 seconds** (one hour). kombu's SQS transport applies the same idea through SQS's own visibility timeout and creates queues with **1800 seconds** unless told otherwise. In both cases the timeout is set through `broker_transport_options`. ## Why delayed payouts ran twice A call such as `send_payout.apply_async((payout_id,), countdown=6 * 3600)` does not park the message on the broker until it is due. The worker fetches it immediately, keeps it in memory, and acknowledges it only when execution begins. For those six hours the message is unacknowledged, so: - after one hour the restore scan puts it back on the queue; - another worker (or the same one) fetches that copy and also waits; - the cycle repeats every hour, so several copies fire when the countdown ends. The same applies to `eta` and to retries scheduled with a long countdown. Tasks running with late acknowledgement for longer than the timeout are redelivered for the same reason. ## What raising it to 12 hours fixes With `visibility_timeout` at 43200 seconds, a message that waits six hours is acknowledged long before the restore scan considers it stale, so the duplicates stop. Celery's Redis documentation gives exactly this remedy, sized to the longest ETA in use, and notes that if several apps share one broker with different values, the shortest one wins. ## What it costs The timeout is also the only recovery path for messages whose worker vanished: | Event | Default 1 hour | 12 hours | |---|---|---| | Warm shutdown (TERM, deploy) | reserved messages requeued at shutdown | same | | Worker killed with SIGKILL or out of memory | redelivered after up to 1 hour | redelivered after up to 12 hours | | Host loses power | redelivered after up to 1 hour | redelivered after up to 12 hours | So the fix trades duplicate executions for long, silent delays after crashes, which for payouts means money that should have moved hours ago. Celery 5.5 added a **soft shutdown** (`worker_soft_shutdown_timeout`, off by default) that gives a stopping worker a window to requeue its messages before a cold shutdown, which narrows but does not close the gap. On SQS there is a harder ceiling: AWS caps the visibility timeout at 12 hours, so a longer countdown cannot be protected this way at all. ## How it shows up in production The duplicate storm is easy to misread as a retry bug. The tell-tale signs: 1. The worker log shows `Task ...[<id>] received` for the **same task id** more than once, often on different workers, about one `visibility_timeout` apart. 2. Several copies of the same payout run at the moment the countdown ends, rather than spread out like retries. 3. On the Redis server, kombu's `unacked` hash and `unacked_index` sorted set stay large, because every waiting countdown occupies an entry. The crash-side symptom of an oversized timeout is quieter: after a worker is killed, its payouts simply do not run, no error is logged, and they reappear hours later. ## Better options for long delays - Keep `countdown` and `eta` to minutes, as Celery's calling guide recommends, and leave `visibility_timeout` near its default. - Store a far-future payout as a row with a due time, and let a periodic task enqueue the payouts that have become due. - Make `send_payout` idempotent, for example by checking the payout's status before paying, because every Celery transport can redeliver. - On RabbitMQ, Celery 5.5+ with quorum queues moves countdown waiting into the broker through native delayed delivery, so no worker holds the message unacknowledged while it waits. A senior answer names both sides: the setting stops a duplicate storm, and in exchange widens the window in which a crashed worker's payouts sit undelivered.

  • Why does a warm shutdown not suffer the 12-hour delay?
    On a warm shutdown the worker stops consuming, finishes what it is running and puts the messages it had reserved but not acknowledged back on the queue, so another worker picks them up at once. The visibility timeout only matters when that step never happens: a SIGKILL, an out-of-memory kill or a power loss. Soft shutdown in 5.5+ exists to give that requeue step more time.
  • Does visibility_timeout affect tasks sent by celery beat?
    Not through countdowns: Celery's docs say periodic tasks are not affected, because beat publishes each run when it is due rather than as a message with a long ETA. A beat-sent task can still be redelivered if it runs with late acknowledgement for longer than the timeout, like any other task.

saying these in an interview costs you the question

  • Redis tells kombu instantly when a worker dies, so the timeout rarely matters.
  • On a Redis broker, a countdown task waits on the broker, not in the worker, until due.
  • Raising visibility_timeout has no downside beyond extra memory in Redis.
  • A killed worker's messages are requeued during its shutdown like a TERM.
  • On SQS any countdown can be protected by raising the timeout high enough.