skip to content

A Celery task that charges a customer's card sets `acks_late=True`; why can the charge now run twice, and what do `reject_on_worker_lost` and `acks_on_failure_or_timeout` change?

level: seniorimportance: must knowfreq 42%

answer

  1. when is the message confirmed
  2. at-most-once versus at-least-once
  3. a killed child is still acked
  4. failures are acked unless told otherwise
  5. idempotency key before late acks

basics

~20 s

Celery acknowledges a message just before running it by default; acks_late=True moves the ack to after the task finishes, so a worker crash mid-charge redelivers the message and the charge runs again. A killed pool child is still acked unless reject_on_worker_lost=True.

solid answer

~50 s

By default the worker acknowledges the message when its pool accepts the task, so a crash mid-charge loses the attempt but never repeats it. `acks_late=True`, or `task_acks_late` for the whole app, acknowledges only after the task returns, raises or schedules a retry. If the worker's main process dies or its broker connection drops in between, the broker hands the message to another worker and `charge_card` runs again, possibly after the first charge went through. `acks_on_failure_or_timeout`, default `True`, still acknowledges a task that raised, so ordinary errors are not redelivered. When a prefork child is killed, by the OOM killer or `SIGKILL`, the worker acknowledges the message and stores `FAILURE` with `WorkerLostError`, unless `reject_on_worker_lost=True` requeues it. Enable late acks only after the charge is idempotent: an idempotency key per payment at the gateway and a status check before charging.

code

python · 19 lines
python
from celery import Celery

from billing import gateway
from billing.models import Payment

app = Celery("shop", broker="amqp://guest@localhost//")


@app.task(acks_late=True, reject_on_worker_lost=True)
def charge_card(payment_id):
    payment = Payment.load(payment_id)
    if payment.captured:
        return payment.charge_id           # redelivered copy: nothing to do
    charge = gateway.charge(
        amount=payment.amount,
        idempotency_key=f"payment-{payment_id}",  # gateway returns the first charge
    )
    payment.mark_captured(charge.id)
    return charge.id

go deeper

for a junior

Recall that Celery acknowledges before running by default and that acks_late moves the ack to after the task, which lets a crash cause a second run.

for a middle

Explain the acknowledgement paths: success, an exception, a killed pool child, and a dead main process, and which setting governs each.

for a senior

Show you would enable late acks only on an idempotent task, add reject_on_worker_lost deliberately, and alert on WorkerLostError before a poison task loops.

for a principal

Decide per task whether a lost attempt or a repeated one costs more, and make that choice explicit in configuration rather than a global task_acks_late.

## When Celery acknowledges a message Brokers such as RabbitMQ, Redis and Amazon SQS keep a delivered message **unacknowledged** until the consumer confirms it. Until that acknowledgement ("ack") arrives, the broker can hand the message to another consumer if the first one disappears; after it, the message is gone. The question Celery answers is only *when* the worker sends the ack. By default it sends it when the worker's pool **accepts** the task — just before the function runs. The Celery docs give the reason: the worker cannot know whether a task is **idempotent** (safe to run more than once with the same effect), so it prefers never to execute a started task twice. The price is **at-most-once** behaviour: if the worker dies halfway through `charge_card`, the message is already acknowledged and that attempt is simply lost. `acks_late=True` on the task, or `task_acks_late = True` for the app, moves the ack to after the task finishes — whether it returned, raised or scheduled a retry. That turns the task **at-least-once**. ## Why late acks make a charge run twice With late acks, the dangerous window lies between the gateway accepting the charge and the worker acknowledging the message. If the worker's main process is killed, its host reboots or its broker connection drops in that window, the message is still unacknowledged. The broker redelivers it — RabbitMQ when the connection closes, Redis and SQS once the transport's visibility timeout passes — and another worker runs `charge_card` again for the same payment. The customer is charged twice unless the task is safe to repeat. ## Worker loss versus a task that raised Once `acks_late` is on, the worker distinguishes several failure paths: 1. **The task raises an ordinary exception.** With `acks_on_failure_or_timeout=True`, the default, the message is acknowledged and the task is stored as `FAILURE`. Late acks do not turn errors into retries. 2. **`acks_on_failure_or_timeout=False`.** An ordinary failure is **rejected without requeue**; on RabbitMQ, a queue with a dead-letter exchange receives it. 3. **A prefork pool child dies** — the kernel's OOM killer, a segfault, `SIGKILL`. The worker's main process sees `WorkerLostError` and, by default, still **acknowledges** the message and stores `FAILURE`. 4. **`reject_on_worker_lost=True`.** The same child death **rejects with requeue**, so the task runs again on this or another worker. 5. **The main process itself dies.** Nothing is acknowledged, and the broker redelivers every unacknowledged message that worker held. | Settings (cumulative) | When the message is acked | Pool child killed mid-charge | Task raises `GatewayError` | |---|---|---|---| | defaults | when the task starts | already acked; `FAILURE` | already acked; `FAILURE` | | `acks_late=True` | after the task finishes | acked; `FAILURE` (`WorkerLostError`) | acked; `FAILURE` | | + `reject_on_worker_lost=True` | after the task finishes | requeued; runs again | acked; `FAILURE` | | + `acks_on_failure_or_timeout=False` | after the task finishes | requeued; runs again | rejected, not requeued | Why acknowledge a killed child by default? The docs list the reasons: a task that segfaults or exhausts memory will probably do it again, an administrator who killed it probably meant it, and a task that always dies on redelivery becomes a high-frequency message loop that can take the system down. `reject_on_worker_lost` carries that warning with it. ## Making the charge safe to repeat Late acks are only safe on an idempotent task. For a card charge that means: - **An idempotency key per payment**, such as `payment-<id>`, sent with the gateway call, so a repeated request returns the first charge instead of creating a second. - **A status check first**: if the payment is already marked captured, return without calling the gateway. - **Recording the gateway's charge id** on the payment as soon as the call returns, so the next attempt's status check sees it. `worker_deduplicate_successful_tasks` (default `False`) skips a redelivered late-ack task whose id is already `SUCCESS` in a persistent result backend. It does not replace the key: the crash that matters happens *before* `SUCCESS` is written. ## Choosing the settings - Leave `acks_late` off for tasks where losing an attempt is cheaper than repeating it. - Turn it on for the charge only together with the idempotency key. - Add `reject_on_worker_lost=True` if a killed child must not silently drop a charge, and alert on repeated `WorkerLostError` so a poison task is noticed. - Keep `self.retry()` or `autoretry_for` for gateway errors; late acks never retry an exception.

  • Does `acks_late=True` make Celery redeliver a task that raised an exception?
    No. With the default `acks_on_failure_or_timeout=True`, a task that raised is acknowledged and stored as `FAILURE`; redelivery covers only a lost worker or connection. For retries on errors you still need `self.retry()` or `autoretry_for`. Setting `acks_on_failure_or_timeout=False` does not requeue either: an ordinary failure is rejected without requeue, which a RabbitMQ dead-letter exchange can capture.
  • Why does Celery acknowledge a task whose prefork child was killed, even with `acks_late=True`?
    The docs give four reasons: a task that forced a segfault would likely do it again, an administrator who killed it probably meant to, a task that triggered the OOM killer may trigger it again, and a task that always dies on redelivery creates a high-frequency message loop. `reject_on_worker_lost=True` opts into redelivery anyway, so pair it with idempotency and alerting on repeated `WorkerLostError`.
  • Does `worker_deduplicate_successful_tasks` make a Celery card-charge task safe under `acks_late`?
    Only partly. It skips a redelivered late-ack task whose id is already `SUCCESS` in a persistent result backend. The duplicate charge comes from a crash after the gateway call but before `SUCCESS` is stored, and that window it cannot see. It saves repeated work; the idempotency key is what prevents the second charge.

A courier who signs the delivery log only after handing over the parcel: if he collapses on the doorstep, the depot sends another courier with the same parcel, and the customer may receive it twice. Signing on departure never duplicates, but a collapse means the parcel never arrives.

saying these in an interview costs you the question

  • acks_late=True makes a Celery task run exactly once.
  • With acks_late, a task that raises an exception is redelivered automatically.
  • acks_late alone redelivers a task whose prefork child was OOM-killed.
  • With default settings, a task running when its worker crashes is redelivered to another worker.
  • reject_on_worker_lost is safe to enable on any task without other changes.