skip to content

Retries & Acknowledgement

self.retry(), autoretry_for with backoff, task states, and acks_late with reject_on_worker_lost. Interviewers pair acks_late with idempotency: redelivery is how a task runs twice.

on this pageshow

explore

questions

5

In Celery, how does `self.retry()` re-run a failing task, and what do its `countdown`, `max_retries` and `exc` arguments control?

level: juniorimportance: must knowfreq 62%

answer

  1. bind=True hands you the task
  2. it raises, so nothing after runs
  3. new message, same task id
  4. defaults: max_retries and default_retry_delay
  5. exc comes back when retries run out

basics

~20 s

A bound Celery task calls raise self.retry(exc=exc, countdown=N) in its except block: Celery publishes a new message for the same task id, delayed N seconds (180 by default), up to max_retries (3), then re-raises exc as the failure.

solid answer

~40 s

Declare the task with `bind=True` so its first argument is the task instance, catch the transient error, and `raise self.retry(exc=exc, countdown=60)`. `retry()` publishes a fresh message for the same task id to the same queue, stores the `RETRY` state if a result backend is configured, and raises the `Retry` exception so the worker ends this run without marking it failed. `countdown` is the delay in seconds (`eta` takes an absolute time); with neither, the task's `default_retry_delay` of 180 seconds applies. `max_retries` defaults to 3, so the body runs at most four times; past the limit, `retry()` re-raises `exc` and the task ends in `FAILURE`, or raises `MaxRetriesExceededError` if no `exc` was given. Retrying forever takes `max_retries=None` on the task itself, not on the `retry()` call.

code

python · 14 lines
python
from celery import Celery

from shipping.client import CarrierTimeout, LabelClient

app = Celery("shop", broker="amqp://guest@localhost//")


@app.task(bind=True, max_retries=5)
def create_shipping_label(self, shipment_id):
    try:
        return LabelClient().create(shipment_id)
    except CarrierTimeout as exc:
        # waits 30, 60, 90, 120, 150 s; the sixth failure re-raises exc
        raise self.retry(exc=exc, countdown=30 * (self.request.retries + 1))

go deeper

for a junior

Recall the shape: bind=True, catch the transient error, raise self.retry(exc=exc, countdown=...). Know the two defaults, three retries and a 180-second delay.

for a middle

Explain that a retry is a new message with the same task id, that the RETRY state is stored, and why max_retries counts retries rather than runs.

for a senior

Show where retries go wrong in production: swallowed Retry exceptions, retried permanent errors, and side effects that repeat on every attempt.

for a principal

Frame retry limits and delays as a policy per dependency: how long a carrier outage the business will wait out, and who is alerted when it lasts longer.

## What `self.retry()` is for A Celery task often calls something outside your process — here, a carrier's API that creates a **shipping label**. That API sometimes times out or returns a temporary error even though the request is fine. Failing the task outright would force someone to re-queue it by hand; trying again a minute later usually just works. `Task.retry()` is Celery's built-in way to say "this attempt failed for a transient reason; schedule another one". To reach `retry()` from inside the function you need the task instance. Declaring the task with `@app.task(bind=True)` passes that instance as the first argument, conventionally named `self`. It also exposes `self.request`, the metadata of the current execution, including `self.request.retries` — the number of retries already made, 0 on the first run. An unbound task can call `retry()` on the task object itself (`create_shipping_label.retry(...)`); `bind=True` is simply the common form. ## How a retry travels When the task runs `raise self.retry(exc=exc, countdown=60)`, Celery does the following: 1. It computes the next retry count, `self.request.retries + 1`, and compares it with `max_retries`. 2. If the limit is not exceeded, it builds a signature from the current request — same arguments, **same task id**, same queue — and publishes it as a **new message** with the requested delay. 3. It stores the **`RETRY`** state in the result backend, if one is configured, together with the exception that caused it. 4. It raises the **`Retry`** exception. The worker treats `Retry` as a signal rather than an error: this run ends, it is not recorded as a failure, and the original message is acknowledged. Because `retry()` raises unless you pass `throw=False`, no code after it runs. The `raise` in front of the call is a convention that makes this obvious to readers and to linters. ## The arguments | Argument | Default | What it controls | |---|---|---| | `countdown` | `None` | Seconds before the next attempt. With neither `countdown` nor `eta`, the task's `default_retry_delay` applies: **180 seconds**. | | `eta` | `None` | An absolute `datetime` for the next attempt instead of a relative delay. | | `max_retries` | the task's `max_retries`, **3** | Overrides the limit for this call. `None` here means "use the task's default", not "forever". | | `exc` | `None` | The exception recorded with the `RETRY` state and re-raised when the retries run out. | | `args` / `kwargs` | the current ones | Replace the arguments for the next attempt. | | `throw` | `True` | Whether to raise `Retry`; `False` returns it and lets the function continue. | `max_retries` and `default_retry_delay` are also task options: `@app.task(bind=True, max_retries=5, default_retry_delay=30)`. To retry forever, set `max_retries=None` on the task itself; passing `max_retries=None` to `retry()` keeps the task's limit. ## When the retries run out `max_retries` counts **retries, not runs**. With the default of 3 the body runs up to four times: the original run plus three retries. On the fourth failure, `retry()` finds the next count (4) above the limit and, instead of publishing, raises: - the `exc` you passed, so the task ends in **`FAILURE`** with the real cause — a `CarrierTimeout`, say; or - **`MaxRetriesExceededError`** if you passed no `exc`. Passing `exc` is therefore worth the habit: the stored failure then names the carrier problem rather than the retry mechanism. ## Mistakes that bite - **A broad `except Exception` around the call.** `Retry` subclasses `Exception`, so the handler swallows it after the new message is already published. This run ends as a success, and the retry still executes later. - **Retrying permanent errors.** An address the carrier rejects is rejected on every attempt; retrying only delays the failure by several minutes. - **Assuming no delay.** Leaving out `countdown` means 180 seconds, not an immediate retry. - **Repeating side effects.** Each attempt runs the whole body again, so anything done before the failing call — writing a row, notifying a warehouse — happens again. - **Reading `max_retries` as total runs.** A limit of 3 allows four executions. ## A worked example The label task in the code example retries carrier timeouts up to five times, waiting 30, 60, 90, 120 and then 150 seconds, because the countdown grows with `self.request.retries`. If the carrier is still timing out on the sixth run, `retry()` re-raises the `CarrierTimeout`, the task ends in `FAILURE`, and the stored traceback points at the carrier call — exactly what the person debugging it needs to see.

  • What happens if a broad `except Exception:` block wraps the `self.retry()` call in a Celery task?
    `Retry` subclasses `Exception` through `CeleryError`, so the broad handler catches it. By then `retry()` has already published the next message, so the current run returns normally and is stored as `SUCCESS`, and the retry still runs later under the same task id. Catch only the transient exception types, or re-raise `Retry` before any catch-all.
  • How do you make a Celery task retry forever, and why does `self.retry(max_retries=None)` not do it?
    In the `retry()` call, `max_retries=None` means "use the task's own limit", so the default of 3 still applies. Set it on the task instead: `@app.task(bind=True, max_retries=None)`. Pair it with a delay and alerting, because an endless retry of a failure that never clears keeps re-queuing the task indefinitely.
  • Inside a bound Celery task, how do you know which attempt is running, and what is it useful for?
    `self.request.retries` is 0 on the first run and grows by one with every retry. Use it to scale the countdown, as in `countdown=30 * (self.request.retries + 1)`, to log which attempt failed, or to switch strategy on a late attempt, such as trying a fallback carrier.

saying these in an interview costs you the question

  • self.retry() re-runs the task body in place, inside the same worker call.
  • max_retries=3 means the task body runs three times in total.
  • Calling self.retry() without a countdown retries the task immediately.
  • Passing max_retries=None to self.retry() makes the task retry forever.
  • When retries run out, Celery drops the task silently with no error recorded.
open as a page

In Celery, how do `autoretry_for`, `retry_backoff`, `retry_backoff_max` and `retry_jitter` combine, and what retry delays do they actually produce?

level: middleimportance: must knowfreq 48%

basics

~20 s

autoretry_for makes Celery call retry() when a listed exception escapes the task. retry_backoff sets the delay to factor x 2^retries seconds, capped by retry_backoff_max (600), and retry_jitter, on by default, draws a random delay between zero and the computed value.

open as a page

A Celery task that charges a customer's card sets `acks_late=True`; why can the charge now run twice, and what do `reject_on_worker_lost` and `acks_on_failure_or_timeout` change?

level: seniorimportance: must knowfreq 42%

basics

~20 s

Celery acknowledges a message just before running it by default; acks_late=True moves the ack to after the task finishes, so a worker crash mid-charge redelivers the message and the charge runs again. A killed pool child is still acked unless reject_on_worker_lost=True.

open as a page

Why does a Celery task's `AsyncResult.state` still read `PENDING` while a worker is running it, and which states can the task report?

level: middleimportance: should knowfreq 36%

basics

~20 s

Celery writes STARTED only when task_track_started, or the task's track_started, is enabled, and it is off by default; PENDING just means the backend has no record of the id. Built-in states: PENDING, STARTED, RETRY, SUCCESS, FAILURE, REVOKED.

open as a page

In Celery, when should a task raise `Reject` or `Ignore` instead of calling `self.retry()`, and what happens to the message and its stored state?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Use self.retry() for transient failures: a new message runs the task again. Raise Reject for a message that should leave the queue, dead-lettered or requeued, which needs acks_late. Raise Ignore when the work is moot: it acks and records nothing.

open as a page