In Celery, how does `self.retry()` re-run a failing task, and what do its `countdown`, `max_retries` and `exc` arguments control?
answer
- bind=True hands you the task
- it raises, so nothing after runs
- new message, same task id
- defaults: max_retries and default_retry_delay
- exc comes back when retries run out
basics
~20 sA bound Celery task calls raise self.retry(exc=exc, countdown=N) in its except block: Celery publishes a new message for the same task id, delayed N seconds (180 by default), up to max_retries (3), then re-raises exc as the failure.
solid answer
~40 sDeclare the task with `bind=True` so its first argument is the task instance, catch the transient error, and `raise self.retry(exc=exc, countdown=60)`. `retry()` publishes a fresh message for the same task id to the same queue, stores the `RETRY` state if a result backend is configured, and raises the `Retry` exception so the worker ends this run without marking it failed. `countdown` is the delay in seconds (`eta` takes an absolute time); with neither, the task's `default_retry_delay` of 180 seconds applies. `max_retries` defaults to 3, so the body runs at most four times; past the limit, `retry()` re-raises `exc` and the task ends in `FAILURE`, or raises `MaxRetriesExceededError` if no `exc` was given. Retrying forever takes `max_retries=None` on the task itself, not on the `retry()` call.
code
python · 14 linesfrom celery import Celery
from shipping.client import CarrierTimeout, LabelClient
app = Celery("shop", broker="amqp://guest@localhost//")
@app.task(bind=True, max_retries=5)
def create_shipping_label(self, shipment_id):
try:
return LabelClient().create(shipment_id)
except CarrierTimeout as exc:
# waits 30, 60, 90, 120, 150 s; the sixth failure re-raises exc
raise self.retry(exc=exc, countdown=30 * (self.request.retries + 1))go deeper
Recall the shape: bind=True, catch the transient error, raise self.retry(exc=exc, countdown=...). Know the two defaults, three retries and a 180-second delay.
Explain that a retry is a new message with the same task id, that the RETRY state is stored, and why max_retries counts retries rather than runs.
Show where retries go wrong in production: swallowed Retry exceptions, retried permanent errors, and side effects that repeat on every attempt.
Frame retry limits and delays as a policy per dependency: how long a carrier outage the business will wait out, and who is alerted when it lasts longer.
## What `self.retry()` is for A Celery task often calls something outside your process — here, a carrier's API that creates a **shipping label**. That API sometimes times out or returns a temporary error even though the request is fine. Failing the task outright would force someone to re-queue it by hand; trying again a minute later usually just works. `Task.retry()` is Celery's built-in way to say "this attempt failed for a transient reason; schedule another one". To reach `retry()` from inside the function you need the task instance. Declaring the task with `@app.task(bind=True)` passes that instance as the first argument, conventionally named `self`. It also exposes `self.request`, the metadata of the current execution, including `self.request.retries` — the number of retries already made, 0 on the first run. An unbound task can call `retry()` on the task object itself (`create_shipping_label.retry(...)`); `bind=True` is simply the common form. ## How a retry travels When the task runs `raise self.retry(exc=exc, countdown=60)`, Celery does the following: 1. It computes the next retry count, `self.request.retries + 1`, and compares it with `max_retries`. 2. If the limit is not exceeded, it builds a signature from the current request — same arguments, **same task id**, same queue — and publishes it as a **new message** with the requested delay. 3. It stores the **`RETRY`** state in the result backend, if one is configured, together with the exception that caused it. 4. It raises the **`Retry`** exception. The worker treats `Retry` as a signal rather than an error: this run ends, it is not recorded as a failure, and the original message is acknowledged. Because `retry()` raises unless you pass `throw=False`, no code after it runs. The `raise` in front of the call is a convention that makes this obvious to readers and to linters. ## The arguments | Argument | Default | What it controls | |---|---|---| | `countdown` | `None` | Seconds before the next attempt. With neither `countdown` nor `eta`, the task's `default_retry_delay` applies: **180 seconds**. | | `eta` | `None` | An absolute `datetime` for the next attempt instead of a relative delay. | | `max_retries` | the task's `max_retries`, **3** | Overrides the limit for this call. `None` here means "use the task's default", not "forever". | | `exc` | `None` | The exception recorded with the `RETRY` state and re-raised when the retries run out. | | `args` / `kwargs` | the current ones | Replace the arguments for the next attempt. | | `throw` | `True` | Whether to raise `Retry`; `False` returns it and lets the function continue. | `max_retries` and `default_retry_delay` are also task options: `@app.task(bind=True, max_retries=5, default_retry_delay=30)`. To retry forever, set `max_retries=None` on the task itself; passing `max_retries=None` to `retry()` keeps the task's limit. ## When the retries run out `max_retries` counts **retries, not runs**. With the default of 3 the body runs up to four times: the original run plus three retries. On the fourth failure, `retry()` finds the next count (4) above the limit and, instead of publishing, raises: - the `exc` you passed, so the task ends in **`FAILURE`** with the real cause — a `CarrierTimeout`, say; or - **`MaxRetriesExceededError`** if you passed no `exc`. Passing `exc` is therefore worth the habit: the stored failure then names the carrier problem rather than the retry mechanism. ## Mistakes that bite - **A broad `except Exception` around the call.** `Retry` subclasses `Exception`, so the handler swallows it after the new message is already published. This run ends as a success, and the retry still executes later. - **Retrying permanent errors.** An address the carrier rejects is rejected on every attempt; retrying only delays the failure by several minutes. - **Assuming no delay.** Leaving out `countdown` means 180 seconds, not an immediate retry. - **Repeating side effects.** Each attempt runs the whole body again, so anything done before the failing call — writing a row, notifying a warehouse — happens again. - **Reading `max_retries` as total runs.** A limit of 3 allows four executions. ## A worked example The label task in the code example retries carrier timeouts up to five times, waiting 30, 60, 90, 120 and then 150 seconds, because the countdown grows with `self.request.retries`. If the carrier is still timing out on the sixth run, `retry()` re-raises the `CarrierTimeout`, the task ends in `FAILURE`, and the stored traceback points at the carrier call — exactly what the person debugging it needs to see.
- What happens if a broad `except Exception:` block wraps the `self.retry()` call in a Celery task?`Retry` subclasses `Exception` through `CeleryError`, so the broad handler catches it. By then `retry()` has already published the next message, so the current run returns normally and is stored as `SUCCESS`, and the retry still runs later under the same task id. Catch only the transient exception types, or re-raise `Retry` before any catch-all.
- How do you make a Celery task retry forever, and why does `self.retry(max_retries=None)` not do it?In the `retry()` call, `max_retries=None` means "use the task's own limit", so the default of 3 still applies. Set it on the task instead: `@app.task(bind=True, max_retries=None)`. Pair it with a delay and alerting, because an endless retry of a failure that never clears keeps re-queuing the task indefinitely.
- Inside a bound Celery task, how do you know which attempt is running, and what is it useful for?`self.request.retries` is 0 on the first run and grows by one with every retry. Use it to scale the countdown, as in `countdown=30 * (self.request.retries + 1)`, to log which attempt failed, or to switch strategy on a late attempt, such as trying a fallback carrier.
saying these in an interview costs you the question
- self.retry() re-runs the task body in place, inside the same worker call.
- max_retries=3 means the task body runs three times in total.
- Calling self.retry() without a countdown retries the task immediately.
- Passing max_retries=None to self.retry() makes the task retry forever.
- When retries run out, Celery drops the task silently with no error recorded.