skip to content

In Celery, how do `autoretry_for`, `retry_backoff`, `retry_backoff_max` and `retry_jitter` combine, and what retry delays do they actually produce?

level: middleimportance: must knowfreq 48%

answer

  1. the decorator writes the except block
  2. factor times two to the retries
  3. a ceiling, then a random draw
  4. ten-minute cap, jitter on by default
  5. no backoff still means a delay

basics

~20 s

autoretry_for makes Celery call retry() when a listed exception escapes the task. retry_backoff sets the delay to factor x 2^retries seconds, capped by retry_backoff_max (600), and retry_jitter, on by default, draws a random delay between zero and the computed value.

solid answer

~40 s

`autoretry_for=(CarrierError,)` wraps the task body so a listed exception becomes `self.retry(exc=exc, **retry_kwargs)`; `dont_autoretry_for` excludes subclasses you never want retried, and `retry_kwargs` passes options such as `max_retries`. `retry_backoff=True` uses a factor of 1, so the delay ceilings are 1, 2, 4, 8 seconds and so on; a number such as `5` gives 5, 10, 20, 40. `retry_backoff_max` caps each ceiling at 600 seconds by default, and `retry_jitter`, which defaults to `True`, replaces each ceiling with a random value from 0 up to it, so a burst of failures does not retry in lockstep. Two traps: with the default `max_retries` of 3, `retry_backoff=True` gives up after at most seven seconds of waiting, and without `retry_backoff` an autoretry is not immediate: it waits the task's `default_retry_delay` of 180 seconds.

code

python · 18 lines
python
from celery import Celery

from shipping.client import CarrierError, CarrierRejectedAddress, LabelClient

app = Celery("shop", broker="amqp://guest@localhost//")


# CarrierRejectedAddress subclasses CarrierError but is never retried
@app.task(
    autoretry_for=(CarrierError,),
    dont_autoretry_for=(CarrierRejectedAddress,),
    retry_backoff=5,         # ceilings 5, 10, 20, 40, 80, 160, 300, 300 s
    retry_backoff_max=300,
    retry_jitter=True,       # the default: random delay from 0 to the ceiling
    max_retries=8,
)
def create_shipping_label(shipment_id):
    return LabelClient().create(shipment_id)

go deeper

for a junior

Recall that autoretry_for lists exceptions Celery retries for you, and that retry_backoff makes the delays grow instead of staying fixed.

for a middle

Explain the formula: factor times two to the retry count, capped by retry_backoff_max, then a random draw below it when retry_jitter is on.

for a senior

Show you size the policy to the dependency: raise max_retries with backoff, retry only transient types, and know the 180-second fallback when backoff is off.

for a principal

Treat retry policy as load control: jittered backoff protects a recovering carrier, and the total retry window should match how long a label may be late.

## What `autoretry_for` does Writing `try` / `except` / `raise self.retry(exc=exc)` in every task that calls a flaky API gets repetitive. `autoretry_for` moves that into the decorator. When the task is registered, Celery wraps its `run` method: if an exception whose type appears in `autoretry_for` escapes the body, the wrapper calls `task.retry(exc=exc, **retry_kwargs)` for you. Everything a manual `self.retry()` does still applies — a new message with the same task id, the `RETRY` state, the `max_retries` limit, and the original exception re-raised when the limit is passed. Three details shape the wrapper: - **`retry_kwargs`** — a dict passed to `retry()`, such as `{"max_retries": 8}`. - **`dont_autoretry_for`** — exception types that are never retried even when they match `autoretry_for`, so you can retry a broad `CarrierError` while failing fast on its `CarrierRejectedAddress` subclass. - **`Ignore` and `Retry` pass straight through**, so an explicit `self.retry()` or `raise Ignore()` inside an autoretry task keeps working. ## The backoff formula `retry_backoff` switches the delay from fixed to exponential. On each failure the wrapper: 1. takes the factor — `1` for `retry_backoff=True`, otherwise the number you gave, truncated to an integer and never below 1; 2. computes a ceiling of `factor × 2^retries`, where `retries` is `self.request.retries`, 0 on the first run; 3. caps that ceiling at `retry_backoff_max`, 600 seconds by default; 4. if `retry_jitter` is on — the default — replaces the ceiling with a random whole number from 0 up to the ceiling; 5. passes the result to `retry()` as `countdown`. | Retry | `retries` value | `retry_backoff=True` | `retry_backoff=5`, `retry_backoff_max=300` | |---|---|---|---| | 1st | 0 | 1 s | 5 s | | 2nd | 1 | 2 s | 10 s | | 3rd | 2 | 4 s | 20 s | | 4th | 3 | 8 s | 40 s | | 6th | 5 | 32 s | 160 s | | 7th | 6 | 64 s | 300 s (capped) | With jitter on, every value in the table is the **most** the task waits, not what it waits. ## Defaults worth memorising | Option | Default | Consequence | |---|---|---| | `autoretry_for` | `()` | Nothing is retried automatically. | | `retry_backoff` | `False` | No exponential delay. | | `retry_backoff_max` | `600` | Ceilings stop at ten minutes. | | `retry_jitter` | `True` | Delays are drawn at random below the ceiling. | | `max_retries` | `3` | Three retries, four runs. | | `default_retry_delay` | `180` | The delay when neither backoff nor a countdown sets one. | Two combinations surprise people. First, `retry_backoff=True` with the default `max_retries=3` gives up after at most 1 + 2 + 4 = 7 seconds of waiting — far too short to ride out a carrier outage. Second, without `retry_backoff` the wrapper passes no countdown, so `retry()` falls back to `default_retry_delay`: an autoretry waits 180 seconds, not zero. The Celery 5.6 user guide describes the no-backoff case as "not delayed"; the pinned source (`celery/app/autoretry.py` and `Task.retry`) shows the 180-second fallback, and the source is what runs. ## Tuning it for a flaky shipping-label API - **List transient types only**: timeouts, connection errors, and an exception your client raises for the carrier's 5xx or 429 replies. Validation errors should fail at once. - **Carve out permanent subclasses** with `dont_autoretry_for` instead of listing every transient class by hand. - **Match the factor to the outage you expect.** A factor of 5 capped at 300 seconds with eight retries has ceilings summing to 915 seconds, about fifteen minutes. - **Keep jitter on** when many label tasks fail together, so their retries spread out instead of hitting the recovering API in the same second. - **Raise `max_retries` whenever you add backoff**; the default of 3 suits a fixed three-minute delay, not a one-second one. ## Pitfalls - `autoretry_for=(Exception,)` retries programming errors and bad input as if they were outages. - `retry_backoff_max` caps each delay, not the number of attempts; `max_retries` still bounds the total. - Full jitter can draw 0, so an individual retry may fire almost at once. - These are task options with no app-wide `task_*` setting of their own; to share one policy, put the attributes on a base `Task` class and pass `base=` to each task. - Every attempt runs the whole body again, so work done before the failing call is repeated.

  • Why does `autoretry_for=(Exception,)` on a Celery task usually do more harm than good?
    It retries every bug as if it were an outage. A `KeyError` from a malformed payload or an address the carrier refuses fails the same way on every attempt, so the task burns its retry budget and the real failure surfaces minutes later. List the transient types instead, and use `dont_autoretry_for` to exclude permanent subclasses of a broad base class.
  • With `retry_jitter=True`, can a Celery autoretry fire with no delay at all?
    Yes. Full jitter draws a random whole number from 0 up to the computed ceiling, inclusive, so any attempt can wait anywhere in that range. That spreads a burst of failing tasks, which is the point, but gives no minimum pause. If every attempt needs a floor, set `retry_jitter=False` or compute your own countdown in an explicit `self.retry()`.

saying these in an interview costs you the question

  • retry_backoff=True starts from the 180-second default delay and doubles it.
  • Without retry_backoff, autoretry_for retries the task immediately with no delay.
  • retry_jitter is off by default, so delays are exact unless you enable it.
  • retry_backoff_max limits how many times the task is retried.
  • autoretry_for=(Exception,) is a safe default for any task that calls an API.