Why does a Celery chord need a result backend, and what happens to the chord body when one thumbnail task in its header fails?
answer
- who counts the finished header tasks
- counter or chord_unlock polling
- the rest still run
- ChordError on the body
basics
~20 sA Celery chord needs a result backend because something must record each header task's completion and result before the body can start. If one header task fails, the others still run, but the body never runs and is marked FAILURE with a ChordError.
solid answer
~50 sThe broker only delivers messages; it does not know when a group of tasks has finished. A chord's body runs only after every header task is done, so Celery joins through the **result backend**: the Redis and Memcached backends keep a per-chord counter updated as each header task returns, while backends without that support, such as the SQLAlchemy database backend, schedule the built-in `celery.chord_unlock` task to poll the group, every `result_chord_retry_interval` (1 second by default). With no backend configured, starting a chord raises `NotImplementedError`, and the `rpc` backend does not support chords at all. When one `make_thumbnail` in the header raises, the other header tasks are not cancelled; once they have all finished, the join fails, `publish_photo` never runs, and its result is stored as `FAILURE` with `ChordError('Dependency <id> raised ...')`. Handle that with an errback linked to the body.
code
python · 22 linesfrom celery import Celery, chord
app = Celery('photos', broker='redis://localhost:6379/0',
backend='redis://localhost:6379/1')
@app.task
def make_thumbnail(photo_id, size):
if size == 640:
raise OSError('corrupt image')
return f'thumbs/{photo_id}_{size}.jpg'
@app.task
def publish_photo(thumb_paths, photo_id):
return photo_id # never runs in this example
@app.task
def mark_upload_failed(request, exc, traceback, photo_id):
print(f'photo {photo_id} failed: {exc!r}') # exc is a ChordError
body = publish_photo.s(42).on_error(mark_upload_failed.s(42))
res = chord(make_thumbnail.s(42, s) for s in (160, 640, 1280))(body)
# res.get() raises ChordError: Dependency <id> raised OSError('corrupt image')go deeper
Recall that a chord needs a result backend and that a failed header task stops the body from running.
Explain the two join mechanisms, native counting versus chord_unlock polling, and what a missing or rpc backend does when a chord starts.
Walk the failure path: header siblings keep running, the body fails with ChordError, errbacks fire on the body, and header errbacks need task_allow_error_cb_on_chord_header.
Weigh all-or-nothing chords against header tasks that return a status, and the backend load and latency chords add, when choosing a workflow design.
## Why the join needs storage A **chord** is a **header** group of tasks plus a **body** task that runs once all of the header has finished, receiving the list of header results. The broker (RabbitMQ, Redis or SQS as a transport) cannot provide that: it delivers each message to a worker and forgets it. Something has to **remember** which header tasks have finished and what they returned. In Celery that is the **result backend**, the store configured with `result_backend`. This is also why Celery's guide says tasks used in a chord must **not ignore their results**: on a backend that joins by polling, a header result that is never stored never shows up, and the body waits forever. ## Two ways Celery joins a chord | Backend | Join mechanism | |---|---| | Redis, Memcached (cache backend) | Native: each finished header task updates a per-chord counter; the last one triggers the body | | SQLAlchemy database backend and others without native support | The built-in `celery.chord_unlock` task polls the group and retries until it is ready | | `rpc` | Not supported: raises `NotImplementedError` ("The \"rpc\" result backend does not support chords!") | | none configured | Starting a chord raises `NotImplementedError` ("Starting chords requires a result backend to be configured.") | The polling path has costs of its own: - `celery.chord_unlock` is a real task: it occupies a worker slot each time it runs and re-queues itself every `result_chord_retry_interval` seconds (default **1.0**). - It is declared with `max_retries=None`, so if a header result never appears it keeps polling. - The body starts up to one interval after the last header task finishes. The error for a missing backend is raised when the chord is started, before any header task is sent, so it surfaces in the caller rather than as a stuck workflow. The same applies to `group(...) | task`, which Celery upgrades to a chord. ## What happens when one header task fails Take the photo pipeline: a header of `make_thumbnail` tasks for 160, 640 and 1280 px, and a `publish_photo` body. Suppose the 640 px task raises `OSError` on a corrupt image. 1. The other header tasks are **not cancelled**. The 160 and 1280 px thumbnails still run and store their results. 2. When the last header task reports, the join finds a failed dependency. 3. `publish_photo` is **never executed**. Its result is stored as `FAILURE` with a `ChordError`, whose message names a failed task id: `Dependency <task-id> raised OSError(...)`. 4. Any **errback linked to the body** is called with that error. 5. The caller waiting on the chord's `AsyncResult`, which is the body's result, gets the `ChordError` from `.get()`. The docs note that the `ChordError` reports one failing dependency, not all of them, so logs from the header tasks are still how you find every failure. ## Reacting to the failure The default reaction point is the **body**: - `publish_photo.s(photo_id).on_error(mark_upload_failed.s(photo_id))` attaches an errback that runs once when the chord fails. (`on_error()` returns the body signature, so it chains; `link_error()` returns the errback.) - Calling `link_error()` on the chord itself links the errback to the body only, because `task_allow_error_cb_on_chord_header` defaults to **False**. - Setting `task_allow_error_cb_on_chord_header = True` also links the errback to every header task, so it can run once per failing thumbnail **and** once for the body. Errbacks used that way must tolerate being called several times. - Celery emits a `CPendingDeprecationWarning` when a chord's `link_error()` runs with the setting at `False`, a sign that linking to the header may become the standard behaviour. Whether the failed thumbnail should be retried before the chord gives up is a question for the task's own retry settings, not for the chord. ## Senior-level judgment - **Pick a backend with native chord support** when chords are hot paths; polling with `chord_unlock` adds latency and queue traffic per chord. - **Keep header results small** (paths, ids), because every result is stored and then delivered to the body in one message. - **Do not set `ignore_result=True`** on header tasks, including through a global `task_ignore_result`. - **Decide what partial success means.** A chord is all-or-nothing: two of three thumbnails done still means no publish. If a partial result is acceptable, make header tasks return a status instead of raising, and let the body decide. - **Watch for chord bodies that never run.** Common causes: a header task that ignores its result on a polling backend, a result that expired before the join, or a header task that never finishes.
- What is celery.chord_unlock, and when does a Celery chord use it?It is a built-in task that result backends without native chord support schedule when a chord starts. It checks whether the header group is ready and, if not, retries itself every `result_chord_retry_interval` seconds (1.0 by default) with no retry limit. When the header is ready it joins the results and sends the body, or fails the body with a `ChordError`.
- How do you run an error handler for each failed thumbnail, not only once for the chord?Either link an errback to each header signature yourself, or set `task_allow_error_cb_on_chord_header = True`, which makes a chord's `link_error()` attach the errback to every header task as well as the body. The errback can then fire several times for one upload, so it must be idempotent.
saying these in an interview costs you the question
- When one header task fails, Celery cancels the other header tasks at once.
- The chord body still runs, with None in place of the failed thumbnail.
- The rpc result backend is enough for chords, since results travel over the broker.
- By default a chord's link_error errback runs once per failed header task.
- The broker tracks which header tasks finished, so no result backend is needed.