skip to content

In Celery, how do the Redis, database and rpc:// result backends differ in who can read a task's result, and for how long?

level: middleimportance: should knowfreq 38%

answer

  1. stored versus sent as a message
  2. keys with a TTL
  3. rows need a cleanup task
  4. read once, by the sender only

basics

~20 s

Celery's Redis and database backends store results any process can read by task id; Redis expires them after result_expires (one day), a database only when celery.backend_cleanup runs from beat. rpc:// sends results as messages, readable once and only by the sending client.

solid answer

~40 s

The **Redis** backend writes each result to a `celery-task-meta-<task_id>` key with a TTL of `result_expires` (one day by default), so any process that knows the task id can read it until it expires, and `get()` waits on pub/sub rather than polling. A **database** backend (`db+` SQLAlchemy URLs, or `django-db` from django-celery-results) keeps a row per task that anyone can query, but rows do not expire on their own: `celery.backend_cleanup`, which beat schedules daily at 4 am, deletes those older than `result_expires`. **`rpc://`** stores nothing; it sends the result as a message back to the client that sent the task, so only that client can read it, only once, and it is transient unless `result_persistent` is on. For a payouts status page served by another process, `rpc://` is ruled out.

code

python · 11 lines
python
from celery import Celery

app = Celery("payouts", broker="amqp://rabbit//")

# stored, readable by any process with the task id, expires by TTL
app.conf.result_backend = "redis://redis:6379/1"
app.conf.result_expires = 6 * 3600  # seconds; default is one day

# alternative: results as messages, readable once by the sending client
# app.conf.result_backend = "rpc://"
# app.conf.result_persistent = True  # survive a broker restart

go deeper

for a junior

Recall that Redis and database backends store results by task id, while rpc:// sends them back as messages to the caller.

for a middle

Explain expiry per backend: a Redis key TTL from result_expires, database rows removed by celery.backend_cleanup when beat runs, and nothing stored at all for rpc://.

for a senior

Match the backend to who reads the result: a status endpoint in another process needs a stored backend, and a database backend needs beat running or its table grows forever.

for a principal

Decide whether results belong in Celery's backend at all, or whether the payout's own record should carry its status and the backend stay short-lived.

## What a result backend stores A **result backend** is where a Celery worker records a task's outcome: its **state** (such as `SUCCESS` or `FAILURE`), its return value or exception, and optionally metadata. Callers read it through an `AsyncResult`, which is little more than a task id plus a handle to the backend. The backend is chosen with `result_backend`, and the three most common choices behave very differently on two questions: **who can read a result** and **how long it lives**. ## Redis: shared keys with a TTL With `result_backend = "redis://..."` the worker writes the result to a key named `celery-task-meta-<task_id>`. - Any process configured with the same backend can read it by building `AsyncResult(task_id)`: the web request that sent the task, a different web process, a script. - The key is written with a TTL equal to `result_expires`, **one day** by default, so Redis removes it without any extra job. - `get()` is notified over Redis pub/sub instead of polling. - Durability is Redis's: results written since the last persistence point can vanish on a restart, which is usually acceptable for results but worth knowing. ## Database backends: queryable rows that need cleaning Celery's own database backend takes `db+` followed by a SQLAlchemy URL; in a Django project, **django-celery-results** adds the `django-db` backend, which stores results in its `TaskResult` model. - Rows are readable by anyone with database access, including the Django admin, and they survive restarts like any other table. - With `result_extended` on, rows also carry the task name, arguments and worker; with `task_track_started` on, the `STARTED` state and its time are recorded. - Rows **do not expire by themselves**. When `result_expires` is set, beat installs a built-in `celery.backend_cleanup` task scheduled daily at 4:00, and that task deletes expired rows. No beat process means the table grows without bound. - Reading state means querying the database; Celery's task guide warns that polling a database for state changes is expensive and suggests longer `get()` intervals. ## rpc://: results as messages The **RPC backend** does not store results at all. The worker sends the state and result as messages back over the broker to a reply queue owned by the client that published the task. 1. Only that client can receive them, so another process asking about the same task id learns nothing. 2. A result can be consumed only once. 3. Messages are transient by default; with `result_persistent = True` they survive a broker restart. 4. The client receives state changes in real time without polling. 5. A client that fires many tasks and reads an old one late can hit `BacklogLimitExceeded`. ## Side by side | Property | Redis | Database (`db+...`, `django-db`) | `rpc://` | |---|---|---|---| | Who can read | any process with the task id | any process, plus ad hoc queries | only the client that sent the task | | How many times | until expiry | until cleanup | once | | Expiry | key TTL from `result_expires` | `celery.backend_cleanup` via beat | none needed; nothing is stored | | Waiting in `get()` | pub/sub, no polling | polling | messages, no polling | ## A common misconfiguration The worker and the web process each read their own Celery configuration. If the worker has `result_backend` set but the web process that builds `AsyncResult(task_id)` does not, the web process uses the disabled backend and cannot read anything; if the two point at different Redis databases, the web process looks in the wrong place and sees every task as `PENDING`, which Celery reports for any id it has no record of. Both processes must share the same `result_backend` value, and an unknown id is indistinguishable from one that is still queued. ## Choosing for a payouts service - A status endpoint running in a different web process than the code that sent `send_payout` needs a **stored** result, so Redis or a database; `rpc://` cannot serve it. - If support staff audit task outcomes, django-celery-results rows in the admin are convenient, provided beat runs the cleanup. - If only the caller that sent the task ever reads the result, `rpc://` gives real-time replies with nothing to clean up. - `result_expires = None` (or 0) disables expiry entirely, which turns the backend into an unbounded log. The common mistake is choosing a backend by speed alone and discovering later that a different process needed to read the result, or that nobody was deleting the rows.

  • Why can a django-db results table keep growing even though result_expires is one day?
    The database backend cannot expire rows on its own, so it relies on the built-in `celery.backend_cleanup` task that beat schedules daily at 4 am to delete rows older than `result_expires`. If no beat process runs, that task never fires and the `TaskResult` table grows without bound. Running beat, or a scheduled job that performs the same cleanup, fixes it.
  • What does result_extended add to stored results?
    With `result_extended = True` the backend stores extra task metadata alongside state and result, such as the task name, arguments and the worker that ran it. It is off by default. django-celery-results fills the matching `TaskResult` columns from it, which makes the admin useful for auditing, at the cost of storing arguments that may be sensitive.

saying these in an interview costs you the question

  • Any web process can read an rpc:// result by building AsyncResult(task_id).
  • Database result rows expire on their own after result_expires.
  • Redis results live forever unless someone calls forget().
  • rpc:// results survive a broker restart by default.
  • result_expires deletes results the moment a caller reads them.