skip to content

What does bind=True change about a Celery task, and what can a bound task read from self.request while it runs?

level: middleimportance: should knowfreq 40%

answer

  1. the first argument changes
  2. a per-execution context object
  3. id, retries, hostname, delivery_info
  4. one task object per worker process

basics

~20 s

With bind=True, Celery passes the task object itself as the first argument, self. Through self.request a running task reads its execution context: task id, retry count, args and kwargs, worker hostname, delivery info, eta and expires.

solid answer

~40 s

`@app.task(bind=True)` makes the task behave like a method: the first parameter, `self`, is the task object, so the body can call `self.retry()`, `self.update_state()` or custom methods from a `base=` class. `self.request` is a context object for the execution in progress: `id`, `retries`, `args`, `kwargs`, `hostname`, `delivery_info` (the exchange and routing key it arrived on), `eta`, `expires`, `is_eager`, `parent_id` and `root_id`. The id survives `self.retry()`, so logs for every attempt at rendering one invoice share it, while a fresh `delay()` gets a new id. The trap is that `self` is not created per call: each worker process registers one task instance and reuses it, so state stored on `self` leaks into the next invoice.

code

python · 16 lines
python
import logging

from celery import Celery

app = Celery('shop', broker='redis://localhost:6379/0')
log = logging.getLogger(__name__)

@app.task(bind=True)
def render_invoice_pdf(self, order_id):
    req = self.request
    log.info('invoice order=%s task=%s attempt=%s host=%s',
             order_id, req.id, req.retries, req.hostname)
    ...

# The caller never passes self:
render_invoice_pdf.delay(42)

go deeper

for a junior

Recall that bind=True makes the task object the first argument and that self.request exposes the task id and retry count.

for a middle

Explain which request fields exist and where they come from, that the id survives self.retry(), and that one task instance serves a whole worker process.

for a senior

Show the production consequences: correlating logs by request.id, not deduplicating by it, and the state-leak bug from storing per-call data on self.

for a principal

Set team conventions for task bodies: what goes into logs from request, where idempotency keys come from, and when a custom base class earns its keep.

## What binding does A Celery task declared with `@app.task` is a plain function wrapped in a **task class**; the worker calls the function with only the arguments from the message. Declaring it with `@app.task(bind=True)` makes the function receive the **task instance** as its first argument, conventionally named `self`, just as a Python method receives its object. Nothing else about publishing or execution changes; the caller still writes `render_invoice_pdf.delay(order_id)` without passing `self`. What binding buys is access to the task's own API from inside the body: - `self.request`, the context of the execution in progress - `self.retry(...)`, which re-publishes the same task (its options belong to the retries topic) - `self.update_state(...)`, which records a custom state such as progress - methods and attributes of a custom base class passed with `base=` ## The request context `self.request` is a `Context` object (defined in `celery/app/task.py`) that the worker fills from the message headers before running the body. The fields a production task actually uses: | Attribute | What it holds | |---|---| | `id` | the task id, the same one the caller's `AsyncResult` carries | | `retries` | how many times this execution has been retried, starting at 0 | | `args`, `kwargs` | the arguments from the message | | `hostname` | the node name of the worker running it | | `delivery_info` | exchange and routing key it arrived with; `self.retry()` reuses them | | `eta`, `expires` | the original schedule and expiry, if any | | `is_eager` | `True` when run locally instead of by a worker | | `called_directly` | `True` when the body was invoked as a plain function | | `parent_id`, `root_id` | the task that called this one and the first task of the workflow | | `headers` | custom message headers, if the caller sent any | In the invoice task, `self.request.id` goes into every log line, `self.request.retries` tells the body whether this is a first attempt, and `self.request.hostname` says which machine produced a bad PDF. ## The id across retries and re-sends Two facts about `request.id` are easy to mix up: 1. `self.retry()` re-publishes the task **with the same id**, so all attempts at one execution share it. 2. A new `delay()` for the same order creates a **new** id. If the checkout view accidentally enqueues the invoice twice, the two runs have different ids. So `request.id` is a good correlation key for logs, but a poor deduplication key for business effects. For "render each order's invoice once", key the check on the order, for example by skipping when the invoice file for that order already exists. ## The task instance is shared The Celery docs are explicit: a task is **not instantiated for every request**. It is registered in the task registry as a global instance, so `__init__` runs once per worker process and the same object serves every message that process handles. For a bound task that means: - attributes you set on `self` persist into the next execution in that process; - they are not shared between processes of the pool, or between worker machines; - a per-process cache on the instance (a PDF template loaded once, a client created lazily) is a legitimate use, per-invoice state is not. Storing `self.current_order = order_id` and reading it later in the body works in a quick test and fails in production: on a thread-based pool a concurrent execution in the same process can overwrite it mid-run, and on any pool an execution that returns early reads or leaves the previous order's value. ## Unbound tasks and the request `request` is a property of the task class, so even an unbound task can read the current context through its module-level name, as in `render_invoice_pdf.request.id`. Binding is simply the idiomatic, explicit way, and it is what makes `self.retry()` and `self.update_state()` read naturally. The `bind=True` flag must be matched by the extra first parameter; forgetting `self` shifts every argument by one and the signature check then rejects the call. ## Common mistakes - Treating `self` as a fresh object per execution and keeping per-call data on it. - Using `request.id` as the idempotency key for a business action that can be enqueued twice. - Assuming `request.retries` counts every delivery; it counts `retry()` calls, not broker redeliveries. - Expecting `request.eta` to be set on every task; it is only there when the caller gave `countdown` or `eta`.

  • Does self.request.retries count broker redeliveries after a worker crash?
    No. `retries` is carried in the message and incremented when the task calls `self.retry()`, which publishes a new message with a higher count. A message the broker redelivers after a worker dies is the same message, so the body sees the same `retries` value as before and cannot tell it was redelivered from that field.
  • When is keeping state on self in a bound Celery task legitimate?
    When the state is a per-process cache that is valid for every execution: a template loaded once, a lazily created HTTP client, a compiled regex. The instance lives as long as the worker process, so anything tied to one order, one request or one retry must stay in local variables or the arguments.

saying these in an interview costs you the question

  • Celery creates a new task object for every message, so attributes on self are private to one call.
  • The caller must pass self when calling a bound task with delay().
  • self.request.id changes on every retry of the same execution.
  • self.request.retries also counts redeliveries after a worker crash.
  • Only bound tasks can ever see the current request context.