skip to content

A Celery invoice task declared with rate_limit='10/m' starts about 40 times a minute across four workers; why, and what does rate_limit actually enforce?

level: seniorimportance: should knowfreq 30%

answer

  1. where the counter lives
  2. a token bucket per task type
  3. per worker instance, not global
  4. excess tasks wait, not fail

basics

~20 s

Celery's rate_limit is a token bucket kept separately by each worker instance for each task type, not a cluster-wide limit. Four workers at '10/m' allow about 40 starts a minute; tasks over the limit wait inside the worker rather than failing.

solid answer

~40 s

`rate_limit='10/m'` means each worker instance starts at most one `render_invoice_pdf` every six seconds: the bucket holds one token, so there is no burst. The bucket lives in the worker's main process, shared by its pool children, but every worker instance keeps its own, and the Celery docs say plainly it is a per-worker limit, not a global one. Four workers therefore give about 40 starts a minute. Tasks over the limit are not rejected: the worker has already taken them off the queue and holds them until a token frees. It limits start rate, not concurrency, and `worker_disable_rate_limits` switches it off. For a real global cap on the PDF service, I route the task to a dedicated queue consumed by a single worker, or enforce a shared limit in the task body.

code

python · 8 lines
python
from celery import Celery

app = Celery('shop', broker='redis://localhost:6379/0')
app.conf.task_routes = {'billing.render_invoice_pdf': {'queue': 'invoices'}}

@app.task(name='billing.render_invoice_pdf', rate_limit='10/m')
def render_invoice_pdf(order_id):
    ...   # calls the external PDF service

go deeper

for a junior

Recall the rate_limit formats, that it limits how often a task starts, and that it is counted per worker.

for a middle

Explain the token bucket in the worker's main process, the lack of burst, and why tasks over the limit wait instead of failing.

for a senior

Diagnose the multiplied rate across worker instances, the prefetch slots held by waiting tasks, and choose between a dedicated single-consumer queue and a shared limiter.

for a principal

Weigh throughput against a provider's hard cap: whether one throttled consumer is an acceptable bottleneck, or the limit belongs in shared infrastructure the whole fleet respects.

## What rate_limit is A Celery task can carry a **rate limit**, set with `@app.task(rate_limit='10/m')`, as the `task_default_rate_limit` setting for all tasks, or changed at runtime with a remote-control command. The value is a number of task starts per unit: - a bare number, such as `2`, means tasks per second; - a string with `/s`, `/m` or `/h` sets the unit, such as `'10/m'` or `'100/h'`; - `None`, the default, means no limit. The docs describe the effect as spacing starts evenly over the period: `'100/m'` enforces at least 600 ms between two starts **on the same worker instance**. ## How the worker enforces it Inside a worker, the consumer, which is the main process that talks to the broker, keeps one **token bucket** per task type (`bucket_for_task` in `celery/worker/consumer/consumer.py`). The bucket refills at the configured rate and has a capacity of one token, so there is no burst allowance: `'10/m'` means one start every six seconds, not ten at once and then a pause. When a message for a rate-limited task arrives: 1. The worker has already received it from the broker, as with any prefetched message. 2. If a token is available, the task goes to the pool and starts. 3. If not, the request waits in the bucket's own list, and a timer retries when the next token is due. Nothing fails and nothing goes back to the broker. The waiting task simply starts later. ## Why four workers give four times the rate The bucket is state in one worker's memory. It is not in the broker and not shared between machines. | Deployment | Effective limit for '10/m' | |---|---| | one worker, `--concurrency 8` | about 10 starts a minute: one bucket in the main process | | four worker instances | about 40 a minute: four independent buckets | | autoscaled to twelve instances | about 120 a minute | So the number that matters is the count of worker instances consuming the task. The Celery docs spell this out: it is a *per worker instance* rate limit, and a global limit needs the task restricted to a given queue. ## What rate_limit does not do - It does **not** cap concurrency. Ten slow PDF renders can still overlap if each takes more than six seconds. - It does **not** protect the queue. Waiting tasks hold prefetch slots in the worker, so a heavily limited task can hold messages that another worker could have run. - It does **not** survive a restart as state; a fresh worker starts with a fresh bucket. - It can be switched off wholesale with `worker_disable_rate_limits = True`. ## Getting a real global limit For a third-party PDF service with a hard account-wide limit: - **Dedicated queue, single consumer.** Route `render_invoice_pdf` to its own queue and run exactly one worker instance on it with the rate limit. The limit is then global because only one bucket exists. The cost is a single point of throughput and failure for that task. - **Shared limiter in the task body.** Keep a counter in a shared store and, when the budget is spent, have the task call `self.retry(countdown=...)` rather than calling the service. This scales across workers at the price of extra code and retries. - **Let the service push back.** Treat the provider's throttling response as a retryable error with backoff. The first option fits a small shop pipeline best; the second fits when invoice volume needs several workers anyway. ## Changing a limit at runtime A limit can be adjusted on running workers without a deploy through the remote-control command `rate_limit`, for example `celery -A shop control rate_limit billing.render_invoice_pdf 5/m`. By default it is broadcast to every worker, and each applies the new value to its own bucket; the change is still per worker and is lost when a worker restarts, because the decorator's value is read again at start-up. It is an incident lever, not configuration. ## Quick diagnosis checklist - Count worker instances consuming the task; multiply the limit by that number. - Check whether `worker_disable_rate_limits` is set on any of them. - Check whether the task's queue is shared with other workers you forgot about, such as a default worker that consumes every queue.

  • Why can a heavily rate-limited Celery task slow other tasks on the same worker?
    The worker takes rate-limited messages off the broker like any other and holds them until a token frees. Those messages occupy prefetch slots, so the worker fetches fewer new messages while they wait. Putting the limited task on its own queue and worker keeps that waiting from affecting unrelated tasks.
  • Does raising --concurrency on one Celery worker raise the effective rate_limit?
    No. The bucket lives in the worker's main process and is shared by every pool process, so one worker instance at `'10/m'` still starts about ten a minute whatever its concurrency. Only adding worker instances multiplies the effective rate.

Each worker is a separate ticket booth with its own turnstile set to one person every six seconds. Opening four booths does not slow each turnstile; it quadruples how many people get through the gate.

saying these in an interview costs you the question

  • Celery's rate_limit is a global limit coordinated through the broker.
  • Each pool process of a worker gets its own rate-limit bucket.
  • Tasks over the rate limit fail and have to be retried.
  • rate_limit='10/m' lets ten tasks start at once, then waits a minute.
  • rate_limit also caps how many of the tasks run at the same time.