skip to content

What does the --concurrency option of a Celery worker control, and what value does it use when you leave it unset?

level: juniorimportance: should knowfreq 45%

answer

  1. one worker, many slots
  2. meaning depends on -P
  3. child processes under prefork
  4. counted from the machine

basics

~20 s

Celery's --concurrency (-c, setting worker_concurrency) sets how many tasks one worker runs at once: child processes under prefork, threads or greenlets under other pools. Unset, it defaults to the number of CPUs the worker detects.

solid answer

~40 s

`celery -A proj worker -c 8` starts one worker that runs up to eight tasks at the same time. What a slot is depends on the pool: a forked child process under the default `prefork` pool, an OS thread under `threads`, a greenlet under `gevent` or `eventlet`; the `solo` pool always runs exactly one task, whatever `-c` says. If neither `-c` nor `worker_concurrency` is set, Celery uses the CPU count it detects (falling back to 2 if it cannot count). For CPU-heavy report rendering keep it near the core count; for I/O-bound webhook calls on a green pool it can run to hundreds. It also sizes the prefetch window: with the default `worker_prefetch_multiplier` of 4, `-c 8` lets the worker hold up to 32 unacknowledged messages.

code

bash · 5 lines
bash
# report rendering: one prefork child per core on an 8-core host
celery -A proj worker -P prefork -c 8 -Q reports

# webhook calls: 200 greenlets in one process
celery -A proj worker -P gevent -c 200 -Q webhooks

go deeper

for a junior

Recall that -c sets task slots inside one worker, that the default pool is prefork, and that an unset value becomes the detected CPU count.

for a middle

Explain what a slot is in each pool, why solo ignores -c, and how concurrency multiplies with worker_prefetch_multiplier into the number of reserved messages.

for a senior

Show that you size concurrency from the task's real bottleneck — cores, RAM per task or downstream connections — and verify what the worker detected inside containers.

for a principal

Frame concurrency as one capacity knob among worker count, pool type and host size, and justify when several small workers beat one large one.

## One worker, several execution slots A **Celery worker** is the long-running process you start with `celery -A proj worker`. It connects to the **broker** (RabbitMQ, Redis or Amazon SQS), receives task messages and runs the task functions. A single worker does not have to run one task at a time: it owns an **execution pool**, and `--concurrency` (short form `-c`, config setting `worker_concurrency`) is the number of **slots** in that pool — the maximum number of tasks that worker executes simultaneously. So `-c 8` does **not** start eight workers. It starts one worker node, registered under one hostname, that can run eight tasks in parallel or concurrently, depending on the pool. ## What a slot is, pool by pool The pool is chosen with `-P` / `--pool`; the default is `prefork`. The same number means different things in each: | Pool (`-P`) | One slot is | What `-c` means | |---|---|---| | `prefork` (default) | a forked child process | number of child processes | | `threads` | an OS thread in a thread pool | number of threads | | `gevent` / `eventlet` | a greenlet (cooperative coroutine) | number of greenlets | | `solo` | the worker's own main process | ignored — always 1 | The `solo` pool sets its limit to 1 internally and runs each task inline, so `-P solo -c 16` still runs one task at a time. ## The default when you set nothing If neither `-c` nor `worker_concurrency` is given, the worker counts the CPUs on the machine and uses that number; if the count is not available it falls back to **2**. The default does not know anything about your tasks: - it is roughly right for **CPU-bound** work on `prefork`, where each child can keep one core busy; - it is far too low for **I/O-bound** work on a green pool, where each slot mostly waits on a socket; - it can be too high for **memory-heavy** tasks, where the limit is RAM, not cores. Inside a container, check what the worker actually detected — the startup banner prints the concurrency — rather than assuming it matches the container's CPU quota. ## Picking a number for a real workload Take one fleet that renders large spreadsheet reports with pandas and also fires thousands of outbound webhooks: 1. **Report workers** (`prefork`): start near the core count. If each report peaks at several gigabytes, divide the RAM you can spare by the peak per task and take the smaller of the two numbers. 2. **Webhook workers** (`gevent`): concurrency is bounded by how many simultaneous HTTP calls you and the receivers can tolerate, not by cores — `-P gevent -c 200` is ordinary. 3. **Measure** throughput and memory, then adjust. The workers guide notes that past a certain point more pool processes hurt, and that several worker instances can outperform one big one. ## How concurrency feeds prefetch Concurrency also sizes how many messages the worker reserves from the broker. The prefetch count is `worker_concurrency × worker_prefetch_multiplier`, and the multiplier defaults to **4**. A worker with `-c 8` therefore asks for up to **32** unacknowledged messages. Raising `-c` raises that reservation too, which matters when tasks are long. ## Where the value comes from The worker resolves concurrency in a fixed order: 1. `-c` / `--concurrency` on the command line, if given; 2. otherwise the `worker_concurrency` setting on `app.conf`; 3. otherwise the detected CPU count, or 2 if it cannot be counted. With `--autoscale=max,min` the pool starts at the minimum and may grow to the maximum, and the startup banner shows both numbers instead of one. Setting the value in configuration keeps it versioned with the code; passing `-c` per deployment lets the same code run on different host sizes. ## Common misreadings - **"-c starts N workers."** It sets slots inside one worker. - **"The default is 1."** The default is the detected CPU count. - **"More is always faster."** CPU-bound tasks gain nothing beyond the core count, and memory-heavy ones can push the machine into swapping or the kernel's OOM killer. - **"-c means threads."** Only under `threads`; under the default `prefork` it means processes.

  • Does starting a Celery worker with -P solo -c 16 run sixteen tasks at once?
    No. The `solo` pool sets its own limit to 1 and runs each task inline in the worker's main process, so `-c` has no effect. It is useful for debugging or on platforms where forking is a problem, not for throughput.
  • Is one Celery worker with -c 16 the same as two workers with -c 8 each?
    The slot count matches, but two workers are two consumers with two prefetch windows and two parent processes. They isolate failures (one crashing parent does not stop all 16 slots) at the cost of extra parent memory. Celery's workers guide notes several instances can perform better than one large one; measure it.

saying these in an interview costs you the question

  • Passing -c 8 starts eight separate Celery worker nodes.
  • A Celery worker's default concurrency is 1 task at a time.
  • Concurrency always means threads, even under the default prefork pool.
  • Raising -c on prefork keeps speeding up CPU-bound tasks past the core count.
  • The solo pool runs as many tasks as -c allows.