skip to content

In Celery, which worker pool suits CPU-heavy report rendering and which suits thousands of outbound webhook calls, and why?

level: middleimportance: must knowfreq 50%

answer

  1. processes versus greenlets
  2. who waits on what
  3. one interpreter lock per process
  4. cooperative yield on sockets

basics

~20 s

Run CPU-heavy rendering on Celery's default prefork pool: separate processes each run their own interpreter and can keep a core busy. Run I/O-bound webhook calls on gevent or eventlet, where hundreds of greenlets wait on sockets cheaply inside one process.

solid answer

~40 s

`prefork` forks child processes, each with its own interpreter and its own GIL, so pandas rendering really uses every core; keep `-c` near the core count. It is also the only pool with child recycling (`worker_max_tasks_per_child`, `worker_max_memory_per_child`). For webhooks, `-P gevent -c 200` or more runs greenlets that yield whenever a patched socket blocks, so hundreds of HTTP calls wait concurrently in one process for little memory. The catch is cooperation: one CPU-bound task on a gevent worker never yields and stalls every other greenlet there. `threads` is a middle option for moderate I/O without monkey patching, and `solo` runs one task at a time for debugging. Because the pool is per worker, the two workloads belong on separate workers.

code

bash · 5 lines
bash
# CPU-heavy pandas reports: processes, one per core
celery -A proj worker -P prefork -c 8 -Q reports --max-memory-per-child=2000000

# I/O-bound webhooks: greenlets; -P on the command line so patching happens first
celery -A proj worker -P gevent -c 300 -Q webhooks

go deeper

for a junior

Recall the five pool names, that prefork is the default, and that prefork suits CPU-bound work while gevent and eventlet suit I/O-bound work.

for a middle

Explain why separate processes use every core while greenlets yield on patched sockets, and why -P must be on the command line for green pools.

for a senior

Show you would never put CPU work on a green worker, would pick threads when patching breaks a client library, and would split workers by workload.

for a principal

Weigh the memory cost of processes against the cooperative-scheduling risk of greenlets when designing a fleet that serves both workloads.

## The pools Celery ships A **Celery worker** runs tasks through an **execution pool**, selected with `-P` / `--pool` on the command line (default `prefork`). Each pool has a different unit of concurrency and suits a different kind of task: | Pool | Unit | CPU-bound tasks | I/O-bound tasks | Notes | |---|---|---|---|---| | `prefork` | child process | good — one core per child | works, but each slot costs a whole process | default; supports child recycling | | `gevent` | greenlet | poor — blocks the whole worker | excellent — hundreds of slots | needs monkey patching via `-P` | | `eventlet` | green thread | poor | excellent | same model as gevent | | `threads` | OS thread | limited by the GIL | good at moderate concurrency | no monkey patching | | `solo` | main process | one task at a time | one task at a time | debugging, special platforms | A **CPU-bound** task spends its time computing (building a DataFrame, writing a large spreadsheet). An **I/O-bound** task spends its time waiting on something else (an HTTP response from a webhook receiver). ## Why reports belong on prefork The `prefork` pool uses billiard, Celery's fork of `multiprocessing`, to start `-c` child processes. Each child is a full Python interpreter with its own **Global Interpreter Lock (GIL)** — the lock that lets only one thread run Python bytecode per process. Separate processes therefore compute in parallel on separate cores, which is exactly what report rendering needs. - Set `-c` near the core count, lower if each report's peak memory times `-c` exceeds the RAM you have. - Use `worker_max_tasks_per_child` or `worker_max_memory_per_child` to recycle children whose memory grew; both work only on `prefork`. - Expect a bigger memory footprint: every child carries its own copy of your imported modules once it starts writing to them. ## Why webhooks belong on a green pool `gevent` and `eventlet` run **greenlets**: lightweight coroutines scheduled cooperatively inside one OS thread. When a greenlet makes a network call on a monkey-patched socket, it yields, and another greenlet runs. A webhook call that waits 800 ms for a response costs almost nothing while it waits, so `-P gevent -c 200` (or more) keeps hundreds of calls in flight from one process. Two conditions make this work: 1. **The standard library must be patched before anything imports it.** The `celery` command applies `gevent.monkey.patch_all()` (or eventlet's patch) as it parses `-P gevent` from the command line. Selecting the pool through the `worker_pool` setting instead makes the worker warn that you must use `-P` so patches are applied early enough. 2. **Every task must yield.** Greenlets are not preempted. A CPU-heavy loop, or a C library doing blocking I/O that the patches cannot reach, holds the only thread until it returns, and every other greenlet on that worker waits. ## The threads pool as a middle option `threads` runs tasks in a `concurrent.futures.ThreadPoolExecutor` sized by `-c`. Blocking I/O releases the GIL, so moderate I/O concurrency works without any monkey patching — useful when a client library misbehaves under gevent. Pure-Python CPU work gains little, because the GIL serialises bytecode; libraries that release the GIL in native code can overlap partially. It does not support child recycling. ## Putting it together Because the pool is a property of a **worker**, not of a task, a fleet with both workloads runs two kinds of worker, each consuming its own queue: - report workers: `-P prefork -c <cores>`; - webhook workers: `-P gevent -c 200`. Mixing both on one gevent worker lets a single report freeze every webhook; mixing both on one prefork worker spends a whole process on each waiting HTTP call. ## A quick decision checklist 1. Does the task spend most of its time computing? Use `prefork`, `-c` near the core count. 2. Does it spend most of its time waiting on the network, and do its client libraries work under monkey patching? Use `gevent` (or `eventlet`) with a high `-c`. 3. Is it I/O-bound but a library breaks under patching? Use `threads` with a moderate `-c`. 4. Are you debugging, or on a platform where forking misbehaves? Use `solo`. Whatever you pick, confirm the choice in the worker's startup banner, which prints the concurrency next to the pool name.

  • Why must a Celery worker choose gevent with -P rather than the worker_pool setting?
    The `celery` command applies gevent's monkey patches as it parses `-P gevent`, before your app and its libraries import `socket` or `ssl`. Set through `worker_pool`, the choice is read too late; the worker emits a warning saying you must use `-P`, and modules imported earlier keep unpatched, blocking sockets.
  • What happens when a pandas report task lands on a Celery gevent worker?
    Greenlets are cooperative, and a computation never yields. For as long as the report runs, every other greenlet on that worker — the webhook calls included — waits, even though `-c` says 300. Route reports to a prefork worker instead of relying on the green pool.
  • When is the Celery threads pool a reasonable choice over gevent?
    When tasks are I/O-bound at moderate concurrency and a client library does not cooperate with monkey patching, or you want to avoid patching altogether. Threads preempt each other, so one slow task does not freeze the rest, but each slot is a real OS thread and child recycling is unavailable.

saying these in an interview costs you the question

  • gevent runs greenlets in parallel across all CPU cores.
  • Setting worker_pool = 'gevent' in config is the same as passing -P gevent.
  • Prefork with -c 500 is the efficient way to run thousands of webhook calls.
  • worker_max_memory_per_child also recycles memory on gevent and threads workers.
  • A CPU-heavy task on a gevent worker only slows its own greenlet.