For a Celery video platform mixing transcodes, password resets and notifications, how would you decide how many queues to run and what each split costs?
answer
- split by what must not wait
- latency class versus task type
- every queue needs its own consumers
- guardrails for routes and topology
basics
~20 sSplit queues by latency and resource class, not per task: a fast lane for resets and notifications, a heavy lane for transcodes, maybe a bulk lane. Each extra queue buys isolation but costs a worker fleet, scaling rules and alerts.
solid answer
~50 sI start from what must not wait for what. Resets and notifications have a seconds-level expectation; transcodes take minutes and need big machines. That gives two lanes, each with its own `task_routes` entries and its own `celery -A video worker -Q ...` fleet, plus perhaps a bulk lane so a re-encode backfill cannot delay fresh uploads. I avoid a queue per task: every queue needs consumers that sit idle when it is empty, its own depth alert and its own scaling rule, and routing drift grows with the count. Priorities alone are weaker — they only reorder within a queue, depend on the broker and are blunted by prefetch. I'd guard it with declared `task_queues`, `task_create_missing_queues = False`, a test that every registered task has a route, and one dashboard of queue depth per lane.
go deeper
Know that separate queues with separate workers keep quick tasks from waiting behind slow ones.
Explain how task_routes and worker -Q implement lanes, and why priorities inside one queue are a weaker substitute.
Describe the operating cost of each lane: idle consumers, per-queue alerts and scaling, overflow direction, and routing drift for new tasks.
Choose the split from latency, resource and blast-radius needs, defend the costs you accept, and set the guardrails and signals that tell you when to split again.
## The question behind the question There is no single correct number of Celery queues. An interviewer asking this wants to see you reason from **what must not wait for what**, and then count the operating cost of each split. A **queue** here is a named line on the broker; a **lane** is a queue plus the worker fleet pinned to it with `-Q`. ## Axes you can split on - **Latency expectation** — a password reset is expected in seconds; a transcode in minutes. Mixing them means the slow work sets the waiting time for the fast work. - **Resource profile** — transcodes want many CPUs, large disks, maybe GPUs; notifications are I/O-bound and cheap. A dedicated lane lets each run on suitable machines with suitable pool settings. - **Blast radius** — a flood of one type (a bulk re-encode of the back catalogue) should not degrade another. - **Tenant fairness** — one customer's backlog starving others. That is a scheduling-design question more than a Celery one, but it is where per-tenant queues come from. ## Candidate topologies | topology | isolation | cost | fits when | |---|---|---|---| | one queue plus priorities | weak: only reorders, no pre-emption | lowest | small app, one worker type | | per latency class (fast / heavy / bulk) | strong between classes | a few fleets and alerts | most production setups | | per task type | strongest | many fleets, idle capacity, config drift | a few tasks with very different hardware | | per tenant | fairness between customers | dynamic queues and consumers to manage | multi-tenant platforms with noisy tenants | For the video platform, a reasonable answer is **three lanes**: `fast` (resets, notifications, other sub-second work), `transcode` (fresh uploads) and `transcode_bulk` (backfills and re-encodes). ## What every split costs 1. **Consumers.** A queue with no consumer strands its messages silently, so each queue needs a fleet that is running even when idle. 2. **Capacity you cannot share.** A fast-lane worker waiting for work cannot help an overloaded transcode lane unless you add overflow consumption, which weakens the isolation. 3. **Scaling and alerting.** Each lane needs its own queue-depth metric, its own scale-up rule and its own alert threshold. 4. **Routing drift.** Each new task must be routed on purpose. A new slow task that falls through to the fast lane re-creates the original problem. 5. **Deploy surface.** More worker commands, more rollout units. ## Overflow, and priorities versus queues Letting heavy-lane workers also consume the fast queue (`-Q transcode,fast`) uses idle capacity. The reverse is a mistake: a fast worker that picks up a transcode stops being fast. Even the safe direction has a subtlety: a heavy worker whose processes are all transcoding may already have reserved fast messages through prefetch, and those wait until a process frees up. Keep prefetch low on mixed consumers, or keep lanes strict. Priorities look cheaper — one queue, one fleet — but they only reorder messages that have not been delivered yet, they never interrupt a running transcode, and their meaning depends on the broker: RabbitMQ needs `x-max-priority` and serves high numbers first, kombu's Redis emulation serves 0 first, and the SQS transport has none. Use queues for isolation between latency classes and priority, if at all, for ordering inside a lane. ## Guardrails, and presenting the decision - Declare the topology in `task_queues` and set `task_create_missing_queues = False`, so a mistyped queue name fails at send time. - Keep all routing in `task_routes`; decorator `queue=` options are invisible to `send_task` callers. - Add a test that every registered task name resolves to an intended lane. - Start every worker with an explicit `-Q`. - Watch queue depth and oldest-message age per lane, not just in aggregate. When presenting the decision, state the lanes, the rule for assigning a new task to one, the fleet and scaling signal for each, and what you would re-examine as the platform grows — for example, splitting `fast` if a marketing notification blast starts to delay resets, or giving `transcode_bulk` spot capacity that can be withdrawn without hurting fresh uploads. The judgement is in the costs you accept, not the count.
- When would you choose one queue with priorities over separate queues?When there is one worker type, modest volume and no long-running tasks, so pre-emption does not matter, and the broker supports priority natively. Then one fleet is simpler to run. Once task durations differ by orders of magnitude, priorities cannot stop fast work waiting behind running slow work, and separate lanes win.
- How do you stop a newly added slow task from quietly landing in the fast lane?Make routing a reviewed contract: keep it in `task_routes`, add a test that walks the registered tasks and asserts each maps to an intended queue, and consider making the default queue a dedicated catch-all lane rather than the fast lane.
saying these in an interview costs you the question
- Each task type should always get its own queue
- One queue with priorities isolates slow tasks as well as separate queues do
- An idle queue costs nothing, so extra queues are free
- Fast-lane workers should also consume transcodes to use spare capacity