skip to content

Background Jobs & Scheduling

Getting work off the request path: task queues, background workers and scheduled jobs in .NET and Python. Asked because retries and at-least-once delivery are where it silently fails.

on this pageshow

explore

questions

page 2 of 2

In Celery, what happens to hourly and nightly beat entries when beat, or every worker, is down for three hours, and which options change that?

level: seniorimportance: should knowfreq 28%

basics

~20 s

If beat was down, each overdue entry is sent once on restart, not once per missed slot. If workers were down, beat kept sending and messages queued. An expires option drops stale copies; beat_cron_starting_deadline skips crontab runs that are too late.

open as a page

Why does a Celery chord need a result backend, and what happens to the chord body when one thumbnail task in its header fails?

level: seniorimportance: should knowfreq 28%

basics

~20 s

A Celery chord needs a result backend because something must record each header task's completion and result before the body can start. If one header task fails, the others still run, but the body never runs and is marked FAILURE with a ChordError.

open as a page

In Celery, you revoke a hung PDF-export task by id, yet it keeps running, and a revoked queued export later runs anyway; why, and what are terminate's risks?

level: seniorimportance: should knowfreq 25%

basics

~20 s

revoke only broadcasts the id to workers, which skip that task when they reach it; a running task continues unless terminate=True signals its pool process. Revoked ids live in worker memory, so a full restart forgets them unless workers use --statedb.

open as a page

During a rolling deploy, a Celery 5.6 worker gets SIGTERM mid-way through a long PDF export and is killed after the grace period; what do warm, soft and cold shutdown do, and how do you keep the task?

level: seniorimportance: should knowfreq 30%

basics

~20 s

SIGTERM starts a warm shutdown that waits, unbounded, for running tasks, so the later SIGKILL loses them. Cold shutdown (SIGQUIT) cancels them; soft shutdown (5.5+) first waits worker_soft_shutdown_timeout. Use REMAP_SIGTERM=SIGQUIT, a timeout under the grace period, and acks_late.

open as a page

In Celery, when should a task raise `Reject` or `Ignore` instead of calling `self.retry()`, and what happens to the message and its stored state?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Use self.retry() for transient failures: a new message runs the task again. Raise Reject for a message that should leave the queue, dead-lettered or requeued, which needs acks_late. Raise Ignore when the work is moot: it acks and records nothing.

open as a page

A Celery app routes transcodes to a new 'transcode' queue, and they pile up unprocessed while other tasks run; why, and what does declaring task_queues change?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Routing creates the queue, but no worker consumes it: a worker without -Q drains only task_queues, by default just celery. Declaring task_queues and disabling task_create_missing_queues makes bare workers consume every declared queue and turns typos into errors.

open as a page

In Celery, why might apply_async(priority=9) fail to move password-reset emails ahead of queued transcodes, and how do RabbitMQ and Redis priorities differ?

level: seniorimportance: should knowfreq 30%

basics

~20 s

On RabbitMQ, priority is ignored unless the queue was declared with x-max-priority, and higher numbers go first. On Redis, kombu emulates it with separate lists where 0 is served first, so 9 is the lowest priority.

open as a page

In Celery, where does a task sent with apply_async(countdown=...) wait before it runs, and why is a countdown of hours risky?

level: seniorimportance: should knowfreq 35%

basics

~20 s

apply_async turns countdown into an eta and publishes at once. On Redis or a classic RabbitMQ queue a worker takes the message immediately and holds it unacknowledged in memory until due, so hours-long waits pile up and can run twice.

open as a page

A Celery invoice task declared with rate_limit='10/m' starts about 40 times a minute across four workers; why, and what does rate_limit actually enforce?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Celery's rate_limit is a token bucket kept separately by each worker instance for each task type, not a cluster-wide limit. Four workers at '10/m' allow about 40 starts a minute; tasks over the limit wait inside the worker rather than failing.

open as a page

For a Celery video platform mixing transcodes, password resets and notifications, how would you decide how many queues to run and what each split costs?

level: principalimportance: should knowfreq 25%

basics

~20 s

Split queues by latency and resource class, not per task: a fast lane for resets and notifications, a heavy lane for transcodes, maybe a bulk lane. Each extra queue buys isolation but costs a worker fleet, scaling rules and alerts.

open as a page

How would you size and split a Celery worker fleet running CPU-heavy report rendering and thousands of webhook calls: pools, concurrency, prefetch and autoscale?

level: principalimportance: should knowfreq 25%

basics

~20 s

Split Celery into two worker deployments on separate queues: prefork reports with concurrency bounded by cores and RAM, prefetch 1 and child recycling; gevent webhooks with hundreds of greenlets and a larger prefetch. Scale mostly by adding workers.

open as a page

In Celery, how do group, chunks() and starmap() differ when resizing 10,000 archived photos, in task messages sent and in parallelism?

level: middleimportance: nice to knowfreq 16%

basics

~20 s

In Celery, a group sends one message per call, so 10,000 photos means 10,000 parallel tasks. starmap() sends one message and runs every call in sequence inside one task. chunks(it, n) sends a group of starmap tasks, each running n calls in sequence.

open as a page

In Celery 5.5 and later, what changes when a RabbitMQ payouts queue becomes a quorum queue, and how does native delayed delivery keep countdown tasks working?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

Celery 5.5+ detects RabbitMQ quorum queues and disables global QoS, so prefetch becomes static and --autoscale stops working. Countdown tasks would then block workers, so Celery enables native delayed delivery: RabbitMQ holds the message until due, which needs a non-direct exchange.

open as a page

showing 31–43 of 43