Background Jobs & Scheduling
Getting work off the request path: task queues, background workers and scheduled jobs in .NET and Python. Asked because retries and at-least-once delivery are where it silently fails.
on this pageshowhide
explore
- Hangfireempty
- Quartzempty
- Coravelempty
- Celery43 questions
- Defining & Calling Tasks6 questions
- Task Canvas6 questions
- Execution Pools & Prefetch5 questions
- Brokers and Result Backends6 questions
- Routing and Queues5 questions
- Retries & Acknowledgement5 questions
- Beat Scheduling5 questions
- Monitoring & Shutdown5 questions
questions
page 2 of 2In Celery, what happens to hourly and nightly beat entries when beat, or every worker, is down for three hours, and which options change that?
basics
~20 sIf beat was down, each overdue entry is sent once on restart, not once per missed slot. If workers were down, beat kept sending and messages queued. An expires option drops stale copies; beat_cron_starting_deadline skips crontab runs that are too late.
Why does a Celery chord need a result backend, and what happens to the chord body when one thumbnail task in its header fails?
basics
~20 sA Celery chord needs a result backend because something must record each header task's completion and result before the body can start. If one header task fails, the others still run, but the body never runs and is marked FAILURE with a ChordError.
In Celery, you revoke a hung PDF-export task by id, yet it keeps running, and a revoked queued export later runs anyway; why, and what are terminate's risks?
basics
~20 srevoke only broadcasts the id to workers, which skip that task when they reach it; a running task continues unless terminate=True signals its pool process. Revoked ids live in worker memory, so a full restart forgets them unless workers use --statedb.
During a rolling deploy, a Celery 5.6 worker gets SIGTERM mid-way through a long PDF export and is killed after the grace period; what do warm, soft and cold shutdown do, and how do you keep the task?
basics
~20 sSIGTERM starts a warm shutdown that waits, unbounded, for running tasks, so the later SIGKILL loses them. Cold shutdown (SIGQUIT) cancels them; soft shutdown (5.5+) first waits worker_soft_shutdown_timeout. Use REMAP_SIGTERM=SIGQUIT, a timeout under the grace period, and acks_late.
In Celery, when should a task raise `Reject` or `Ignore` instead of calling `self.retry()`, and what happens to the message and its stored state?
basics
~20 sUse self.retry() for transient failures: a new message runs the task again. Raise Reject for a message that should leave the queue, dead-lettered or requeued, which needs acks_late. Raise Ignore when the work is moot: it acks and records nothing.
A Celery app routes transcodes to a new 'transcode' queue, and they pile up unprocessed while other tasks run; why, and what does declaring task_queues change?
basics
~20 sRouting creates the queue, but no worker consumes it: a worker without -Q drains only task_queues, by default just celery. Declaring task_queues and disabling task_create_missing_queues makes bare workers consume every declared queue and turns typos into errors.
In Celery, why might apply_async(priority=9) fail to move password-reset emails ahead of queued transcodes, and how do RabbitMQ and Redis priorities differ?
basics
~20 sOn RabbitMQ, priority is ignored unless the queue was declared with x-max-priority, and higher numbers go first. On Redis, kombu emulates it with separate lists where 0 is served first, so 9 is the lowest priority.
In Celery, where does a task sent with apply_async(countdown=...) wait before it runs, and why is a countdown of hours risky?
basics
~20 sapply_async turns countdown into an eta and publishes at once. On Redis or a classic RabbitMQ queue a worker takes the message immediately and holds it unacknowledged in memory until due, so hours-long waits pile up and can run twice.
A Celery invoice task declared with rate_limit='10/m' starts about 40 times a minute across four workers; why, and what does rate_limit actually enforce?
basics
~20 sCelery's rate_limit is a token bucket kept separately by each worker instance for each task type, not a cluster-wide limit. Four workers at '10/m' allow about 40 starts a minute; tasks over the limit wait inside the worker rather than failing.
For a Celery video platform mixing transcodes, password resets and notifications, how would you decide how many queues to run and what each split costs?
basics
~20 sSplit queues by latency and resource class, not per task: a fast lane for resets and notifications, a heavy lane for transcodes, maybe a bulk lane. Each extra queue buys isolation but costs a worker fleet, scaling rules and alerts.
How would you size and split a Celery worker fleet running CPU-heavy report rendering and thousands of webhook calls: pools, concurrency, prefetch and autoscale?
basics
~20 sSplit Celery into two worker deployments on separate queues: prefork reports with concurrency bounded by cores and RAM, prefetch 1 and child recycling; gevent webhooks with hundreds of greenlets and a larger prefetch. Scale mostly by adding workers.
In Celery, how do group, chunks() and starmap() differ when resizing 10,000 archived photos, in task messages sent and in parallelism?
basics
~20 sIn Celery, a group sends one message per call, so 10,000 photos means 10,000 parallel tasks. starmap() sends one message and runs every call in sequence inside one task. chunks(it, n) sends a group of starmap tasks, each running n calls in sequence.
In Celery 5.5 and later, what changes when a RabbitMQ payouts queue becomes a quorum queue, and how does native delayed delivery keep countdown tasks working?
basics
~20 sCelery 5.5+ detects RabbitMQ quorum queues and disables global QoS, so prefetch becomes static and --autoscale stops working. Countdown tasks would then block workers, so Celery enables native delayed delivery: RabbitMQ holds the message until due, which needs a non-direct exchange.
showing 31–43 of 43