skip to content

A SaaS runs `celery -A proj worker -B` in each of three identical containers and bills every customer three times nightly; why, and how should Celery beat be deployed?

level: seniorimportance: must knowfreq 45%

answer

  1. count the schedulers, not workers
  2. what -B starts inside a worker
  3. no lock between beat processes
  4. one standalone beat, one replica

basics

~20 s

The -B flag embeds a beat scheduler in every worker, so three containers run three independent schedulers and each sends the billing task at 02:00. Celery has no lock between beat processes: run exactly one standalone celery beat.

solid answer

~50 s

`worker -B` starts an embedded beat alongside that worker, so three identical containers run three schedulers. Each keeps its own state, with the default `PersistentScheduler` a separate `celerybeat-schedule` file in each container, and nothing coordinates them: Celery's docs say you must ensure only one scheduler runs for a schedule, and Celery ships no lock to enforce that. At 02:00 all three decide the billing entry is due and each sends a message, so workers run billing three times. The fix is to drop `-B` from the workers and run one dedicated `celery -A proj beat` process, deployed as a single replica and rolled out stop-then-start so an old and a new beat never overlap. Switching to django-celery-beat's `DatabaseScheduler` does not fix it by itself; it has no lock either. Keep the billing task idempotent as a second line of defence.

code

bash · 8 lines
bash
# Before: every container embeds its own scheduler
celery -A saas worker -B -l INFO

# After: every worker container runs only a worker
celery -A saas worker -l INFO

# After: one dedicated beat process, a single replica
celery -A saas beat -l INFO -s /var/lib/celery/celerybeat-schedule

go deeper

for a junior

Recall that beat is the scheduler and that -B puts a scheduler inside each worker; three schedulers mean three sends.

for a middle

Explain that each beat keeps its own state and that nothing in Celery coordinates two beat processes, so the fix is structural, not a setting.

for a senior

Lay out the deployment: plain workers, one beat replica with stop-before-start rollouts, a persisted schedule file, alerting on beat, and an idempotent billing task.

for a principal

Decide whether a single beat as a scheduling single point of failure is acceptable for billing, or whether the job needs an external scheduler with claims.

## What `-B` actually starts A Celery **worker** consumes task messages and runs them. **Beat** is the scheduler that sends periodic task messages when entries in the schedule come due. Normally they are separate processes. The worker's **`-B` (`--beat`) option** folds the two together: the worker starts an **embedded beat service**, a child process by default, that runs the same scheduler loop `celery beat` would. Celery's own guide describes `-B` as convenient "if you'll never run more than one worker node" and not recommended for production. The code adds two hard limits: on Windows the option is rejected outright, and with the `gevent` or `eventlet` pools the worker raises `ImproperlyConfigured` and tells you to use a standalone beat instead. ## Why three containers bill three times The scenario: a SaaS with a nightly billing run, an hourly cache warm-up and per-customer report schedules, deployed as three identical containers, each started with `celery -A proj worker -B`. 1. Each container starts a worker, and each worker starts its own embedded beat. 2. Each beat loads the same `beat_schedule` from the shared code and opens **its own** state file, `celerybeat-schedule`, inside its own container. 3. At 02:00 each scheduler independently finds the `nightly-billing` entry due and publishes a task message. Each logs `Scheduler: Sending due task nightly-billing ...`. 4. The broker now holds three billing messages. Workers are doing their job correctly: they run all three. 5. The hourly warm-up and every customer report are also tripled; billing is simply where it hurts first. Nothing here is a broker redelivery or a retry. Three schedulers each did exactly what they were told. Celery's periodic-tasks guide says it plainly: you have to ensure **only a single scheduler is running for a schedule at a time**, otherwise you get duplicate tasks. The centralized design is deliberate, so beat needs no locks, and Celery therefore provides none. ## Fixes that do not fix it | Candidate fix | Why it still duplicates | |---|---| | Switch to django-celery-beat's `DatabaseScheduler` | Three beats read the same rows, and each still sends; the scheduler has no lock | | Mount one shared `celerybeat-schedule` file | A `shelve` file stores last-send times; it is not a lock, and each process still decides for itself | | Scale the containers down to two | Two schedulers, two bills | | Rely on `acks_late` or retries | They govern how one message is acknowledged or re-run; these are three distinct messages | ## The deployment that works - **Workers without `-B`.** All three containers run plain `celery -A proj worker`, and they can scale freely. - **One standalone beat.** A separate process or deployment runs `celery -A proj beat`, pinned to **exactly one replica**. - **Stop before start on deploy.** A rollout that starts the new beat before stopping the old one briefly runs two schedulers; if an entry falls due in that window it is sent twice. Use a replace-style rollout for the beat deployment. - **Persist or accept losing the state file.** With `PersistentScheduler`, keep `celerybeat-schedule` on storage that survives restarts (set its path with `-s` or `beat_schedule_filename`). If it is lost, entries restart with "last sent = now" and a run missed during the outage is not caught up. - **Monitor the one beat.** A single beat is a single point of failure for scheduling: if it is down at 02:00, billing is simply not sent until it comes back. Alert on it. ## Defence in depth Even with one beat, a billing task that can charge a customer twice for one period is fragile: a deploy mistake, a manual re-run or an operator starting a second beat "to test" would repeat the damage. Make the task idempotent per billing period, for example by recording a unique (customer, period) charge before calling the payment provider, so a duplicate send becomes a no-op. The general theory of firing a schedule once across a fleet (leader election, per-tick claims) belongs to distributed cron; the Celery-specific answer is simply one beat. ## How to spot it - Beat logs a `Scheduler: Sending due task` line per send, so three such lines per night, one per container, is the smoking gun. - Count processes titled `celery beat` across the fleet: the answer must be one.

  • Why not keep `-B` on just one of the three Celery worker containers?
    It works until the containers stop being identical in practice: that one worker becomes special, so scaling, restarting or replacing it now also restarts or duplicates the scheduler. Celery's guide recommends `-B` only when you will never run more than one worker node, and the option is refused with the `gevent` and `eventlet` pools. A dedicated beat deployment makes the one-scheduler rule explicit.
  • With one standalone Celery beat, what happens if it is down at 02:00?
    Nothing is sent at 02:00. When beat returns with its `celerybeat-schedule` file intact, the overdue billing entry is sent once, straight away, unless `beat_cron_starting_deadline` tells beat to skip runs older than that many seconds. If the file was lost, entries restart with last send equal to now, and that night's run is skipped silently.
  • Does django-celery-beat's `DatabaseScheduler` let several Celery beat processes share the work safely?
    No. It moves the schedule into database tables that an admin screen can edit, and each beat process reads those rows and sends whatever is due. Nothing in it claims an entry before sending, so two beats against one database still send every entry twice.

Three identical alarm clocks set for 02:00, each wired to the same doorbell: each rings on its own, so the bell rings three times. The fix is one alarm clock, not a smarter doorbell.

saying these in an interview costs you the question

  • Beat processes coordinate through the broker so only one sends each entry
  • Switching to DatabaseScheduler makes several beat processes safe
  • worker -B is the recommended production way to run beat
  • Duplicate billing means the broker redelivered the message
  • Running three beat replicas gives scheduling high availability for free