skip to content

Beat Scheduling

celery beat sends periodic tasks from beat_schedule with crontab, timedelta or solar entries, or from the database via django-celery-beat. Asked because two beat processes send every task twice.

on this pageshow

explore

questions

5

In Celery, what does the celery beat process do, and how do you declare a periodic task in the beat_schedule setting?

level: juniorimportance: must knowfreq 55%

answer

  1. a clock that only sends
  2. workers still do the work
  3. entry name, task, schedule
  4. seconds, timedelta, crontab, solar
  5. celerybeat-schedule file

basics

~20 s

celery beat is a scheduler process that sends a task message whenever a beat_schedule entry comes due; workers run it. Each entry names a registered task, a schedule (seconds, timedelta, crontab or solar) and optional args, kwargs and options.

solid answer

~40 s

`celery beat` is a separate process that runs no task code. It reads the schedule, and when an entry is due it sends an ordinary task message to the broker, just as `apply_async()` would; whichever worker consumes that queue executes it. By default the entries come from `app.conf.beat_schedule`, a dict keyed by a unique entry name. Each value holds `task` (the registered task name, as a string), `schedule` (a number of seconds, a `timedelta`, a `crontab(...)` or a `solar(...)`), and optionally `args`, `kwargs` and `options`, which accepts any `apply_async()` option such as `queue` or `expires`. You start it with `celery -A proj beat`. The default `PersistentScheduler` records last-send times in a local `celerybeat-schedule` file, and only one beat may run per schedule.

code

python · 19 lines
python
from datetime import timedelta

from celery import Celery
from celery.schedules import crontab

app = Celery("saas", broker="redis://localhost:6379/0")

app.conf.beat_schedule = {
    "nightly-billing": {
        "task": "billing.tasks.run_nightly_billing",
        "schedule": crontab(hour=2, minute=0),
    },
    "hourly-cache-warmup": {
        "task": "cache.tasks.warm_cache",
        "schedule": timedelta(hours=1),
        "kwargs": {"scope": "pricing"},
        "options": {"queue": "maintenance", "expires": 50 * 60},
    },
}

go deeper

for a junior

Recall that beat only sends messages on a timetable and workers do the work. Be able to write one beat_schedule entry with a task name and a crontab or timedelta.

for a middle

Explain the entry fields, including options passing apply_async arguments, the crontab wildcard defaults, and how a timedelta schedule counts from the last send.

for a senior

Show you know beat records sends, not outcomes, so overlaps and queued backlogs are yours to handle, and that exactly one beat must run per schedule.

for a principal

Weigh a static beat_schedule in code against a database-backed schedule that operators edit, and who owns correctness when schedules change at runtime.

## What beat is, and what it is not **`celery beat`** is Celery's periodic scheduler. It is a long-running process of its own, started with `celery -A proj beat`, and its whole job is to decide *when* a message should be sent. It does **not** execute task bodies. When an entry comes due, beat publishes an ordinary task message to the broker, and from that moment the message is indistinguishable from one sent by `apply_async()` in a web request: it lands in a queue, a **worker** consuming that queue picks it up, and the worker runs the function. That split explains most beat behaviour: - If no worker is running, beat still sends on time and the messages wait in the broker. - Beat does not wait for a run to finish before sending the next one, so a slow task can overlap with its own next run. - Beat records the time it **sent** an entry (`last_run_at`), not whether the task succeeded. In the reserved scenario, a SaaS product, beat sends the nightly billing run, the hourly cache warm-up and the customers' report jobs; the workers in the application's containers do the actual billing, warming and reporting. ## Declaring entries in `beat_schedule` The default source of entries is the **`beat_schedule`** setting: a dict whose keys are unique entry names and whose values describe one periodic send each. | Field | Required | Meaning | |---|---|---| | `task` | yes | The registered task name as a string, such as `'billing.tasks.run_nightly_billing'` | | `schedule` | yes | A number of seconds, a `datetime.timedelta`, a `crontab(...)` or a `solar(...)` | | `args` | no | Positional arguments, a list or tuple | | `kwargs` | no | Keyword arguments, a dict | | `options` | no | Any `apply_async()` option: `queue`, `routing_key`, `expires`, `priority` and so on | | `relative` | no | For `timedelta` schedules, round the period to the clock instead of counting from beat's start | Two small traps live in this table. `task` is the task's **registered name**, not the function object; by default that name is built from the module path, but it is a name all the same. And a one-item `args` tuple needs its trailing comma, `(42,)`, because `(42)` is just the integer 42. ## The schedule types | Type | Example | Fires | |---|---|---| | Number or `timedelta` | `timedelta(hours=1)` | One interval after the last send; a new entry is first sent one interval after beat starts | | `crontab` | `crontab(hour=2, minute=0)` | At matching wall-clock times in the app's `timezone` (UTC by default) | | `solar` | `solar('sunset', -37.81753, 144.96715)` | At a sun event for a latitude and longitude, computed in UTC | `crontab` fields default to `'*'`, which surprises people: `crontab(day_of_week='sunday')` fires **every minute** on Sundays, because minute and hour were left as wildcards. Write `crontab(minute=0, hour=2)` when you mean 02:00 every day. Entries can also be added in code with `app.add_periodic_task(schedule, signature, name=...)`, usually from an `on_after_configure` or `on_after_finalize` signal handler. It writes into `beat_schedule` behind the scenes, and two entries built from the same signature need distinct `name`s or the second replaces the first. ## Running beat and where it keeps state 1. Start one beat process: `celery -A proj beat -l INFO`. 2. Start workers separately: `celery -A proj worker -l INFO`. 3. Let beat write its state file. The default scheduler, `celery.beat:PersistentScheduler`, keeps each entry's last send time in a `shelve` file named `celerybeat-schedule` (the `beat_schedule_filename` setting, or `-s` on the command line), so it needs a writable directory. The scheduler class is itself a setting, `beat_scheduler`. The common alternative is `django_celery_beat.schedulers:DatabaseScheduler`, which reads entries from database tables that an admin screen can edit. ## Only one beat per schedule Celery's documentation is explicit that you must make sure **only a single scheduler runs for a schedule at a time**, or tasks are sent more than once. Beat processes do not coordinate with each other and Celery ships no lock between them. That is why the worker's `-B` flag, which embeds a beat in each worker, is fine on a single node and dangerous on a fleet. ## Mistakes that show up in interviews - Saying beat "runs" the task: it only sends it. - Expecting beat to skip a send because the previous run is still going. - Leaving `crontab` fields as wildcards by accident. - Starting beat in every container "for redundancy".

  • Does Celery beat wait for the previous run of a periodic task to finish before sending the next one?
    No. Beat only sends messages on its timetable and never hears back from the workers. If the hourly warm-up takes ninety minutes, the next message is sent on time and a second worker can start it while the first is still running. When overlap matters, the task itself has to guard against it, for example with a lock it takes at the start.
  • Why does `crontab(day_of_week='sunday')` in a Celery beat entry fire far more often than once a week?
    Every `crontab` field you leave out defaults to `'*'`. With only `day_of_week` set, minute and hour are wildcards, so the entry matches every minute of every Sunday. A weekly job needs explicit fields, such as `crontab(minute=0, hour=3, day_of_week='sunday')`.
  • What does the `options` key of a Celery beat entry accept?
    Any keyword that `apply_async()` accepts: `queue`, `routing_key`, `exchange`, `priority`, `expires` and so on. Beat passes them through when it sends the message, so an entry can go to a dedicated queue or be discarded by the worker if it is not started before its expiry.

Beat is a school bell on a timetable: it rings at the scheduled time and has no idea whether the teachers are in the room. The teaching is done by the teachers, as the task is done by workers.

saying these in an interview costs you the question

  • celery beat executes the periodic task code itself
  • Every worker automatically runs its own copy of the beat schedule
  • crontab(day_of_week='sunday') runs once each Sunday
  • An entry's task field takes the decorated function object
  • Beat waits for the previous run to finish before sending again
  • Running beat in every container gives redundancy
open as a page

A SaaS runs `celery -A proj worker -B` in each of three identical containers and bills every customer three times nightly; why, and how should Celery beat be deployed?

level: seniorimportance: must knowfreq 45%

basics

~20 s

The -B flag embeds a beat scheduler in every worker, so three containers run three independent schedulers and each sends the billing task at 02:00. Celery has no lock between beat processes: run exactly one standalone celery beat.

open as a page

How does django-celery-beat's DatabaseScheduler let an admin screen change Celery beat schedules at runtime, and when can an edit go unnoticed for a while?

level: middleimportance: should knowfreq 33%

basics

~20 s

DatabaseScheduler builds beat's schedule from PeriodicTask rows and checks a one-row PeriodicTasks change marker, bumped by save and delete signals, about every 5 seconds. Bulk update() or raw SQL skips those signals, so beat notices only at its next full reload.

open as a page

In Celery, how do the timezone and enable_utc settings decide when a beat entry with crontab(hour=2, minute=0) fires, and what must change afterwards?

level: middleimportance: should knowfreq 35%

basics

~20 s

crontab fields are matched against the wall clock of the app's timezone; with timezone unset and enable_utc True, that is UTC, so hour=2 means 02:00 UTC. PersistentScheduler resets itself when a stored zone changes; DatabaseScheduler needs a manual reset.

open as a page

In Celery, what happens to hourly and nightly beat entries when beat, or every worker, is down for three hours, and which options change that?

level: seniorimportance: should knowfreq 28%

basics

~20 s

If beat was down, each overdue entry is sent once on restart, not once per missed slot. If workers were down, beat kept sending and messages queued. An expires option drops stale copies; beat_cron_starting_deadline skips crontab runs that are too late.

open as a page