In Celery, what does the celery beat process do, and how do you declare a periodic task in the beat_schedule setting?
answer
- a clock that only sends
- workers still do the work
- entry name, task, schedule
- seconds, timedelta, crontab, solar
- celerybeat-schedule file
basics
~20 scelery beat is a scheduler process that sends a task message whenever a beat_schedule entry comes due; workers run it. Each entry names a registered task, a schedule (seconds, timedelta, crontab or solar) and optional args, kwargs and options.
solid answer
~40 s`celery beat` is a separate process that runs no task code. It reads the schedule, and when an entry is due it sends an ordinary task message to the broker, just as `apply_async()` would; whichever worker consumes that queue executes it. By default the entries come from `app.conf.beat_schedule`, a dict keyed by a unique entry name. Each value holds `task` (the registered task name, as a string), `schedule` (a number of seconds, a `timedelta`, a `crontab(...)` or a `solar(...)`), and optionally `args`, `kwargs` and `options`, which accepts any `apply_async()` option such as `queue` or `expires`. You start it with `celery -A proj beat`. The default `PersistentScheduler` records last-send times in a local `celerybeat-schedule` file, and only one beat may run per schedule.
code
python · 19 linesfrom datetime import timedelta
from celery import Celery
from celery.schedules import crontab
app = Celery("saas", broker="redis://localhost:6379/0")
app.conf.beat_schedule = {
"nightly-billing": {
"task": "billing.tasks.run_nightly_billing",
"schedule": crontab(hour=2, minute=0),
},
"hourly-cache-warmup": {
"task": "cache.tasks.warm_cache",
"schedule": timedelta(hours=1),
"kwargs": {"scope": "pricing"},
"options": {"queue": "maintenance", "expires": 50 * 60},
},
}go deeper
Recall that beat only sends messages on a timetable and workers do the work. Be able to write one beat_schedule entry with a task name and a crontab or timedelta.
Explain the entry fields, including options passing apply_async arguments, the crontab wildcard defaults, and how a timedelta schedule counts from the last send.
Show you know beat records sends, not outcomes, so overlaps and queued backlogs are yours to handle, and that exactly one beat must run per schedule.
Weigh a static beat_schedule in code against a database-backed schedule that operators edit, and who owns correctness when schedules change at runtime.
## What beat is, and what it is not **`celery beat`** is Celery's periodic scheduler. It is a long-running process of its own, started with `celery -A proj beat`, and its whole job is to decide *when* a message should be sent. It does **not** execute task bodies. When an entry comes due, beat publishes an ordinary task message to the broker, and from that moment the message is indistinguishable from one sent by `apply_async()` in a web request: it lands in a queue, a **worker** consuming that queue picks it up, and the worker runs the function. That split explains most beat behaviour: - If no worker is running, beat still sends on time and the messages wait in the broker. - Beat does not wait for a run to finish before sending the next one, so a slow task can overlap with its own next run. - Beat records the time it **sent** an entry (`last_run_at`), not whether the task succeeded. In the reserved scenario, a SaaS product, beat sends the nightly billing run, the hourly cache warm-up and the customers' report jobs; the workers in the application's containers do the actual billing, warming and reporting. ## Declaring entries in `beat_schedule` The default source of entries is the **`beat_schedule`** setting: a dict whose keys are unique entry names and whose values describe one periodic send each. | Field | Required | Meaning | |---|---|---| | `task` | yes | The registered task name as a string, such as `'billing.tasks.run_nightly_billing'` | | `schedule` | yes | A number of seconds, a `datetime.timedelta`, a `crontab(...)` or a `solar(...)` | | `args` | no | Positional arguments, a list or tuple | | `kwargs` | no | Keyword arguments, a dict | | `options` | no | Any `apply_async()` option: `queue`, `routing_key`, `expires`, `priority` and so on | | `relative` | no | For `timedelta` schedules, round the period to the clock instead of counting from beat's start | Two small traps live in this table. `task` is the task's **registered name**, not the function object; by default that name is built from the module path, but it is a name all the same. And a one-item `args` tuple needs its trailing comma, `(42,)`, because `(42)` is just the integer 42. ## The schedule types | Type | Example | Fires | |---|---|---| | Number or `timedelta` | `timedelta(hours=1)` | One interval after the last send; a new entry is first sent one interval after beat starts | | `crontab` | `crontab(hour=2, minute=0)` | At matching wall-clock times in the app's `timezone` (UTC by default) | | `solar` | `solar('sunset', -37.81753, 144.96715)` | At a sun event for a latitude and longitude, computed in UTC | `crontab` fields default to `'*'`, which surprises people: `crontab(day_of_week='sunday')` fires **every minute** on Sundays, because minute and hour were left as wildcards. Write `crontab(minute=0, hour=2)` when you mean 02:00 every day. Entries can also be added in code with `app.add_periodic_task(schedule, signature, name=...)`, usually from an `on_after_configure` or `on_after_finalize` signal handler. It writes into `beat_schedule` behind the scenes, and two entries built from the same signature need distinct `name`s or the second replaces the first. ## Running beat and where it keeps state 1. Start one beat process: `celery -A proj beat -l INFO`. 2. Start workers separately: `celery -A proj worker -l INFO`. 3. Let beat write its state file. The default scheduler, `celery.beat:PersistentScheduler`, keeps each entry's last send time in a `shelve` file named `celerybeat-schedule` (the `beat_schedule_filename` setting, or `-s` on the command line), so it needs a writable directory. The scheduler class is itself a setting, `beat_scheduler`. The common alternative is `django_celery_beat.schedulers:DatabaseScheduler`, which reads entries from database tables that an admin screen can edit. ## Only one beat per schedule Celery's documentation is explicit that you must make sure **only a single scheduler runs for a schedule at a time**, or tasks are sent more than once. Beat processes do not coordinate with each other and Celery ships no lock between them. That is why the worker's `-B` flag, which embeds a beat in each worker, is fine on a single node and dangerous on a fleet. ## Mistakes that show up in interviews - Saying beat "runs" the task: it only sends it. - Expecting beat to skip a send because the previous run is still going. - Leaving `crontab` fields as wildcards by accident. - Starting beat in every container "for redundancy".
- Does Celery beat wait for the previous run of a periodic task to finish before sending the next one?No. Beat only sends messages on its timetable and never hears back from the workers. If the hourly warm-up takes ninety minutes, the next message is sent on time and a second worker can start it while the first is still running. When overlap matters, the task itself has to guard against it, for example with a lock it takes at the start.
- Why does `crontab(day_of_week='sunday')` in a Celery beat entry fire far more often than once a week?Every `crontab` field you leave out defaults to `'*'`. With only `day_of_week` set, minute and hour are wildcards, so the entry matches every minute of every Sunday. A weekly job needs explicit fields, such as `crontab(minute=0, hour=3, day_of_week='sunday')`.
- What does the `options` key of a Celery beat entry accept?Any keyword that `apply_async()` accepts: `queue`, `routing_key`, `exchange`, `priority`, `expires` and so on. Beat passes them through when it sends the message, so an entry can go to a dedicated queue or be discarded by the worker if it is not started before its expiry.
Beat is a school bell on a timetable: it rings at the scheduled time and has no idea whether the teachers are in the room. The teaching is done by the teachers, as the task is done by workers.
saying these in an interview costs you the question
- celery beat executes the periodic task code itself
- Every worker automatically runs its own copy of the beat schedule
- crontab(day_of_week='sunday') runs once each Sunday
- An entry's task field takes the decorated function object
- Beat waits for the previous run to finish before sending again
- Running beat in every container gives redundancy