How does django-celery-beat's DatabaseScheduler let an admin screen change Celery beat schedules at runtime, and when can an edit go unnoticed for a while?
answer
- schedules as database rows
- a one-row change marker
- signals bump it, bulk updates don't
- five-second loop, five-minute resync
basics
~20 sDatabaseScheduler builds beat's schedule from PeriodicTask rows and checks a one-row PeriodicTasks change marker, bumped by save and delete signals, about every 5 seconds. Bulk update() or raw SQL skips those signals, so beat notices only at its next full reload.
solid answer
~40 sWith `-S django_celery_beat.schedulers:DatabaseScheduler` (or the `beat_scheduler` setting), beat builds its schedule from `PeriodicTask` rows, each pointing at one `IntervalSchedule`, `CrontabSchedule`, `SolarSchedule` or `ClockedSchedule` and carrying JSON `args`/`kwargs`, `queue`, `enabled`, `one_off` and `start_time`. Saving or deleting those rows through the ORM fires signals that bump `PeriodicTasks.last_update`, a single-row change marker. On every tick beat compares that timestamp with the last one it saw and reloads the enabled rows when it moved; its loop sleeps at most 5 seconds by default, so admin edits land within seconds. A `QuerySet.update()` or raw SQL fires no signals, so beat misses it until the forced full reload that 2.9.0 does every five minutes, or until you call `PeriodicTasks.update_changed()`. It is still one scheduler: run one beat.
code
python · 23 linesimport json
from django_celery_beat.models import CrontabSchedule, PeriodicTask, PeriodicTasks
def save_report_schedule(customer_id, hour, tz_name):
cron, _ = CrontabSchedule.objects.get_or_create(
minute="0", hour=str(hour), day_of_week="1",
day_of_month="*", month_of_year="*", timezone=tz_name,
)
PeriodicTask.objects.update_or_create(
name=f"weekly-report-{customer_id}",
defaults={
"task": "reports.tasks.send_customer_report",
"crontab": cron,
"kwargs": json.dumps({"customer_id": customer_id}),
},
) # a save: signals bump the change marker
def pause_all_reports():
PeriodicTask.objects.filter(name__startswith="weekly-report-").update(enabled=False)
PeriodicTasks.update_changed() # update() fired no signalsgo deeper
Recall that django-celery-beat stores periodic tasks as database rows that an admin can edit, read by DatabaseScheduler.
Explain the change marker: signals bump PeriodicTasks.last_update, beat polls it on a short loop, and bulk updates skip the signals.
Build admin tooling that saves through the ORM or calls update_changed(), and still deploy exactly one beat against the database.
Decide whether customer-editable schedules belong in beat at all, or in an application table a single periodic task scans.
## Why a database scheduler at all The default Celery scheduler, `PersistentScheduler`, takes its entries from the `beat_schedule` setting in code. Changing a schedule means a deploy and a beat restart. In the reserved scenario, a SaaS lets each customer choose when their report is emailed, from an admin screen, and support staff can pause a customer's schedule. That needs schedules as **data**, which is what **django-celery-beat** provides: Django models for schedules plus a beat scheduler class, `DatabaseScheduler`, that reads them. ## The models | Model | Holds | |---|---| | `PeriodicTask` | One entry: unique `name`, `task` name, JSON `args` and `kwargs`, `queue`, `priority`, `enabled`, `one_off`, `start_time`, `expires`/`expire_seconds`, plus bookkeeping `last_run_at` and `total_run_count` | | `IntervalSchedule` | `every` plus `period` (for example every 2 days) | | `CrontabSchedule` | Cron fields plus its own `timezone`, so rows can fire in different zones | | `SolarSchedule` | A sun event at a latitude and longitude | | `ClockedSchedule` | One exact time; the task must be `one_off` | | `PeriodicTasks` | A single row (`ident=1`) with `last_update`: the change marker | A `PeriodicTask` points at exactly one schedule; the model's validation rejects a row with more than one. Entries still declared in `beat_schedule` are not ignored: at start-up the scheduler writes them into `PeriodicTask` rows too. ## How beat learns about an edit 1. Beat runs with `celery -A proj beat -S django_celery_beat.schedulers:DatabaseScheduler`, or with that class path in the `beat_scheduler` setting. 2. When the admin screen saves or deletes a `PeriodicTask`, or saves a schedule row, Django model signals call `PeriodicTasks.update_changed()`, which stamps `last_update` with the current time. 3. Each time beat's loop reads the schedule, the scheduler compares `last_update` with the value it saw last. If it moved, it reloads all enabled `PeriodicTask` rows and rebuilds its internal queue of due times. 4. The loop never sleeps longer than `max_interval`. For `DatabaseScheduler` that defaults to **5 seconds** (the Celery default scheduler uses 300), unless `beat_max_loop_interval` overrides it, so an admin edit takes effect within seconds. Beat's own bookkeeping writes, updating `last_run_at` and `total_run_count` after each send, are flagged so they do not bump the marker; otherwise every send would force a reload. ## When an edit goes unnoticed - **Bulk `QuerySet.update()`**, such as pausing 500 customers' reports with `.update(enabled=False)`, fires no model signals, so the marker does not move. The same goes for raw SQL, a data migration writing directly, or an import script that bypasses `save()`. - In **django-celery-beat 2.9.0** the scheduler also forces a full reload every **five minutes**, so such an edit is picked up late rather than never. Until then beat keeps sending the old schedule. - The package's README gives the fix: after a bulk change, call **`PeriodicTasks.update_changed()`** yourself. ## Other runtime behaviours worth knowing - **`enabled=False`** stops an entry without deleting the row; saving it disabled also clears its `last_run_at`, and re-enabling takes effect on the next reload. - **`one_off=True`** disables the row after its first send; `ClockedSchedule` requires it. - **`start_time`** holds an entry back until that moment. - Each sent message carries a `periodic_task_name` header naming the row, useful when tracing which schedule sent a task. - Changing Celery's app `timezone` does **not** reset stored `last_run_at` values under this scheduler; reset them and call `update_changed()`. ## Designing the admin screen around it 1. Save through the ORM, one row at a time (`save()`, `update_or_create()`), so the signals fire; the model's `save()` also runs its own validation. 2. Give each customer's entry a stable, unique `name`, such as `weekly-report-<customer id>`, so edits update the same row rather than piling up new ones. 3. Reuse schedule rows with `get_or_create()`; many customers can point at the same `CrontabSchedule`. 4. After any bulk operation, call `PeriodicTasks.update_changed()` in the same code path, not as a manual afterthought. 5. Keep `args` and `kwargs` to plain JSON values such as a customer id: the row stores them as JSON text. The scheduler also installs Celery's `celery.backend_cleanup` entry, daily at 04:00, whenever `result_expires` is set, so do not be surprised to find a row for it in the admin. ## What it does not change A database scheduler is still **one scheduler**. Several beat processes pointed at the same database each read the rows and each send what is due; the scheduler takes no lock. Run exactly one beat, whatever the storage.
- Why does django-celery-beat's DatabaseScheduler default to a 5-second loop when Celery's default scheduler uses 300 seconds?Celery's default scheduler reads a schedule that only changes on restart, so between due times it can sleep long. A database schedule can change at any moment from outside the process, and beat only notices on a tick. A short maximum sleep bounds how stale the in-memory schedule can be. `beat_max_loop_interval` overrides either value.
- Can a customer's django-celery-beat crontab row fire in the customer's own time zone rather than Celery's?Yes. `CrontabSchedule` has its own `timezone` field, defaulting to the Celery time zone setting or UTC, and the scheduler evaluates the cron fields on that zone's clock. One beat can send a 07:00 New York report and a 07:00 Tokyo report from separate rows.
saying these in an interview costs you the question
- DatabaseScheduler re-reads every row only when beat restarts
- A bulk QuerySet.update() on PeriodicTask is picked up immediately
- Entries in beat_schedule are ignored once DatabaseScheduler is used
- Storing schedules in the database lets several beat processes share the work
- The admin save pushes a control message to the running beat