In Celery, what happens to hourly and nightly beat entries when beat, or every worker, is down for three hours, and which options change that?
answer
- beat remembers sends, not results
- overdue collapses into one
- messages queue while workers sleep
- expires in the entry's options
- a deadline for late cron runs
basics
~20 sIf beat was down, each overdue entry is sent once on restart, not once per missed slot. If workers were down, beat kept sending and messages queued. An expires option drops stale copies; beat_cron_starting_deadline skips crontab runs that are too late.
solid answer
~50 sBeat stores each entry's last send time. When it restarts with its `celerybeat-schedule` file intact, an entry whose due time has passed is sent **once**, immediately, however many slots were missed; a `timedelta` entry then counts its next hour from that send, and a `crontab` entry waits for its next matching time. `beat_cron_starting_deadline` (default `None`) changes the crontab case: a missed run older than that many seconds is skipped rather than sent late. If the file is lost, entries restart with "last sent = now", and the missed runs vanish silently. If beat was up but the workers were down, beat kept sending every hour and three warm-ups sit in the queue to run back to back; set `options={'expires': ...}` below the interval so the worker discards stale copies. The nightly billing run should usually not expire: late beats never.
code
python · 19 linesfrom datetime import timedelta
from celery import Celery
from celery.schedules import crontab
app = Celery("saas", broker="redis://localhost:6379/0")
app.conf.beat_cron_starting_deadline = None # billing: late beats never
app.conf.beat_schedule = {
"nightly-billing": {
"task": "billing.tasks.run_nightly_billing",
"schedule": crontab(hour=2, minute=0),
},
"hourly-cache-warmup": {
"task": "cache.tasks.warm_cache",
"schedule": timedelta(hours=1),
"options": {"expires": 50 * 60}, # stale copies are discarded
},
}go deeper
Recall that beat sends messages on a timetable whether or not any worker is running to consume them.
Explain that an overdue entry is sent once on restart and that expires in an entry's options lets workers drop stale copies.
Decide per entry: expire the warm-up, never expire billing, persist the schedule file, and remember the cron starting deadline is app-wide, so it cannot treat billing and reports differently.
Set the organisation's policy for missed periodic work, late or skipped, per job class, and make it visible in alerts.
## Two different outages The reserved scenario is a SaaS with a **nightly billing run** at 02:00 (`crontab`), an **hourly cache warm-up** (`timedelta(hours=1)`) and per-customer report schedules, deployed as three identical containers with one beat. "Down for three hours" means two very different things depending on which process was down, because **beat only sends messages and remembers when it sent them**; it never learns whether a task ran. ## Case 1: beat was down Beat's default `PersistentScheduler` keeps every entry's **last send time** (`last_run_at`) in its `celerybeat-schedule` file. When beat starts again: 1. It loads each entry with its stored last send time. 2. For each entry it asks the schedule how long remains until the next run. Three hours after the last hourly send, that remaining time is negative, so the entry is **due**. 3. It sends the entry **once**, then stores the new send time. There is no loop over missed slots, so three missed warm-ups become one. 4. The next send is computed from the new last send time: a `timedelta` entry now runs one hour after the restart (its cadence has shifted to the restart minute), while a `crontab` entry waits for its next matching wall-clock time. For billing, if beat was down from 01:30 to 04:30, the 02:00 run is sent at 04:30, late but once. ### Skipping cron runs that are too late The setting **`beat_cron_starting_deadline`** (default `None`) bounds this for `crontab` schedules only. With `None`, a past-due cron run always runs immediately. With a number of seconds, beat checks the latest run it missed; if that run is older than the deadline, the entry is **not** sent and waits for its next time. Celery's docs warn against setting it above 3600. | Setting | Billing missed at 02:00, beat back at 04:30 | |---|---| | `beat_cron_starting_deadline = None` (default) | Sent once at 04:30 | | `beat_cron_starting_deadline = 1800` | Skipped; next run 02:00 tomorrow | ### When the state file is lost or stale - If `celerybeat-schedule` lived on a container's ephemeral disk and the container was replaced, entries are rebuilt with "last sent = now": the missed billing run is **silently skipped**, with no error anywhere. - The file is written periodically (every three minutes by default) and on a clean shutdown. If beat is killed hard shortly after sending, the last send may not be on disk, and the entry can be sent again after restart. `beat_sync_every = 1` writes after every send and narrows that window. ## Case 2: beat was up, the workers were down Beat carried on as normal: every hour it sent a warm-up message and recorded the send. Nobody consumed them, so the queue now holds **three warm-ups**, and when the workers return they run all three back to back, plus any report jobs that piled up. Beat will not hold back because a previous message is unconsumed; it has no way to know. The tool here is the message's **`expires`**, passed through the entry's `options`: - `"options": {"expires": 50 * 60}` makes each warm-up expire 50 minutes after it is sent. - A worker that receives an expired message does not run it; it logs the discard and marks the task **`REVOKED`** (in the result backend, when results are stored). - Choose a value **below the interval**, so at most one live copy exists at a time. Do not copy that onto billing blindly. A billing run that expires after an outage is a billing run that never happens. For work that must happen, prefer late over never, and make the task itself safe to run late. ## Choosing per entry The right behaviour after an outage is a property of each job, not of beat as a whole: - **Cache warm-up**: only the latest copy is useful. Expire it below the interval, and let a restart's single catch-up send refill the cache. - **Nightly billing**: must happen, ideally once. Keep the default `None` deadline so a late run is still sent, persist the schedule file, and make the task safe to run late and idempotent per billing period. - **Customer reports**: a report arriving six hours late may be worse than none. The built-in lever, `beat_cron_starting_deadline`, is **app-wide**: it applies to every crontab entry, billing included. If billing must still run late, the reports need their own guard instead, such as the report task checking whether its reporting window has already passed and returning early. Write the choice down next to the entry; it is invisible in the schedule itself. ## Summary | Outage | Default outcome | Option that changes it | |---|---|---| | Beat down, file kept | One catch-up send per overdue entry | `beat_cron_starting_deadline` for crontab entries | | Beat down, file lost | Missed runs skipped silently | Persist the file (`-s` / `beat_schedule_filename`) | | Workers down, beat up | Messages accumulate and run in a burst | `expires` in the entry's `options` |
- Why is `beat_cron_starting_deadline` no help for the hourly Celery warm-up declared with `timedelta(hours=1)`?The deadline is checked only in `crontab`'s due calculation. A `timedelta` schedule has no notion of a missed wall-clock slot: it is simply overdue by some amount, and beat sends it once. To stop stale warm-ups, use `expires` on the message instead.
- After workers come back, how can you tell that a Celery warm-up was dropped by `expires` rather than lost?The worker logs `Discarding revoked task`, emits a `task-revoked` event flagged as expired, which Flower shows, and marks the task `REVOKED` in the result backend when results are stored. A run that was never sent leaves no trace at all, which is the silent case a lost `celerybeat-schedule` file produces.
saying these in an interview costs you the question
- Beat replays every missed slot when it restarts
- Beat stops sending while no worker is consuming the queue
- beat_cron_starting_deadline also limits timedelta schedules
- Expired messages still run, just later
- Losing celerybeat-schedule makes beat resend everything at startup