skip to content

A systemd timer you deployed is not firing when you expected. How do you check what systemd thinks your OnCalendar= expression means, when the timer last and next elapses, and whether the problem is in the timer or in the service it triggers?

level: seniorimportance: should knowfreq 45%

answer

  1. two halves fail for different reasons
  2. normalize the expression first
  3. which column proves it ever ran?
  4. run the unit, not the script
  5. daemon-reload after every edit

basics

~20 s

Normalize the schedule with systemd-analyze calendar, check arming and elapse times with systemctl list-timers --all and systemctl status on the .timer, then separate the halves: run the paired service by hand with systemctl start and read its status and journal.

solid answer

~50 s

Work down the chain. First, is the expression what you meant? `systemd-analyze calendar 'Mon *-*-* 02:00:00'` prints the normalized form and the next elapse, in local time and UTC — which catches both syntax mistakes and time-zone assumptions. Second, is the timer armed? `systemctl list-timers --all` shows `NEXT`, `LEFT`, `LAST`, `PASSED` and the unit each timer `ACTIVATES`; a timer missing from the list is not active, and `systemctl status foo.timer` will say whether it is `active (waiting)`, disabled, masked, or failed to load on a bad `[Timer]` setting. Remember `systemctl daemon-reload` after any edit. Third, if the timer *is* triggering, the failure is downstream: `systemctl status foo.service` shows the last run's result and `journalctl -u foo.service` its output. Run `systemctl start foo.service` by hand to reproduce the job in its real environment. `systemd-analyze verify foo.timer` catches unit-file problems before any of this.

code

bash · 12 lines
bash
# 1. does the schedule mean what I meant?
systemd-analyze calendar 'Mon *-*-* 02:00:00'
timedatectl

# 2. is the timer armed, and what does it activate?
systemctl daemon-reload
systemctl list-timers --all
systemctl status backup.timer

# 3. is the failure in the job itself?
systemctl start backup.service
systemctl status backup.service

go deeper

for a junior

Know the two commands that start the investigation: systemctl list-timers --all to see whether the timer is armed and when it next runs, and systemctl status on the timer unit.

for a middle

Explain what systemd-analyze calendar normalizes and why its UTC line matters, what each list-timers column proves, and why daemon-reload is required after editing a unit.

for a senior

Drive the split between timer and service deliberately, reproduce the job with systemctl start to get the real environment, and recognize randomized delay, catch-up runs and overlapping executions as legitimate causes of an unexpected elapse.

for a principal

Make this diagnosable by default across a fleet: conventions that make scheduled jobs observable, and alerting on failed units so a silently broken schedule is not discovered weeks later.

## Split the problem in two before you touch anything A timer-driven job has two halves that fail for different reasons: the timer may not be arming or not elapsing when you think, or the timer may be firing perfectly and the service may be failing. The first question to answer is which of those you are in, because it halves the search space immediately. `systemctl list-timers` answers it in one line: if the `LAST` column shows a recent elapse, the timer is doing its job and you should stop looking at the schedule. ## Step 1 — does the expression mean what you meant? ```bash systemd-analyze calendar 'Mon *-*-* 02:00:00' systemd-analyze calendar '*-*-* 0/6:00:00' systemd-analyze calendar daily ``` The output gives the original form, systemd's normalized form, and the next elapse — printed both in local time and in UTC. Two classes of bug fall out here. Syntax you got wrong (a mis-ordered field, a repetition you thought meant something else) shows as a parse error or an obviously wrong normalization. And the local/UTC pair exposes the time-zone assumption: calendar expressions are local time unless suffixed with `UTC`, so a host whose zone is not what you assumed will fire at a different absolute moment. `timedatectl` confirms the host's zone and whether the clock is synchronized. `systemd-analyze timespan 15min` does the equivalent sanity check for the monotonic directives and for `RandomizedDelaySec=`. ## Step 2 — is the timer actually armed? ```bash systemctl daemon-reload systemctl list-timers --all systemctl status backup.timer ``` `list-timers` without `--all` shows only active timers; `--all` includes inactive ones, which is how you discover that the timer exists but was never started. The columns are `NEXT`, `LEFT`, `LAST`, `PASSED`, `UNIT` and `ACTIVATES`. Read them in that order: a plausible `NEXT` proves arming and schedule; `LAST`/`PASSED` prove whether it has ever run; `ACTIVATES` proves the pairing resolved to the service you intended rather than to a name that does not exist. `systemctl status backup.timer` adds the state — `active (waiting)` is healthy — and a `Trigger:` line with the next firing. It is also where a bad directive surfaces: an unparseable `[Timer]` setting makes the unit fail to load and status reports the offending line. `systemd-analyze verify /etc/systemd/system/backup.timer` performs that check without deploying. The recurring causes at this step, in rough order of frequency: nobody ran `daemon-reload` after writing the files; the `.service` was enabled instead of the `.timer`; the timer was enabled but never started and the host has not rebooted; the unit is masked; the timer's name does not pair with the service's and no `Unit=` was set. ## Step 3 — the timer fires and the job still does not happen If `LAST` shows recent elapses, hand the investigation to the service half: ```bash systemctl status backup.service journalctl -u backup.service --since today ``` Status reports the outcome of the most recent run — a `oneshot` unit that succeeded shows `inactive (dead)` with a zero exit, while a failure shows `failed` with the exit status or signal. The journal carries whatever the job wrote to stdout and stderr. The highest-value move here is to run the job yourself: ```bash systemctl start backup.service ``` This is not the same as running the script in your shell. It executes the unit with the same `User=`, `WorkingDirectory=`, environment, resource limits and cgroup that the scheduled run gets, which is precisely where the "works when I run it, fails on schedule" class of bug lives — a `PATH` the script assumed, a credential only present in your login session, a relative path resolved from your home directory. ## Step 4 — the subtler cases - **The job runs but takes longer than the interval.** systemd will not start a second copy while the unit is still active, so triggers appear to be skipped. `LAST` advances less often than the schedule suggests, and status shows a long-running activation. - **A randomized delay is in play.** With `RandomizedDelaySec=` set, the elapse legitimately drifts from the nominal time; `systemctl show backup.timer -p RandomizedDelayUSec` confirms what is configured. - **DST or a clock step.** A calendar time inside a daylight-saving transition can be skipped or duplicated. `journalctl` timestamps around the transition make it obvious. - **A persistent timer already caught up.** With `Persistent=true`, the run you were waiting for may have happened right after boot rather than at the nominal hour; the stamp under `/var/lib/systemd/timers` and the `LAST` column agree on when. ## The order matters Expression, then arming, then the service. Each step is one command, each rules out a whole class of cause, and running them in that order stops you from tuning a schedule that was firing correctly all along.

  • systemctl list-timers does not show your timer at all. What does that tell you?
    That the timer unit is not active. Re-run with `--all` to include inactive timers: if it appears there, it exists but was never started — typically because only `enable` was run without `--now`, or because the `.service` was enabled instead. If it is absent even with `--all`, the unit is not loaded at all: a missing `daemon-reload`, a wrong path, or a masked unit.
  • The timer's LAST column shows it fired on schedule, yet the work never happened. Where do you look next?
    At the service half. `systemctl status backup.service` gives the last run's exit status, and `journalctl -u backup.service` gives its output. Then run `systemctl start backup.service` to reproduce the run in the same user, environment and working directory the schedule uses — most "fires but does nothing" cases are a path or credential the script only finds in an interactive shell.
  • Why can a timer legitimately elapse at a different time from the one in its OnCalendar= expression?
    Several reasons, all configured: `RandomizedDelaySec=` adds a random offset; `AccuracySec=` permits firing anywhere in a window that defaults to a minute; `Persistent=true` may have fired a catch-up run right after boot instead of at the nominal hour; and a local-time expression shifts against absolute time across a daylight-saving transition.

saying these in an interview costs you the question

  • Edits unit files and never runs daemon-reload
  • Tests the job by running the script in an interactive shell
  • Assumes OnCalendar= is evaluated in UTC
  • Tunes the schedule without checking whether the timer already fired
  • Reads a missing timer in list-timers as a systemd bug

context