skip to content

A Kubernetes CronJob stopped firing after the cluster control plane was unavailable for several hours, and the controller logs mention too many missed start times. Explain the mechanism behind that, and what role `startingDeadlineSeconds` plays.

level: seniorimportance: should knowfreq 38%

answer

  1. Controller reconciles ~every 10s from lastScheduleTime
  2. More than 100 missed points → refuses to schedule, sticky
  3. Log: 'too many missed start times'
  4. startingDeadlineSeconds = lateness bound + lookback bound
  5. Recovery: recreate CronJob; alert on lastSuccessfulTime age

basics

~20 s

The CronJob controller counts every schedule point missed since the last run; if it finds more than 100 it gives up and logs an error instead of guessing. startingDeadlineSeconds limits how far back it looks and how late a missed run may still start — setting it below 100 intervals prevents the lock-up.

solid answer

~50 s

The controller polls roughly every 10 seconds and asks: which schedule points have passed since `.status.lastScheduleTime`? After a multi-hour outage on a frequent schedule that list is huge. Kubernetes refuses to guess — if more than **100** missed start times accumulate, it stops scheduling entirely and logs `Cannot determine if job needs to be started: too many missed start times`. The CronJob then requires manual intervention. `startingDeadlineSeconds` fixes both halves: - It bounds **how late** a missed occurrence may still be started; older ones are abandoned as `missed`. - It bounds **how far back** the controller counts missed schedules, so the 100 limit is never reached. Set it to a small multiple of the interval — for `*/5 * * * *`, something like 120–300 seconds. Then after an outage the controller starts at most the most recent occurrence and resumes normally. Caveat: a value under ~10 seconds risks missing runs entirely, because the poll interval itself can exceed the deadline.

code

yaml · 19 lines
yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: metrics-rollup
spec:
  schedule: '*/5 * * * *'
  timeZone: 'UTC'
  startingDeadlineSeconds: 200      # ~2 intervals; caps lookback and lateness
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      backoffLimit: 1
      activeDeadlineSeconds: 240
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: rollup
              image: registry.example.com/rollup:1.9.2

go deeper

for a junior

Know that a CronJob can stop firing after a long outage and that startingDeadlineSeconds limits how late a missed run may start.

for a middle

Explain the reconcile-from-lastScheduleTime model, the 100-missed-schedules guard, and both effects of startingDeadlineSeconds.

for a senior

Cover diagnosis from status and events, why the state is sticky and how to recover, sensible values per interval, and outcome-based alerting instead of object-state alerting.

for a principal

Decide whether cron-on-Kubernetes has adequate delivery semantics for the workload at all, versus a workflow engine or durable queue, and define fleet-wide defaults plus a heartbeat convention for every scheduled workload.

## How the controller decides to fire The CronJob controller does not hold a timer per object. It reconciles: on each pass (roughly every 10 seconds) it takes the CronJob's `schedule`, its `.status.lastScheduleTime`, and the current time, and enumerates the schedule points that fall between them. For each such point it decides whether to create a Job. That design is what makes the controller stateless and restartable — but it means a gap in reconciliation (an unreachable API server, a crashed or evicted controller-manager, an etcd outage, a suspended CronJob later resumed, a clock jump) leaves a **backlog** of schedule points to consider on the next pass. ## The 100-missed-schedules guard Enumerating that backlog is unbounded work: a `* * * * *` schedule that has been dark for two days is nearly 3,000 points. Rather than iterate indefinitely — or, worse, fire thousands of Jobs in a burst — the controller stops counting once it exceeds **100** missed start times, refuses to schedule, and records an event and log line: `Cannot determine if job needs to be started: too many missed start times`. The crucial operational fact is that this state is **sticky**. The CronJob does not heal itself when the control plane comes back, because `lastScheduleTime` is still far in the past and the backlog is still over 100 on every subsequent pass. `kubectl get cronjob` shows a healthy-looking object with an old `LAST SCHEDULE`, and nothing fires. Teams typically discover it days later. Recovery is manual: either update `.status.lastScheduleTime` to a recent value, or — far more commonly and safely — recreate the CronJob (delete and re-apply), which resets the reference point. Then set `startingDeadlineSeconds` so it cannot recur. ## What startingDeadlineSeconds actually does The field has two effects, and both matter here. **Lateness bound.** If a schedule point was missed, the controller starts a Job for it only if the current time is within `startingDeadlineSeconds` of that point. Otherwise the occurrence is abandoned and counted in `.status` as missed. This is what you want for time-sensitive work: a report due at 02:00 is usually worthless if it starts at 09:00. **Lookback bound.** When counting missed schedules, the controller only considers points inside that window instead of walking all the way back to `lastScheduleTime`. That is what keeps the count below 100 and prevents the lock-up entirely. With the field **unset**, there is no lateness bound and no lookback bound — the controller enumerates everything since the last run, which is precisely the path into the 100-schedule guard. ## Choosing a value Set it to a small multiple of the schedule interval, long enough to survive a brief control-plane hiccup, short enough that a stale run never fires: - every 5 minutes → 120–300 seconds - hourly → 600–1800 seconds - daily → 3600 seconds or so, if a late start is acceptable at all Avoid values below about 10 seconds. The controller's poll interval is around 10 seconds, so a deadline shorter than that can expire between passes and the occurrence is skipped even with a perfectly healthy cluster. ## What happens after recovery With a sane `startingDeadlineSeconds`, an outage produces at most **one** catch-up run — the most recent missed occurrence, if it is still within the deadline. Everything older is abandoned. That is almost always the behaviour you want; the alternative (a thundering herd of backlogged Jobs all starting at once) can take out the database the jobs talk to. If you genuinely need every missed occurrence executed, CronJob is the wrong tool — that is a workflow engine or a durable queue with its own backlog semantics. ## Detecting the failure Because the object looks healthy, alert on **outcome**, not on object state: - age of `.status.lastSuccessfulTime` exceeding a few intervals, - age of `.status.lastScheduleTime`, - the "too many missed start times" event on the CronJob, - a heartbeat written by the job payload itself into a system you already monitor. The last one is the most robust, because it survives the CronJob object being deleted, suspended, or misconfigured, and it also catches jobs that "run" but do nothing useful. ## Related causes of the same symptom The same lock-up appears without any outage if a CronJob is left `suspend: true` for a long time and then resumed, or if the controller-manager's clock jumps backwards or forwards significantly. And a CronJob whose Jobs pile up under `concurrencyPolicy: Forbid` with a hung run stops firing for an entirely different reason — worth distinguishing during diagnosis, since the remedy differs.

  • After the control plane recovers with a sane `startingDeadlineSeconds`, how many catch-up Jobs do you get?
    At most one — the most recent missed occurrence, and only if the current time is still within `startingDeadlineSeconds` of it. Every older occurrence is abandoned as missed. That is deliberate: firing the whole backlog at once would create a thundering herd against whatever the job talks to.
  • Why is a `startingDeadlineSeconds` of 5 a bad idea?
    The CronJob controller reconciles roughly every 10 seconds, so a 5-second deadline can expire between two passes. The occurrence is then treated as too late to start and skipped, even on a completely healthy cluster, producing intermittent missed runs that look random. Keep the value comfortably above the poll interval — tens of seconds at minimum.
  • How would you alert on this failure mode, given the CronJob object still looks healthy?
    Alert on outcomes rather than object state: the age of `.status.lastSuccessfulTime` (or `lastScheduleTime`) exceeding a few schedule intervals, plus the "too many missed start times" event. The most robust signal is a heartbeat written by the job payload itself into your monitoring system, since that survives the CronJob being suspended, deleted or misconfigured, and also catches runs that succeed without doing useful work.

An alarm clock that, after a long power cut, refuses to guess which of the hundred alarms you missed should ring now — and will not ring again until you reset it.

saying these in an interview costs you the question

  • Expecting the CronJob to resume on its own once the control plane is healthy — the state is sticky.
  • Believing every missed occurrence is executed after recovery.
  • Leaving `startingDeadlineSeconds` unset and assuming a bounded catch-up.
  • Setting `startingDeadlineSeconds` below the ~10-second reconcile interval and causing random skips.
  • Monitoring only Job failures, which never fire when no Job is ever created.

context