How do the `schedule` and `concurrencyPolicy` fields of a Kubernetes CronJob work, and what do the values `Allow`, `Forbid` and `Replace` each do when the previous run is still going?
answer
- CronJob creates Jobs, never Pods directly
- 5-field cron; timeZone (IANA) stable in 1.27, else UTC
- Allow (default) = overlap; Forbid = skip, not queue; Replace = kill old
- Forbid + hang + no activeDeadlineSeconds = silently stops
- history limits: 3 successful / 1 failed; suspend to pause
basics
~20 sschedule is a five-field cron expression; at each firing the CronJob controller creates a Job. concurrencyPolicy decides what happens if the previous Job is still running: Allow (default) starts another anyway, Forbid skips the new run, Replace kills the running Job and starts the new one.
solid answer
~50 sA CronJob is a template plus a schedule; it creates a **Job** object at each firing and never runs Pods directly, so all Job semantics apply to each run. `schedule` is standard five-field cron (minute, hour, day-of-month, month, day-of-week), with `@hourly`/`@daily` style macros. Since Kubernetes 1.27 you can set `timeZone` (an IANA name such as `Europe/Berlin`); without it the controller uses the kube-controller-manager's own time zone, historically UTC. `concurrencyPolicy` handles overlap: - **`Allow`** (default) — fire regardless; runs may overlap. Fine for idempotent, short work; dangerous for anything that takes a lock or writes the same rows. - **`Forbid`** — if the previous Job is still active, skip this occurrence entirely. It is skipped, not queued — it will not run later. - **`Replace`** — delete the currently running Job (killing its Pods) and create the new one. Only sensible when the freshest result is what matters and partial work is disposable. Other fields: `suspend`, `successfulJobsHistoryLimit` (3), `failedJobsHistoryLimit` (1), `startingDeadlineSeconds`.
code
yaml · 21 linesapiVersion: batch/v1
kind: CronJob
metadata:
name: invoice-sync
spec:
schedule: '*/5 * * * *'
timeZone: 'Europe/Berlin'
concurrencyPolicy: Forbid
startingDeadlineSeconds: 120
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
backoffLimit: 1
activeDeadlineSeconds: 240 # shorter than the interval: no permanent block
template:
spec:
restartPolicy: Never
containers:
- name: sync
image: registry.example.com/invoice-sync:2.0.4go deeper
Know that a CronJob creates Jobs on a cron schedule and name what each of the three concurrency policies does.
Add the defaults (Allow, history limits), the time-zone behaviour and field, and the fact that Forbid skips rather than queues.
Discuss the failure modes: overlap pile-up under Allow, silent stoppage under Forbid with a hung run and no deadline, partial writes under Replace, and how you would alert on missed runs.
Decide whether the workload belongs on a CronJob at all versus a workflow engine or queue consumer, and define platform defaults for time zone, concurrency policy, deadlines and missed-run alerting.
## What a CronJob actually is A CronJob is a *Job factory*. It holds a `jobTemplate` and a schedule; the CronJob controller wakes up periodically (roughly every 10 seconds), compares the current time to `.status.lastScheduleTime`, and for each schedule point that has passed creates a **Job** object. From there, everything is ordinary Job behaviour: `backoffLimit`, `parallelism`, `restartPolicy: Never|OnFailure`, `activeDeadlineSeconds`, `ttlSecondsAfterFinished`. A CronJob adds no execution semantics of its own — it only decides *when* and *whether* a Job is created. That two-level structure explains a lot of behaviour. A failing CronJob is really "the Jobs it created failed"; you debug by listing the Jobs it owns and then the Pods those own. ## The schedule field `.spec.schedule` uses the classic five-field cron syntax: minute, hour, day-of-month, month, day-of-week. `*/15 * * * *` is every fifteen minutes; `0 3 * * 1` is 03:00 on Mondays. Macros `@yearly`, `@monthly`, `@weekly`, `@daily`, `@hourly` are supported. Cron's historical day-of-month/day-of-week OR semantics apply when both are restricted. Time zone is the classic trap. Without `.spec.timeZone`, the schedule is interpreted in the kube-controller-manager's local time zone — in practice UTC on nearly all clusters — so "midnight" in a manifest is rarely local midnight. `.spec.timeZone`, taking an IANA name like `America/New_York`, became stable in Kubernetes 1.27 and should be set explicitly for any business-hours schedule. Note the consequence: a DST-shifted schedule can skip or repeat an hour on transition days, so schedules that must run exactly once per day are safer at a time that is not near the shift. A second timing caveat: the controller polls, so firings can be a few seconds late, and a schedule granularity finer than the poll interval is not achievable. ## concurrencyPolicy The question the field answers is "the previous Job for this CronJob is still active — now what?" **`Allow` (default)** — create the new Job regardless. Runs overlap. This is safe only when the payload is idempotent and re-entrant. The classic incident is a five-minute schedule whose job starts taking seven minutes: overlapping runs pile up, each consuming CPU and database connections, until the cluster or database saturates. Because it is the default, this bites teams that never considered the field at all. **`Forbid`** — if a previous Job is still active, do not create a new one for this occurrence. Critical detail: the occurrence is **skipped, not deferred**. There is no queue; the run simply never happens. If a job routinely overruns its interval you silently lose executions, so pair `Forbid` with an alert on missed runs (for example, on `lastSuccessfulTime` age). **`Replace`** — delete the active Job (which cascades to its Pods, terminating them) and create a fresh one. The new run wins. Appropriate when only the newest result matters and partial work is disposable — refreshing a cache or a materialised view. Wrong for anything that must complete atomically or that leaves partial writes behind, because the killed Pod gets SIGTERM and then, after the grace period, SIGKILL. Note what `concurrencyPolicy` does *not* do: it governs concurrency **within one CronJob only**. Two different CronJobs that touch the same resource can still overlap; that needs an application-level lock. ## The supporting fields - **`suspend: true`** — stop creating new Jobs; existing ones continue. The clean way to pause a schedule during an incident or migration. - **`successfulJobsHistoryLimit`** (default 3) and **`failedJobsHistoryLimit`** (default 1) — how many finished Job objects to retain per outcome, for debugging. Combined with `ttlSecondsAfterFinished` in the job template, whichever prunes first wins. - **`startingDeadlineSeconds`** — how late a missed firing may still be started; it also bounds how far back the controller looks when counting missed schedules. ## Choosing a policy Ask what happens if two copies run at once. If the answer is "nothing bad", `Allow` and move on. If it is "duplicate rows / lock contention / double charges", use `Forbid` and alert on skips. If it is "the older run's output is now worthless", use `Replace`. And in all three cases, set `activeDeadlineSeconds` on the job template so a hung run cannot block every subsequent occurrence indefinitely under `Forbid` — that combination (Forbid plus a hang plus no deadline) is how a nightly job silently stops running for a week. ## Verifying behaviour `kubectl get cronjob` shows `SCHEDULE`, `SUSPEND`, `ACTIVE`, `LAST SCHEDULE`. `kubectl get jobs --selector` or simply listing Jobs whose owner is the CronJob shows the run history. `kubectl create job --from=cronjob/<name> <name>-manual` triggers an out-of-band run using the same template, which is the standard way to test a schedule's payload without waiting.
- With `concurrencyPolicy: Forbid`, is a skipped run executed later?No. The occurrence is dropped entirely — there is no queue and no catch-up for it. If the job regularly overruns its interval you lose executions silently, so pair `Forbid` with an alert on the age of `.status.lastSuccessfulTime` and set `activeDeadlineSeconds` on the job template so a hung run cannot suppress every subsequent occurrence.
- Why can a CronJob with `concurrencyPolicy: Forbid` stop running altogether?If one Job hangs instead of failing and the job template has no `activeDeadlineSeconds`, that Job stays Active forever. Every subsequent firing sees an active Job and is skipped, so the schedule silently stops producing runs while the CronJob object looks healthy. The fix is a deadline on the job template, plus alerting on last successful completion time rather than on Job failures.
- Your CronJob is scheduled for `0 2 * * *` and runs at 03:00 local time. What happened?The schedule is being interpreted in the controller's time zone, almost always UTC, not the local one — so 02:00 UTC lands at 03:00 in a UTC+1 zone. Set `.spec.timeZone` to the IANA zone you mean (stable since 1.27). Also plan around DST: a schedule near the transition can be skipped or repeated on changeover days, so pick a time away from the shift for once-a-day work.
saying these in an interview costs you the question
- Believing a run skipped by `Forbid` is queued and executed later.
- Assuming schedules are interpreted in local time when `.spec.timeZone` is not set.
- Leaving the default `Allow` on a job that takes a lock or writes shared rows.
- Thinking `concurrencyPolicy` prevents overlap between *different* CronJobs.
- Using `Replace` for work that must complete atomically, where a killed run leaves partial writes.