skip to content

A CronJob has produced nothing for two days and no kubectl command ever failed — what happened?

level: seniorimportance: should knowfreq 46%

answer

  1. nothing a human submitted was rejected
  2. walk the chain down to the Pod
  3. look for the absence of pods, not failed ones
  4. check events on the newest Job
  5. audit trail outlives the events

basics

~20 s

Most likely a Pod-level admission rule is rejecting the job's Pods. The CronJob and its Jobs are admitted fine; only the Pod create is denied, by a controller, so the failure lands in a Job event and never in anyone's terminal.

solid answer

~50 s

Work the chain backwards. The CronJob is admitted when applied; on each tick the cronjob controller creates a Job, which is also admitted; the job controller then tries to create a Pod and that create is rejected by a rule matching Pods. Because the requester is a controller, the rejection is recorded as a create-failure event on the Job and retried with backoff — no Pod object is ever created, so `kubectl get pods` shows nothing and there is nothing to describe. Confirm it from events on the most recent Job (short-lived — events expire, typically within an hour), the Job status showing no pods in any state, and the API server audit trail. Then fix the visibility, not just the manifest: evaluate the rule at the workload kinds too so the next one fails at apply time, and alert on denials.

go deeper

for a junior

Know that a rejected create leaves no object behind, so an empty pod list can mean blocked rather than never attempted. Check the parent object's events before assuming nothing happened.

for a middle

Trace the chain out loud — CronJob to Job to Pod — and say which hop was rejected and why the rejection is recorded against the Job rather than returned to a user.

for a senior

Show a real diagnostic order: newest Job's events, then the Job status showing no pods in any state, then the audit trail once events have expired, then reproduce by submitting the Pod yourself.

for a principal

Treat this as a feedback-design failure rather than a manifest bug. Decide what the platform owes a team whose workload a new guardrail stops, and what alerting makes silent enforcement impossible to sustain for days.

## The shape of the failure This is the canonical "blocked but invisible" outcome, and the batch path is where it hides longest. Nothing failed loudly because nothing a human did was rejected. The chain is: CronJob (admitted when someone applied it, possibly weeks ago) → Job, created by the cronjob controller on each schedule tick → Pod, created by the job controller. A rule matching only Pods lets the first two through and rejects the third. The requester on that rejected create is a control-plane controller, so the denial has nowhere to be printed. ## Why nobody noticed - **No Pod exists.** A rejected create produces no object, so `kubectl get pods` — including with `--show-all`-style filters or by label — returns nothing. There is no crash, no `Error`, no `ImagePullBackOff`, nothing to describe. Engineers looking for a failed pod find an empty list and assume the schedule did not fire. - **The Job looks empty rather than failed.** Its status typically shows no active, succeeded or failed pods at all — because "failed" counts pods that ran and terminated badly, and none ever ran. - **The evidence expires.** The create failure is an event on the Job, and clusters expire events after a short retention window, an hour by default in a stock configuration. Two days later the events are long gone. - **The evidence is also garbage-collected.** A CronJob keeps only a small number of finished Jobs by its history limits, so older Jobs — with whatever events they had — are deleted outright. - **The CronJob itself looks healthy.** Its last-schedule time keeps advancing, because scheduling worked; only the Pod create did not. ## How to confirm it quickly 1. Describe the most recent Job and read its events — a create failure quoting the rule's message is the smoking gun while it is still in the retention window. 2. Check the Job status for the tell: no pods in any state, rather than pods that failed. 3. Go to the API server audit trail and search for Pod create requests in that namespace made by the job controller's service account, and their response. This is the durable record when events have expired. 4. Check the admission engine's own denial records or metrics, if you run something that keeps them, filtered by namespace and resource. 5. Reproduce deterministically: create a Pod by hand from the job's pod template. As a human-submitted request, the same rule now returns the denial straight to your terminal with the message — which both confirms the cause and gives you the exact wording to hand to the owner. ## The wider lesson to give the interviewer The manifest fix is trivial. The real finding is that a guardrail was enforcing correctly for two days while the only evidence lived in an object nobody watches and a log that expires. Three durable improvements: - **Evaluate the same intent at the workload kinds as well**, so the next such change is rejected on the developer's `kubectl apply` or in CI, where a human is looking. Keep the Pod-level rule as the authoritative control. - **Alert on admission denials**, broken out by namespace and resource. A denial rate that goes from zero to steady in one namespace is exactly this failure, and it is visible within minutes. - **Alert on the workload, not just the gate.** A scheduled job that has not completed within its expected window should page regardless of the reason — that alert would have caught this on day one, and catches every other reason a batch job stops running too. ## A note on rollout sequencing Any new rule that will be evaluated against controller-created Pods should be watched at the workload level before it is allowed to reject. The failure mode above is not exotic; it is the default outcome of enforcing on an object no human submits, in a namespace whose owners were never told a rule was added.

  • Why does kubectl get pods show nothing at all rather than a failed pod?
    Because the create was rejected at admission, so no Pod object was ever persisted. There is nothing to list, describe or read logs from. Engineers used to runtime failures look for a crashed pod and conclude the schedule never fired, which sends the investigation in the wrong direction entirely.
  • The Job's events have expired. Where is the durable record?
    The API server audit trail. Search for Pod create requests in that namespace made by the job controller's service account and read the response — a rejection at admission is recorded there with the reason. If your admission engine keeps its own denial log or metrics, that is the other durable source, and it is usually easier to query by rule.
  • What single alert would have caught this on day one?
    An alert on the workload's own outcome: a scheduled job that has not completed successfully within its expected window. It is reason-agnostic, so it catches admission denials, quota exhaustion, a broken image and a paused schedule alike. Denial-rate alerting on the gate is a good second layer, not a substitute.
  • How do you reproduce the denial in a form you can show the owner?
    Submit a Pod yourself from the job's pod template. Because you are the requester, the rejection comes straight back on your connection with the rule's message, which confirms the diagnosis and gives you the exact text to send. It is also the fastest way to check that a proposed manifest change actually passes.

A delivery is refused at the loading dock every night. The order system says the order was placed and dispatched; nobody in the office ever hears a complaint, because the only record is a note the driver leaves on a clipboard that gets wiped every hour.

saying these in an interview costs you the question

  • Looks only for failed or crashed pods
  • Concludes the schedule never fired
  • Assumes admission denials always reach the user
  • Fixes the manifest without fixing the visibility
  • Expects Job status to count the rejected pod as failed

context