skip to content

Why does a Job under a Helm chart's templates/ apply after the chart's Deployments?

level: seniorimportance: should knowfreq 40%

answer

  1. Two independent reasons, not one
  2. Look at where the kind sits in the table
  3. Renaming the file changes nothing here
  4. Created is not the same as finished
  5. Waiting gates the release, not each kind

basics

~20 s

Job sits near the end of Helm's hard-coded install-order table, behind Deployment and StatefulSet, and nothing in a chart can move it. Even if it sorted first, Helm submits manifests without waiting, so the Job would still not be finished.

solid answer

~50 s

Two separate facts bite here. First, Helm's install order is a fixed table of kinds, and `Job` sits near its end — after `Deployment`, `HorizontalPodAutoscaler` and `StatefulSet` — so an ordinary Job in `templates/` is submitted last among the workloads regardless of its file name or where it lives in the tree. Second, and more important, the sort is a *submission* order with no readiness barrier between kinds: even a Job that sorted first would only be created before the Deployment, never completed before it. Any waiting Helm does is a gate at the end of the whole release, not a checkpoint between kinds. So a schema migration modelled as a plain Job in `templates/` is racing the app it is supposed to precede. The supported way to sequence it is the `helm.sh/hook` annotation family, which takes the manifest out of the sorted set and runs it at a chosen point.

code

bash · 6 lines
bash
# Job is submitted after the workloads, whatever the file is called
helm template batch ./batch-job | grep '^kind:'
# kind: ConfigMap
# kind: Service
# kind: Deployment
# kind: Job

go deeper

for a junior

Recall the two-part shape of the answer: Job is near the end of Helm's fixed kind order, and Helm does not wait between kinds. Knowing that file names cannot fix it is the practical takeaway.

for a middle

Be able to place Job in the table relative to Deployment and StatefulSet, and explain why the path tie-break inside a kind cannot move it. Say plainly that creation order is not completion order.

for a senior

Diagnose the incident: separate the ordering fact from the no-waiting fact, explain why an occasional green install proves nothing, and state exactly what the sort does and does not guarantee before proposing a fix.

for a principal

Frame the policy: charts should not encode sequencing the format cannot honour. Decide when the ordering belongs in the workload's own startup behaviour rather than in release mechanics, and hold that line across teams.

### The scenario A team ships a batch-job chart whose `templates/` directory contains a queue-worker Deployment, a Service, a ConfigMap, and a Job that applies a database schema change. They expect the Job to go first — it is a prerequisite, after all — and instead they watch the worker pods roll, crash on a missing column for a while, and settle only once the Job's pod finally finishes. Nothing is broken; this is exactly what Helm promises. ### Fact one: where Job sits in the table Helm orders a release by sorting every rendered manifest against a hard-coded table of kinds. Near the end of that table, the workload kinds run `DaemonSet`, `Pod`, `ReplicaSet`, `Deployment`, `HorizontalPodAutoscaler`, `StatefulSet`, `Job`, `CronJob`, followed only by `Ingress` and `APIService`. `Job` is therefore behind every long-running workload kind, not in front of them. That placement is defensible: an ordinary Job in a chart is far more often a periodic or auxiliary task than a prerequisite, and the kinds a workload actually needs — ServiceAccount, ConfigMap, Secret, PVC, RBAC, Service — are all much earlier. But it is not adjustable. There is no flag, values key or `Chart.yaml` field that reorders the table, and file naming does not help: file paths only break ties *within* one kind. Renaming the template to `00-migration.yaml` moves nothing, because the comparison never gets as far as the path when the kinds differ. ### Fact two: order is submission, not completion This is the one that matters, and it is where candidates who know the table still get caught. Helm walks the sorted list and sends each manifest to the API server one after another. It does not wait for the ConfigMap to be observed, for the Service to have endpoints, or for a Job's pod to reach `Succeeded` before sending the next thing. So even if `Job` sorted ahead of `Deployment`, all you would have bought is a Job *object* created a few milliseconds earlier — its pod would still be pulling an image while the worker pods started. Nor does asking Helm to wait fix it. Waiting is a gate applied to the release as a whole once everything has been submitted, not a barrier inserted between kinds. In Helm 4 the wait flag is strategy-valued and, when it is omitted, the release does not wait for workloads at all; in Helm 3 it was a boolean, off by default. In both versions waiting is an end-of-release check, so it can tell you the migration finished but cannot make the worker start afterwards. ### What the sort can and cannot express State the boundary crisply in an interview. The kind sort gives you exactly one guarantee: *the objects the platform's own machinery needs in order to admit a workload — the namespace, the account, the config, the claim, the roles — exist before the workload object is submitted.* It gives you no guarantee about running processes: not that a container has started, not that a Service has endpoints, not that a task has completed. Anything phrased as "X must have finished before Y begins" is outside what a sorted single apply can say. ### The remedies Helm's own answer is the `helm.sh/hook` annotation family. Annotating the migration manifest lifts it out of the sorted set entirely and runs it at a chosen point in the release, with Helm waiting on it before continuing — which is precisely the sequencing the sort cannot express, and why the migration in this chart is modelled as a pre-upgrade hook rather than a plain Job. Outside Helm, the alternatives are the ordinary Kubernetes ones: make the worker tolerate the missing schema and back off, or gate its startup on the migration's completion within the pod spec. Those choices belong to the workload's design rather than to the chart's ordering, and they are usually the more robust answer anyway, because they also survive a pod rescheduled onto a fresh node three hours after the release finished — something no install-time ordering can help with. ### The trap to avoid The failure to watch for is a team that discovers the table, moves the migration into a differently-named file, sees a green install once because the timing happened to work, and concludes it is fixed. It is not: nothing changed, and the race will reappear on a slower cluster or a bigger image.

  • The team renames the migration template to 00-migration.yaml and reports that install order improved. What do you tell them?
    That nothing changed. File paths are only a tie-break between manifests of the same kind; across kinds the comparison never reaches the path. The Job still applies after the Deployment, and any success they saw was timing, which will reverse on a slower cluster or a larger image.
  • If Job sorted ahead of Deployment, would the problem be solved?
    No. Helm submits objects in order without waiting between them, so the Job object would merely be created first — its pod would still be scheduling and pulling an image while the Deployment's pods started. Creation order is not completion order, and no rearrangement of the table would change that.
  • What does Helm's install ordering actually guarantee, stated precisely?
    That the objects a workload references — namespace, ServiceAccount, Secret, ConfigMap, PersistentVolumeClaim, RBAC, Service — are submitted to the API server before the workload manifest is. It guarantees nothing about processes: not that a container started, not that a Service has endpoints, not that a task completed.

saying these in an interview costs you the question

  • Thinks renaming the template file changes apply order
  • Says Jobs always run before Deployments in a chart
  • Believes Helm waits for each object between kinds
  • Assumes the wait flag inserts per-kind barriers
  • Claims the install-order table can be overridden per chart

context