On a fresh cluster, a Flux Kustomization that deploys an app fails because the operator supplying its CRDs is not installed yet. How do dependsOn and healthChecks on the Kustomization resource fix the ordering?
answer
- ordering is a graph, not a file order
- one Kustomization waits on another
- applied is not the same as running
- name the objects that must be healthy
- timeout bounds the wait
basics
~20 sSplit the work into two Kustomizations and give the app one a dependsOn pointing at the infrastructure one. Flux then waits until the dependency reports Ready, and healthChecks or wait: true make Ready mean the named objects actually became healthy rather than merely applied.
solid answer
~50 sFlux does not order objects across Kustomizations by itself, so you express the order as a dependency graph. Put the operator and its CRDs in one `Kustomization` — say `infra-controllers` — and the app in another, and give the app `dependsOn: [{name: infra-controllers}]`. kustomize-controller will not apply the app until the dependency's `Ready` condition is true, and while it waits the app's Ready condition reports that it is blocked on the dependency. On its own, though, "Ready" for the infra Kustomization means only that the apply succeeded — the operator's Pod may still be starting, so the app can still race. That is what `healthChecks` fixes: list the specific objects (apiVersion, kind, name, namespace) that must reach a healthy state before the Kustomization is Ready, or set `wait: true` to require readiness for everything it applied, with `timeout` bounding the wait.
code
yaml · 33 linesapiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: infra-controllers
namespace: flux-system
spec:
interval: 30m
path: ./infrastructure/controllers
prune: true
timeout: 5m
sourceRef:
kind: GitRepository
name: fleet
healthChecks:
- apiVersion: apps/v1
kind: Deployment
name: cert-manager-webhook
namespace: cert-manager
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: apps
namespace: flux-system
spec:
interval: 10m
path: ./apps/production
prune: true
sourceRef:
kind: GitRepository
name: fleet
dependsOn:
- name: infra-controllersgo deeper
Know that Flux can order one Kustomization after another with dependsOn, and that the usual layout is infrastructure first, applications second.
Explain that dependsOn gates on the dependency's Ready condition, and that Ready by default only means the apply succeeded unless healthChecks or wait: true is set.
Show the failure you are preventing: an operator applied but not yet serving, so the app races and fails. Choose named healthChecks over blanket wait for speed, and set timeout so a stuck layer fails loudly instead of hanging.
Own the layering itself — how many Kustomizations, where the tenant boundaries sit, and the convergence-time cost of deep dependency chains on a cold cluster, balanced against the reproducibility guarantee that unattended bootstrap gives you.
## Why the ordering problem exists at all Kubernetes has no general ordering primitive. Apply a CustomResource before its CustomResourceDefinition exists and the API server rejects it as an unknown kind. Apply a workload before the operator that mutates it is running and it starts unconfigured. A single `kubectl apply -f` of a directory hits this constantly, and Flux inherits it: within one `Kustomization`, kustomize-controller applies in a sensible built-in order (namespaces and CRDs before things that use them) and retries the reconciliation on failure, but it cannot know that your app needs *this* operator's controller Pod to be running first. The Flux answer is to make ordering a property of the resources you already have, rather than an annotation buried in manifests. ## dependsOn: order between Kustomizations ```yaml apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: apps namespace: flux-system spec: interval: 10m path: ./apps/production prune: true sourceRef: kind: GitRepository name: fleet dependsOn: - name: infra-controllers ``` `dependsOn` is a list of references to other Kustomization objects (name, and namespace if it differs). Before reconciling, kustomize-controller checks each dependency's `Ready` condition. If any is not ready, it does not apply — it requeues and reports that it is waiting on the dependency, so the reason is visible in `flux get kustomizations` rather than hidden in a crash loop. This is what makes the common three-tier layout work: `infra-controllers` (cert-manager, ingress controller, CRDs) → `infra-configs` (ClusterIssuers, the custom resources those CRDs define) → `apps`. Each layer names the one below. `HelmRelease` has the same field, so a chart can wait for another release. Two properties worth stating in an interview. First, a dependency is a gate, not a trigger: the dependent still reconciles on its own interval, it just refuses to proceed while blocked. Second, the graph must be acyclic — a cycle leaves both objects waiting forever, and the fix is to look at the reported reason rather than at logs. ## healthChecks: making "Ready" mean something Here is the subtlety that separates people who have run Flux from people who have read about it. By default, a Kustomization becomes Ready when the apply succeeded. Applied is not running. The operator's Deployment exists but has zero available replicas; its webhook is not serving yet; the CRD is registered but the controller that reconciles those custom resources is still pulling its image. The dependent Kustomization sees Ready, applies the app, and the app fails — the same failure you were trying to avoid, one step later. `healthChecks` changes the definition of Ready for that Kustomization: ```yaml spec: interval: 10m path: ./infrastructure/controllers prune: true wait: false timeout: 5m healthChecks: - apiVersion: apps/v1 kind: Deployment name: cert-manager-webhook namespace: cert-manager sourceRef: kind: GitRepository name: fleet ``` Each entry names an object by apiVersion, kind, name and namespace. The Kustomization is not Ready until those objects are assessed healthy — for a Deployment, that means its rollout has completed. `timeout` bounds the wait; exceeding it fails the reconciliation with a health-check failure rather than hanging. `wait: true` is the broad-brush version: require readiness of *everything* the Kustomization applied. It is convenient and it is honest about the whole layer, but it is slower, it makes one perpetually-unready object block the layer, and it needs a realistic `timeout`. Named `healthChecks` is the surgical version — you list the handful of objects the next layer genuinely depends on. Note that the two are not mutually exclusive in intent: `wait: true` supersedes the need for an explicit list, so pick one deliberately. ## What this buys you operationally - **A fresh cluster converges unattended.** Bootstrap a brand-new cluster and the layers come up in order without anyone running things by hand. That is the real test — a GitOps setup that only works because the cluster already has the operators installed is not reproducible. - **Failures point at a cause.** "apps: dependency 'infra-controllers' is not ready" is a diagnosis. A CrashLoopBackOff on the app is a symptom. - **Retries are free.** Because reconciliation is level-triggered and interval-driven, a dependency that becomes ready five minutes later just unblocks the next tick. There is nothing to re-run. ## Limits to acknowledge dependsOn orders *Kustomizations*, not individual objects, so the granularity of your ordering is the granularity of your layer split — over-splitting to express fine ordering produces a brittle dependency graph. Health assessment covers the object kinds Flux knows how to assess plus custom resources that report standard readiness conditions; an operator whose CR never reports a condition cannot be waited on meaningfully. And long chains multiply worst-case convergence time on a cold cluster, since each layer's timeout can be spent in sequence.
- What is the difference between setting wait: true and listing explicit healthChecks?`wait: true` requires every object the Kustomization applied to become ready, so one perpetually-unready resource blocks the whole layer and convergence is slower. Explicit `healthChecks` names only the objects the next layer actually needs — usually the operator Deployment or webhook — which is faster and fails for a reason you can act on. Both need a realistic `timeout`.
- Two Kustomizations each list the other in dependsOn. What happens?Neither ever reconciles: each requeues reporting that it is waiting on a dependency that will never become Ready. Flux does not break the cycle for you. You see it immediately in `flux get kustomizations`, where both show a dependency-not-ready reason, and the fix is to restructure the layers so the graph is acyclic.
- Do you still need dependsOn for CRDs and the custom resources that use them inside one Kustomization?Not usually. Within a single Kustomization, kustomize-controller applies in a built-in order that puts CRDs and namespaces ahead of dependent objects, and it retries on failure, so a transient unknown-kind error resolves itself. dependsOn earns its place when the wait is on something *running* — an operator or webhook — rather than merely registered.
saying these in an interview costs you the question
- Thinks Flux applies manifests in alphabetical or file order
- Assumes Ready means the workloads are actually running
- Believes dependsOn can order individual objects, not Kustomizations
- Sets wait: true everywhere without a realistic timeout
- Expects Flux to detect and break a dependency cycle