A GitOps repository contains an operator's CustomResourceDefinitions plus custom resources that use them, and the first sync into a fresh cluster fails on those custom resources. Why does that happen, and what are the ways to order components in GitOps?
answer
- not a script, one applied set
- unregistered kind is rejected
- retry converges eventually
- applied is not the same as ready
- edges only for hard prerequisites
basics
~20 sThe controller applies the whole set in one pass, so a custom resource whose CRD is not yet registered is rejected by the API server. GitOps handles this either by retrying until it converges, or by declaring explicit ordering so the CRDs land first.
solid answer
~50 sA GitOps controller does not read your directory as a script; it computes a desired set and applies it, largely in one pass. On a fresh cluster the custom resource is submitted before the API server knows that kind exists, so it is rejected and the sync reports failure. There are three legitimate answers. First, let it converge: the controller retries on the next interval, and by then the CRDs are registered, so the second or third attempt succeeds — cheap, but noisy and slow. Second, split the repo into separate reconciled units and declare an explicit dependency edge, so the dependent unit is not applied until its prerequisite is not just applied but healthy. Third, use in-unit ordering hints where the tool supports phases. The distinction that matters is applied versus ready: a webhook-backed operator needs running pods, not just accepted objects.
go deeper
Be able to say that a GitOps controller applies the set rather than running the files in order, so a custom resource can arrive before its CRD is registered and gets rejected.
Explain the three ordering options — converge through retries, split into units with an explicit dependency edge, or use in-unit ordering hints — and name the applied-versus-ready distinction that decides whether an edge actually helps.
Show the operational judgment: retries are acceptable noise for a rare bootstrap and unacceptable when you create clusters daily. Diagnose an intermittent failure behind a declared edge as gating on acceptance rather than readiness.
Own the layering policy for a fleet — which platform layers exist, which edges are sanctioned, and why deep application-to-application chains are rejected in favour of components that tolerate a missing dependency and retry.
## Why the first sync fails It is tempting to read a GitOps repository top-to-bottom like a shell script, but that is not what happens. The controller renders a desired set of objects and submits them to the API server, mostly concurrently, and then reports what stuck. On a cluster that already has the operator installed, the custom resources apply fine — the kind is registered. On a *fresh* cluster the CRDs and the custom resources are in the same batch, and a resource of an unregistered kind is rejected outright. Hence the classic symptom: works on the old cluster, fails on the new one, succeeds on the retry. The same shape appears anywhere a prerequisite must exist first: - a namespace before the objects inside it; - an operator's CRDs before its custom resources; - a validating or mutating webhook backend actually **running** before objects it intercepts can be created; - a secret-management operator before the workload whose secret it materialises; - a storage class or an ingress controller before things that reference them. ## Three ways to order ### 1. Let it converge (retry) Reconciliation is level-triggered: the controller re-evaluates the whole desired state every interval. Failing objects are simply retried. So ordering is often *emergent* — attempt one registers the CRDs, attempt two creates the custom resources. This is a real answer and frequently the right one for small setups. Its costs are: slow bootstrap (you wait out intervals and backoff), noisy status (the unit is degraded for a while, so "red" stops meaning anything), and misleading alerts during cluster creation. It also relies on the failure being genuinely transient — a resource rejected for a reason that will never resolve retries forever. ### 2. Split into units with explicit dependency edges The robust answer is to stop putting the prerequisite and the dependant in the same reconciled unit. Make the operator (or the platform baseline) its own unit, make the consuming workloads another, and declare that the second depends on the first. Every mainstream GitOps tool has such an edge — the spelling differs per tool — and the important semantics to ask about is what the edge waits for: merely *applied*, or *healthy*. Waiting for applied is not enough for a webhook-backed operator, because the objects exist while the pods are still starting; waiting for health is what actually removes the race. This is also why platform layering shows up in fleet repos: `crds` → `platform` → `apps`, three units, two edges, each layer reused unchanged across clusters. ### 3. Ordering hints inside one unit Some controllers let you annotate objects with a phase or wave so one unit is applied in ordered batches, with the next batch waiting for the previous to settle. This keeps everything in one unit while still handling the CRD-then-CR case. It is lighter than splitting, but it hides ordering inside object metadata where it is easy to miss, and it does not give you the independent status of separate units. ## Applied vs ready — the crux Candidates who have only read about this say "add a dependency and you are done". The follow-up that separates them: *what counts as satisfied?* Kubernetes acknowledges an object long before the thing it describes is usable. A Deployment exists immediately; its pods are ready later. A CRD is registered quickly; the operator that serves its conversion or admission webhook may not be up for another minute. If your ordering primitive only proves the manifests were accepted, you have narrowed the race rather than removed it — which is exactly why health-aware ordering, and readiness signals on the prerequisite, matter. ## How much ordering to declare Ordering constraints are not free. Each one serialises the rollout, lengthens cluster bootstrap, and creates a way for one stuck component to hold everything behind it hostage — a dependency chain of six layers means a broken layer two blocks four layers of unrelated work. The usual advice is to declare edges only for genuine hard prerequisites (CRDs, webhooks, namespaces, cluster-wide platform pieces) and let everything else converge through retries. Deep dependency chains between application components are usually a sign that the applications should tolerate a missing dependency and retry, the way well-behaved controllers already do. ## Debugging checklist When a first sync fails on a fresh cluster: check whether the failing kind is registered at all; check whether the prerequisite unit reports healthy or merely applied; check whether the failure resolves on the next interval (transient ordering) or repeats forever (a real error wearing an ordering costume).
- If retries eventually converge anyway, why bother declaring ordering at all?Because "eventually" is expensive: bootstrap drags out over several intervals, the unit sits degraded so red status loses meaning, and alerting fires on every new cluster. It also assumes the failure is genuinely transient. Explicit edges make bootstrap deterministic and let a real failure stand out, which matters most when you are creating clusters routinely rather than once a year.
- A dependency edge is declared but the dependent component still fails intermittently. What is the likely cause?The edge is satisfied by the prerequisite being applied rather than healthy. The CRDs are registered and the operator's objects exist, but its webhook or controller pods are not serving yet, so requests are rejected in the gap. The fix is health-aware ordering: gate on readiness of the prerequisite, not on the API server having accepted its manifests.
- When are long dependency chains between components a design smell?When they serialise things that could converge independently. Each edge lengthens bootstrap and lets one stuck layer block everything behind it. Hard prerequisites — CRDs, namespaces, webhook backends, cluster-wide platform pieces — deserve edges; application-to-application chains usually mean the app should tolerate a missing dependency and retry, like every well-behaved controller does.
saying these in an interview costs you the question
- Believes the controller applies files in directory or alphabetical order
- Says a dependency edge is satisfied as soon as manifests are accepted
- Adds an ordering edge between every pair of components
- Treats a permanent error as something retries will eventually fix
- Thinks putting CRDs in the same folder guarantees they apply first