Argo CD reports a third-party custom resource as Healthy the moment it is created, even while its controller is still provisioning the underlying thing. How do you make Argo CD assess that kind's health properly?
answer
- unknown kinds get a free pass
- teach Argo what ready means
- a script keyed by group and kind
- reads status conditions, returns a verdict
- default to Progressing, not Healthy
basics
~20 sRegister a custom health check for that kind. Argo CD evaluates a Lua script, configured in the argocd-cm ConfigMap under resource.customizations.health.<group>_<kind>, which reads the resource's status and returns Healthy, Progressing, Degraded or Suspended plus a message.
solid answer
~50 sArgo CD only knows how to assess kinds it has a health check for — the built-in Kubernetes kinds plus a library of bundled checks for popular CRDs. Anything else has no check, so it is assumed Healthy as soon as it exists, which is why an Application can go green while a controller is still working. The fix is to register a check yourself: a Lua script keyed by `resource.customizations.health.<group>_<kind>` in the `argocd-cm` ConfigMap. The script receives the live object as `obj`, inspects `obj.status` — usually the `conditions` array the controller writes — and returns a table with `status` set to one of Argo CD's health values and a human-readable `message`. Once registered, the Application's aggregate health reflects the CRD, sync waves and `PostSync` hooks wait for it correctly, and "Healthy" starts meaning something. The alternative to writing one is contributing it upstream, since Argo CD ships checks for many well-known CRDs already.
code
yaml · 24 linesapiVersion: v1
kind: ConfigMap
metadata:
name: argocd-cm
namespace: argocd
data:
resource.customizations.health.example.com_FooDatabase: |
hs = {}
hs.status = "Progressing"
hs.message = "No status reported yet"
if obj.status ~= nil then
if obj.metadata.generation ~= obj.status.observedGeneration then
hs.message = "Controller has not observed the latest spec"
return hs
end
if obj.status.phase == "Ready" then
hs.status = "Healthy"
hs.message = "Database is ready"
elseif obj.status.phase == "Failed" then
hs.status = "Degraded"
hs.message = obj.status.reason
end
end
return hsgo deeper
Know that Argo CD can only judge kinds it has a check for, and that a custom resource with no check simply counts as Healthy once it exists.
Explain where the check is registered and what it does: a Lua script per group and kind in argocd-cm that reads the object's status and returns one of Argo CD's health values with a message.
Show the habits that make one safe in production: default to Progressing, nil-guard every field, respect observedGeneration, and check whether Argo CD already bundles a check before writing your own.
Own the platform consequence: health is the signal your waves, hooks and release gates all key off, so decide who maintains these scripts, how they are tested, and whether the right move is contributing checks upstream rather than accumulating cluster-local Lua.
## Why the default is "Healthy" Argo CD's health engine is a lookup by resource kind. It has built-in logic for core Kubernetes kinds and ships a library of checks for widely used custom resources. For a kind with no registered check, there is nothing to evaluate — Argo CD cannot know what a `FooDatabase` considers success — so it does not block the Application, and the resource counts as Healthy once it exists in the cluster. That default is defensible (the alternative would leave every unknown CRD stuck at Unknown), but it produces a specific failure: your Application reports green while the controller behind the CRD is still creating a cloud database, or has already given up with an error condition nobody is looking at. Any `PostSync` smoke test that fires when the app is Healthy will fire too early, and any `argocd app wait` in a pipeline returns success prematurely. ## Registering a Lua health check Custom health checks live in the `argocd-cm` ConfigMap, keyed by API group and kind with an underscore separator: ```yaml apiVersion: v1 kind: ConfigMap metadata: name: argocd-cm namespace: argocd data: resource.customizations.health.example.com_FooDatabase: | hs = {} hs.status = "Progressing" hs.message = "Waiting for the controller to report a condition" if obj.status ~= nil and obj.status.conditions ~= nil then for i, condition in ipairs(obj.status.conditions) do if condition.type == "Ready" and condition.status == "False" then hs.status = "Degraded" hs.message = condition.message return hs end if condition.type == "Ready" and condition.status == "True" then hs.status = "Healthy" hs.message = condition.message return hs end end end return hs ``` The contract is small and worth memorising: - The live object is available to the script as `obj`, exactly as the API server returns it. - The script returns a table with `status` and `message`. - `status` must be one of Argo CD's health values — `Healthy`, `Progressing`, `Degraded`, `Suspended`. Returning anything else leaves the resource Unknown. - `message` is what shows in the UI and CLI, so put the controller's own error text there; it is the difference between "Degraded" and "Degraded: quota exceeded in region eu-west-1". ## Writing one that behaves Three rules keep custom checks from causing more trouble than they solve. **Default to Progressing, not Healthy.** The script runs constantly, including in the window right after creation when `obj.status` is still nil. Starting from `Progressing` means a resource that has not reported yet correctly blocks downstream waves and hooks; starting from `Healthy` recreates the exact bug you are fixing. **Guard every field access.** Lua indexing a nil value throws, and a script that errors leaves the resource `Unknown` — which is a health assessment failure that is easy to miss because it is neither green nor red. Check `obj.status ~= nil` before reaching into it, every time. **Compare generations where the CRD supports it.** Well-behaved controllers write `status.observedGeneration`. If it lags `metadata.generation`, the status you are reading describes the *previous* spec, and reporting Healthy from it means you approved the old configuration. Return `Progressing` in that case. ## Where else checks come from Before writing one, check whether Argo CD already bundles it — the project maintains health checks for many popular CRDs in its resource-customizations library, and the bundled check is usually better tested than a hand-rolled one. Contributing yours upstream is the maintainable end state. Argo CD also supports the reverse operation, `resource.customizations.ignoreDifferences`, but that is a diffing concern, not a health one; do not confuse the two — one changes whether the Application is Synced, the other whether it is Healthy. ## What this unlocks Once the kind is assessed properly, everything downstream that keys off health starts working: the Application's rolled-up status, sync waves that must wait for a database to exist before the app deploys, `PostSync` hooks that should only run against a genuinely ready system, and CI steps that wait on Application health before declaring a release done. That chain is the reason interviewers ask — a candidate who has written a custom health check has usually hit the "green but not ready" failure in production and understood why it happened.
- Why should a custom health script default to Progressing rather than Healthy?Because it runs from the moment the object exists, before its controller has written any status. Defaulting to Healthy reproduces the original bug — the Application goes green while nothing is ready — and lets sync waves and PostSync hooks proceed too early. Progressing keeps the Application honest until the controller actually reports something, and Degraded is reserved for a reported failure.
- What happens if the Lua script throws an error, for example by indexing a nil status?The resource ends up Unknown rather than crashing anything, which is deceptively bad: it is neither green nor red, so alerting keyed on Degraded misses it entirely and health-based waits behave unpredictably. Guard every field access with a nil check, and test the script against a freshly created object, not just a settled one.
- How is a custom health check different from ignoreDifferences?They act on the two different status axes. A health check decides whether a live resource is working, affecting Healthy/Degraded. ignoreDifferences tells the diffing engine to disregard specific fields — typically ones a controller or webhook mutates — so the Application does not sit permanently OutOfSync. Reaching for one when you need the other is a common mix-up.
saying these in an interview costs you the question
- Thinks Argo CD infers CRD health from the pods it creates
- Assumes an unknown kind shows Unknown rather than Healthy
- Defaults the script to Healthy before any status exists
- Returns arbitrary status strings the engine does not recognise
- Confuses a health check with ignoreDifferences