Your Helm chart's new crds/ field works on fresh clusters but not on upgraded ones — why?
answer
- Fresh clusters differ from upgraded ones
- One code path ran on some and not others
- The old schema is still the judge
- A green upgrade can still lose the field
- One step must come before the other
basics
~20 scrds/ is applied only by helm install, so a cluster that already ran the release keeps the old definition and validates new custom resources against it — the added field is pruned or rejected. Apply the definition, then upgrade.
solid answer
~50 sThe split between fresh and existing clusters is the diagnostic signature: `crds/` runs on install and never on upgrade, so every cluster that already had the release is validating your new custom resources against the schema it received the first time. Depending on the old definition, the new field is silently dropped or the object is refused — and a silently dropped field gives you a **green upgrade with the old behaviour**, which is the worse of the two. Confirm by reading the live definition and comparing it with `helm show crds` for the exact chart version you are deploying; `helm get manifest` will show your rendered custom resource carrying the field, which proves Helm did its part. The fix is to make the definition its own ordered step: apply it from the pinned chart version, verify it, then run the upgrade. Keep schema changes additive so the definition can safely run ahead of the release.
code
bash · 13 lines# 1. Helm rendered it: the field is in the stored manifest
helm get manifest digestflow -n digest | grep retryBudget
# 2. the cluster disagrees: read the live definition
kubectl get crd digestschedules.digest.example.com -o yaml | grep -c retryBudget
# 3. what the chart version actually carries
helm show crds digestflow/digestflow --version 2.6.1 | grep -c retryBudget
# 4. definition first, then the release, both pinned
helm show crds digestflow/digestflow --version 2.6.1 | kubectl apply -f -
helm upgrade digestflow digestflow/digestflow --version 2.6.1 \
-n digest -f digestflow-values.yaml --rollback-on-failurego deeper
The takeaway to hold on to: crds/ is applied at install and never at upgrade, so a cluster that already runs the release keeps the definition it first received. Never suggest deleting a CustomResourceDefinition to refresh it.
Be able to prove the diagnosis rather than assert it: check the stored release manifest, read the live custom resource back to see whether the field survived, and compare the cluster's definition with the chart's copy.
Own the rollout: definition first from the pinned chart version, verification that the field is really there, then the release upgrade — and explain why a silently pruned field is more dangerous than a rejected object.
Make it a property of the delivery system rather than a runbook: one pinned version driving both stages, additive-only schema changes as a review rule, and a check that fails the pipeline when a cluster's definition lags the chart.
### The incident A chart called `digestflow` wraps a vendor image for an email-digest builder and drives it from a values-generated config file. Version 2.6.1 adds a `retryBudget` field to the `DigestSchedule` definition in `crds/`, and the chart renders a `DigestSchedule` that sets it from an 11-value override file. The rollout goes to 23 clusters. Six of them are new and behave perfectly. The other 17 report a successful upgrade and keep the old retry behaviour. ### Why the split `crds/` is applied by `helm install` and by nothing else. The six new clusters ran an install, so they received the 2.6.1 definition. The 17 existing ones ran an upgrade, which never touches `crds/`, so they are still on the definition installed months ago. Their API server validates and stores your new custom resource against **that** schema, not the one in the chart you just shipped. What happens next depends on the old definition: a field the schema does not describe is dropped rather than kept, and a field that violates a constraint the schema does describe is refused. The first case is the dangerous one, because Helm sees a successful apply, marks the revision deployed, and nothing anywhere is red. ### Confirming it in about two minutes Three reads settle it: 1. `helm get manifest digestflow -n digest` — does the stored release manifest contain the custom resource with `retryBudget`? If yes, Helm rendered and applied what you expected, and the problem is downstream. 2. Read the live custom resource back. If the field is absent from an object you just applied with it, the schema pruned it. 3. Compare the cluster's definition with the chart's: `helm show crds digestflow/digestflow --version 2.6.1` against the live CustomResourceDefinition. The chart's copy has `retryBudget`; the cluster's does not. The fresh-versus-upgraded split is itself strong evidence — whenever a chart change works on new clusters and vanishes on old ones, look at `crds/` first. ### Fixing the 17 clusters Apply the definition from the exact chart version you are deploying, then upgrade: ```bash CHART=digestflow/digestflow VER=2.6.1 helm show crds "$CHART" --version "$VER" | kubectl apply -f - helm upgrade digestflow "$CHART" --version "$VER" -n digest \ -f digestflow-values.yaml --rollback-on-failure ``` What to avoid, in descending order of damage: - **Deleting the definition and reinstalling.** This destroys every `DigestSchedule` in the cluster, including ones other teams created. It is the single most common wrong answer. - **Uninstalling and reinstalling the release** to get the install path. It works, and it takes an outage plus everything else the release owns. - **Reaching for `--force-replace` or `--force-conflicts`.** Neither has any path to `crds/`; the first replaces release resources, the second overrides field-manager conflicts under server-side apply. - **Applying the definition after the upgrade.** The release is already sitting on pruned objects; you then have to re-apply the custom resources anyway. Definition first, always. ### Making it not happen again **Order the two steps explicitly.** The delivery pipeline gets a CRD stage before the chart stage, both pinned to the same chart version so they can never drift apart. Reading the definition out of the chart with `helm show crds` rather than keeping a hand-maintained copy is what keeps them honest. **Keep schema changes additive.** Adding an optional field means the definition is safe to run ahead of the release: old custom resources still validate, and the new field is simply unused until the chart that sets it arrives. Removals and un-serving a version are not upgrades — they are migrations, and they need their own plan. **Verify rather than assume.** A post-apply check that reads a known field back off the live definition costs nothing and turns a silent prune into a failed pipeline stage. `--dry-run=server` on the upgrade catches outright rejections, but it will not catch a silently pruned field, so it is not sufficient on its own. **Do not trust a rendered diff.** `helm template` leaves `crds/` out of its output unless you pass `--include-crds`, so a CI job that diffs rendered manifests reports "no CRD change" for exactly the change that is about to bite you. ### What to say in the interview Name the mechanism first — `crds/` is install-only — then the signature (fresh works, upgraded does not), then the confirmation reads, then the ordered fix, and finish on additive-only schema changes as the property that makes the ordering safe. Candidates who jump straight to "delete the CRD and reinstall" have told you they have never run this in a shared cluster.
- Why is applying the definition before the chart upgrade safe, but applying it after is not?An additive schema change is forward-compatible: a definition that knows about `retryBudget` still validates every custom resource that omits it, so the cluster is correct in the window between the two steps. The reverse order leaves the release applying objects against a stale schema, which prunes or rejects them, and you have to re-apply the custom resources once the definition catches up.
- How would you catch this skew before the rollout rather than after?Make the CRD stage read the definition out of the pinned chart version and then verify it: after applying, read a known new field back off the live definition and fail the stage if it is missing. `--dry-run=server` on the upgrade catches objects the API server would reject, but a field pruned by a schema that simply does not mention it comes back clean, so a dry run alone is not enough.
- Someone proposes deleting the CustomResourceDefinition and reinstalling the release to get the new schema. What do you say?That deletes every custom resource of that kind across the entire cluster, not just the ones this release created, and there is no undo. It is also unnecessary: applying the new definition over the old one is an ordinary update the API server handles, and the release can then be upgraded in place. Deletion is only ever a deliberate, announced migration with backups, never a rollout step.
saying these in an interview costs you the question
- Reinstalls the release to trigger the crds/ path
- Reaches for --force-replace expecting a CRD update
- Deletes the CRD and loses every custom resource
- Assumes a green upgrade means the field landed
- Applies the new definition after the chart upgrade
- Trusts a rendered diff that omits crds/