Your Helm pre-upgrade migration Job failed midway and every retry fails on a change that already exists — how do you make it re-runnable?
answer
- The retry is not an exception, it is the path
- Ask what finished, not what exists
- Half-applied steps are the real enemy
- Only one writer may hold the schema
- Long work does not belong inside a gate
basics
~20 sMake the migration decide from recorded state rather than assumption: a ledger table of applied versions, one transaction per step where the engine allows it, and a lock so only one runner migrates at a time. A retry then resumes instead of repeating.
solid answer
~50 sThe retry is the normal path for a Helm hook, so the Job has to be written for it. Use a migration tool that records applied versions in a ledger table and skips what is already there, and keep each step small enough to be a single transaction on an engine with transactional DDL — a step that half-applies is what leaves the ledger disagreeing with the schema. Take an application-level lock at the start of the run, because Helm's `--timeout` ends Helm's wait and not the Pod: a second attempt can put a second migration Pod next to a first that is still writing. Move long backfills out of the hook and into a job that runs after the release, so the hook stays short and the blocking window stays small. Then the recovery is boring: fix, re-run `helm upgrade`, watch it skip what is done.
code
bash · 4 lineshelm upgrade myrelease ./platform-chart --timeout 20m
# first attempt gave up at 5m; check before retrying
kubectl get jobs -l app.kubernetes.io/instance=myrelease
kubectl logs -f job/myrelease-migrate-27go deeper
Know the goal in plain terms: running the migration twice must be safe, because a failed Helm upgrade is simply re-run. Migration tools achieve that by recording which migrations have already been applied.
Explain the mechanics of the ledger and of per-step transactions, and why a step that half-applies leaves the ledger and the schema disagreeing. Be able to say why guard clauses alone are not idempotence.
Demonstrate the operational insight: Helm's timeout ends the wait, not the Pod, so two migration Pods can run at once and a lock is mandatory. Show how you split fast schema change from long backfill and how you rehearse a mid-run kill.
Set the standard others follow: which migration tool, what the lock and timeout conventions are, and where backfills are allowed to run. Own the tradeoff between a strong pre-deploy gate and the deploy latency that a blocking hook adds.
### Why the retry is the normal path Helm gives you no command to resume a release. When a `pre-upgrade` hook fails you fix the cause and run `helm upgrade` again, which re-renders the chart and creates the hook Job from scratch. Every migration shipped as a Helm hook is therefore a program that will, sooner or later, be started against a database it has partly changed. Idempotence is not a nicety here; it is the contract. ### The scenario An internal platform chart owned by the infrastructure team — a 3.4 MB tarball — deploys a document-indexing workload. Revision 27 adds a `content_hash` column and backfills it for 41,208 documents. The migration takes about 11 minutes; the upgrade ran with the default five-minute `--timeout`. Helm gave up at five minutes and reported the release failed. The engineer on call re-ran the upgrade immediately. The second hook Job started while the first Pod was still backfilling, and it failed on `column "content_hash" already exists`. Now the release is failed twice, the column exists, the backfill is partial, and nobody is sure which Pod wrote what. Every defect in that story is a design defect in the Job, not in Helm. ### Decide from recorded state The foundation is a ledger: a table in the target database holding the identifiers of migrations that have completed. Every mainstream migration tool ships one. The run reads the ledger, computes what is outstanding, applies only that, and records each success. Re-running then does nothing rather than failing, because the migration never asks "does this column exist?" — it asks "did step 0031 finish?". Hand-rolled SQL run by a shell script has no ledger, which is why the guard-clause style (`ADD COLUMN IF NOT EXISTS`) shows up. Guards make a single statement re-runnable and they are worth having, but they are not a substitute: they cannot tell a half-finished multi-statement migration from an unstarted one, and they quietly succeed when the column exists with the wrong type. ### Make each step atomic, or make it resumable On an engine with transactional DDL, wrapping a step in a transaction means it either lands entirely or not at all, and the ledger row commits with it. That is the cleanest possible re-run story. Where that is unavailable — an online index build, a change the engine will not do inside a transaction, a data change too large for one transaction — the step has to be resumable instead: batched, with progress recorded, so a second run picks up from the last committed batch. The batch cursor becomes part of the state the migration reads on startup. ### One runner at a time This is the part specific to Helm. `--timeout` bounds how long Helm waits; it does not stop the Job. Kubernetes was never told the release failed, so the Pod continues. A retry launched by a human, or by a pipeline configured to retry, creates a second hook Job whose Pod runs concurrently with the first. So the migration must take an application-level lock — an advisory lock, or a row in the ledger claimed with a conditional update — as its first act, and exit or wait if it cannot get one. Locks must be released on crash, which usually means a session-scoped lock the database drops when the connection dies, not a flag the process is trusted to clear. Two migrations racing on the same schema is the failure this prevents, and it is a genuinely bad one: interleaved DDL, deadlocks, and a ledger that no longer describes the schema. ### Keep the hook short The 11-minute backfill is doing the real damage in this story, because it is inside the gate. While the hook runs, the release is blocked and the old version is serving. The standard fix is to split: the schema change (fast, transactional, additive) stays in the hook; the backfill moves to a separate job started after the release, written to be restartable and to make progress in batches. The application in the meantime tolerates both the filled and unfilled state — which it must anyway, because the old version is still running during the hook. If the migration genuinely must be inside the hook and genuinely takes 11 minutes, then `--timeout` must be raised well above that. Leaving it at five minutes and retrying is how you get two writers. ### Verify it, cheaply The test that matters is not "does the migration work on an empty database". It is: run it against a copy of production, kill it at a random point, run it again, and assert the schema and ledger agree. That is the exact sequence Helm will produce in an incident, and it is the only way to find the step that half-applies. ### Failing loudly beats guessing A migration that cannot tell whether it is safe to proceed should refuse and say so, rather than guessing. A hook that exits non-zero blocks the release, and a blocked release with the old version serving is a far better outcome than a schema in a state nobody can name.
- Why is ADD COLUMN IF NOT EXISTS not enough on its own?A guard makes one statement re-runnable but cannot distinguish a half-finished multi-statement migration from an unstarted one, and it succeeds silently when the column exists with the wrong type. You still need a ledger recording which migrations completed, so the run decides from recorded state rather than from inspection.
- Where would you put a 41,000-row backfill instead of the hook?In a separate batched job that runs after the release, restartable from its last committed batch. The hook keeps only the fast additive schema change. That keeps the blocking window small and means the backfill's own failure does not fail a deploy — the application has to tolerate both states anyway, since the old version runs throughout the hook.
- How do you test that the migration is genuinely re-runnable?Run it against a production-shaped copy, kill it at a random point, and run it again, asserting the schema and the ledger agree. That reproduces exactly what Helm produces in an incident. Testing only against an empty database proves nothing about the half-applied case.
saying these in an interview costs you the question
- Assumes the previous hook Pod stopped when Helm timed out
- Relies on IF NOT EXISTS guards instead of a ledger
- Puts a long backfill inside the blocking hook
- Trusts a process to release its lock on crash
- Retries the upgrade immediately without checking the running Job
- Tests the migration only against an empty database