skip to content

A Helm rollback restores the previous revision's manifests but not the database it migrated — how do you keep releases reversible?

level: principalimportance: should knowfreq 44%

answer

  1. Restoring pods is not restoring the world
  2. Which code must run against which schema
  3. Two steps, separated by a release
  4. The second step is the one nobody does
  5. Some doors only open one way

basics

~20 s

Only ship migrations the previous release's code can still run against: additive changes now, destructive ones in a later release once the old version is gone. Helm restores manifests in seconds; whether that restoration actually works is decided entirely by the schema.

solid answer

~50 s

Treat reversibility as a schema property, not a Helm feature. The rule to enforce is that revision N-1's image must run correctly against revision N's schema — expand first (add the column, dual-write, backfill), contract later (drop the old column) in a release that ships only after the version needing it is gone everywhere. That turns one risky deploy into two or three boring ones, and the cost is real: longer-lived dual paths and contract steps that get forgotten, so the cleanup needs an owner and a deadline. Some changes are genuinely one-way — a destructive data transformation, a vendor migration. Those should be named as one-way in advance, with a different safety story: rehearsed restore, a maintenance window, an explicit no-rollback marker on the release. The failure to avoid is discovering at 02:00 that the rollback command succeeded and the application still cannot start.

go deeper

for a junior

Take away the core fact: rolling a Helm release back restores the old pods, not the old database. If the migration removed something the old code needs, the rollback leaves you broken in a new way.

for a middle

Be able to explain expand/contract concretely — add and dual-write, backfill, switch reads, drop later — and why splitting a rename across releases is what makes each individual release reversible.

for a senior

Show you enforce the N-1 compatibility rule in review and rehearse rollback in staging, and that you can identify which changes are genuinely one-way and give those a restore-and-window plan instead.

for a principal

Own the tradeoff across the estate: what expand/contract costs in deploy latency and cleanup debt, who funds the contract step, whether migrations belong inside the release at all for each service class, and when fix-forward is the declared strategy rather than an accident.

### The claim to interrogate "We can always roll back, it's one Helm command." `helm rollback` re-applies a previous revision's stored manifests and records the result as a new revision — quick, reliable, and entirely about Kubernetes objects. If the release that is being undone also changed a database, the rollback restores the old pods and leaves them pointing at the new schema. Whether that combination works is decided by how the migration was written, weeks earlier, by someone who was not thinking about this moment. So the real design question is not "can we roll back" but "what is the compatibility contract between adjacent releases". ### The contract worth enforcing State it as a rule teams can check in review: **the previous release's code must run correctly against this release's schema.** If that holds, the rollback path is real. If it does not, the release is a one-way door whether or not anyone labelled it one. The technique that satisfies the rule is expand/contract. A change like renaming a column becomes a sequence: add the new column and write to both; backfill; switch reads; and only in a later release, once no running version reads the old column, drop it. Each individual release is reversible, because the version before it still works against the schema in front of it. ### What it costs, honestly Three releases instead of one. Application code that carries both shapes for a while, which is real complexity and real test surface. And a contract step that ships after the interesting work is done, which is exactly the kind of task that never gets prioritised — estates end up carrying columns nobody reads and dual-write paths nobody remembers. If you mandate expand/contract you have to fund the cleanup: an owner, a due date, and a way to see which contractions are outstanding. There is also a latency cost. A team that could have shipped a rename on Tuesday now ships it across three deploys. For a low-traffic internal service that may be poor value, and saying so is part of the judgement. ### Decoupling the migration from the release A Helm `pre-upgrade` hook binds the migration to the deploy, which is what gives you the ordering guarantee — the new code cannot start against the old schema. It also means every deploy is a schema event, and a migration failure is a deploy failure. The alternative is to run migrations as their own step: a separate release, a pipeline stage, or an operator, with the application deploy following once the schema is known good. That buys independent failure domains and lets a schema change be reviewed and timed on its own; it costs you the guarantee, because now nothing structurally prevents someone deploying the code first. Which model you pick can legitimately differ by service class — the shared platform chart the infrastructure team owns and the fast-moving product services need not answer this the same way. ### One-way doors Some changes cannot be made reversible at a sensible price: a destructive transformation, a data-format change with no inverse, an engine migration. The mistake is not having them. The mistake is having them unknowingly. Name them in advance. A one-way release gets a different safety story: a rehearsed restore with a measured restore time, a window, a switched-off traffic path, sometimes a shadow write and a comparison run before the cutover. It should also be visible operationally, so the person paged at 02:00 learns from the release itself that rolling back will not help, rather than from trying it. ### Rehearse the thing you claim to have A rollback path nobody has exercised is a belief, not a capability. The cheap version is a staging environment where you deploy N, then run `helm rollback` to N-1 and assert the old version serves correctly against the new schema. It is a small test and it catches the class of error — a `NOT NULL` column with no default, a dropped column still referenced by the previous image — that otherwise surfaces only in an incident. ### Forward fix versus rollback Mature teams often conclude that for stateful services the primary recovery is fix-forward, and rollback is the exception rather than the default. That is a defensible position, but only if it is a decision rather than a discovery. If fix-forward is the plan, invest where it pays: small releases, fast pipelines, feature flags that decouple activation from deploy, and enough observability to know within minutes that the last release is the problem. ### What to say in an interview The strongest answer distinguishes three layers. Helm restores manifests. The schema decides whether restored code runs. The organisation decides which changes are allowed to break that, and pays for the alternative. A candidate who only says "use expand/contract" has the technique; a candidate who also names its cleanup debt, its latency cost, and the one-way cases it does not cover has the judgement.

  • Give a concrete expand/contract sequence for renaming a column.
    Release 1 adds the new column and writes to both. Release 2 backfills and switches reads to the new column. Release 3, shipped only once no running version reads the old one, drops it. Every adjacent pair is compatible, so any single release can be rolled back without the previous image breaking.
  • When would you accept a migration that is not reversible?
    When making it reversible costs more than the risk it removes — a destructive transformation with no inverse, or an engine migration. The condition is that it is declared one-way in advance and carries a different safety story: a rehearsed and timed restore, a window, and an operational marker so nobody wastes an outage attempting a rollback that cannot work.
  • How do you stop the contract step from being forgotten across dozens of services?
    Give it an owner and a date at the moment the expand ships, and make outstanding contractions visible — a tracked item per pending drop, reviewed like any other backlog of debt. Without that, dual-write paths and unread columns accumulate until nobody can prove which are safe to remove.
  • Does running migrations outside the Helm release make things safer?
    It separates failure domains and lets a schema change be timed and reviewed on its own, but it removes the guarantee a pre-upgrade hook gives you — that new code cannot start against an old schema. You are trading a structural ordering guarantee for process discipline, and that trade is only worth it where the discipline actually exists.

saying these in an interview costs you the question

  • Says helm rollback undoes the migration
  • Writes down migrations and assumes they will work
  • Ships the additive and destructive change in one release
  • Never rehearses a rollback in a staging environment
  • Treats every migration as reversible in principle
  • Mandates expand/contract with no owner for the contract step

context