A Kamal deploy has shipped a bad release. What does `kamal rollback` actually do, and in which situations will it not save you?
answer
- reverse cutover, no rebuild
- the old container is still on the host
- only a bounded number are retained
- app container only, never the database
- env and secrets do not roll back
basics
~20 skamal rollback takes a previous version tag, brings that container back up on each host, and switches the proxy to it — no rebuild, no registry round trip, so it is fast. It restores only the app container: not the database, not accessories, and not host state.
solid answer
~50 sYou run `kamal rollback <version>`, where the version is the git revision of a release you already deployed; `kamal app containers` shows which ones the hosts still have. Because that container is usually still present, stopped, from an earlier deploy, the rollback is a health-gated cutover in reverse and takes seconds rather than a full build-and-push cycle. It fails to help in several concrete cases: the old container has already been reaped, because Kamal only retains a bounded number of previous containers; a host that was added since that version never had it; a schema migration ran forward and the old code cannot read the new schema; environment variables or secrets changed and rolling the app back does not roll those back; and anything outside the app container — accessories, proxy configuration, third-party state — is untouched. Rollback is a fast undo of one thing, not of a release.
code
bash · 3 lineskamal app containers
kamal details
kamal rollback 9f3c1a2go deeper
Know that rolling back means naming a previous version and that Kamal switches traffic back to a container it already has, rather than rebuilding anything.
Explain why it is fast — retained containers on the host — and be able to list what a rollback leaves untouched: the database, accessories, environment variables and secrets.
Demonstrate incident judgment: check which versions each host still holds, ask what irreversible change shipped alongside, and know when fixing forward is the safer call than rolling back into a migrated schema.
Own the reversibility policy. Decide how many versions to retain, mandate expand-and-contract migrations, keep config changes out of code releases, and be clear with the business about which classes of change simply cannot be undone.
## What the command does `kamal rollback <version>` takes the version identifier of a release you previously deployed — the git revision that tagged the image — and makes that the running version again. Mechanically it is the ordinary cutover run backwards: on each host, bring up the container for that version, wait for its healthcheck, switch the proxy to it, then drain and stop the container that was serving. There is no build, no push, and usually no pull, because that container is typically still sitting stopped on the host from the earlier deploy. That is the reason rollback is measured in seconds while a fix-forward deploy is measured in minutes, and it is the single most valuable property of the tool during an incident. You need the version string. `kamal app containers` lists what each host actually has, and `kamal details` shows what is currently running — do not guess a SHA under pressure. ## Why it is fast: retained containers Kamal does not delete the previous container the moment it stops serving. It keeps a bounded number of recent ones per host — the `retain_containers` setting — precisely so that rollback is a local operation. That bound is also the first way rollback lets you down: **deploy often enough after a bad release and the good version ages out**. Ship five hotfixes trying to patch a broken release and the version you actually want to return to may no longer be on the host. It can still be pulled from the registry if the tag exists there, but that is a slower path and depends on your registry retention policy not having pruned it. ## The five ways it does not save you 1. **The container is gone.** Aged out by `retain_containers`, or pruned by a housekeeping job, or the image tag expired from the registry. 2. **The host never had it.** A server added to `servers:` after that version shipped has no such container. Mixed fleets after a partially failed deploy have the same shape of problem: some hosts can roll back locally, others cannot. 3. **A migration already ran forward.** This is the big one. Kamal rolls back the *application container*; it does nothing to your database. If the bad release renamed a column or dropped one, the old code now runs against a schema it does not understand, and you have converted a broken feature into a hard outage. The defence is entirely upstream: expand-and-contract migrations, never destructive in the same release as the code that stops using the column. 4. **Configuration moved.** Environment variables and secrets are managed separately from the image. If the bad release also changed an env var or rotated a secret, rolling the container back leaves the *new* configuration in place against the *old* code. 5. **Everything outside the app container.** Accessories — Postgres, Redis — are not versioned with the app and are not touched. Neither is the proxy's own configuration, nor anything you changed on the host by hand, nor any side effect already committed to a third-party system. ## How to talk about it in an interview The strong answer separates two things that candidates habitually merge: **the deploy is reversible, the release is not**. Kamal makes reversing the container trivially cheap, and that genuinely raises how confidently a small team can ship. It does nothing about the irreversible parts — schema changes, consumed messages, emails sent, money moved — and those are the parts that actually decide whether an incident is a two-minute blip or an afternoon. So the operational discipline that surrounds Kamal is the same one that surrounds any rolling release: make migrations backward compatible, keep configuration changes in their own release, and know the version string before you need it. ```bash kamal app containers # which versions each host still has kamal rollback 9f3c1a2 # health-gated cutover back to that version ``` One more practical note: rollback takes the same deploy lock as a deploy, so it will refuse to run while another deploy is in flight — which is a feature, not an obstacle, during an incident with several people at keyboards.
- Why is a rollback usually faster than deploying the previous commit again?Deploying the previous commit rebuilds the image, pushes it, and every host pulls it — minutes. A rollback reuses a container already present on the host, so it is just a health-gated proxy switch — seconds. It also removes the build from the critical path during an incident, which is exactly when your CI queue is least predictable.
- How do you keep a destructive migration from making rollback impossible?Split it across releases: expand first (add the new column, write to both), deploy the code that reads the new shape, then contract in a later release once no running version depends on the old shape. At every point the previous container can still serve against the current schema, which is the only thing that keeps a fast rollback meaningful.
- What should you check before running kamal rollback during an incident?Which versions the hosts actually still hold, with `kamal app containers`, and what is running now, with `kamal details`. Then ask whether anything irreversible shipped with the bad release — a migration, a config or secret change, an accessory change — because those are exactly what the rollback will not undo and what will turn a quick reversal into a worse outage.
saying these in an interview costs you the question
- Assumes rollback also reverts database migrations
- Thinks every previous version stays available forever on the host
- Believes rolling back the container also restores old env vars
- Says rollback rebuilds and redeploys the older commit
- Treats a reversible deploy as a reversible release