What is a recreate deployment strategy, what does it cost, and when is it still the right choice?
answer
- stop everything, then start
- the gap is the strategy, not a bug
- two versions never coexist
- exclusive lock, single writer
- undo is a second outage
basics
~20 sRecreate stops every old instance before starting any new one, so there is a deliberate downtime window and never two versions running at once. Choose it when mixed versions would corrupt shared state or fight over an exclusive resource.
solid answer
~50 sRecreate is the stop-then-start strategy: drain and terminate the whole old fleet, then start the new version and wait for it to become healthy. The gap between the last old instance stopping and the first new one passing its health check is real, user-visible downtime, and its length is dominated by process startup plus any migration, not by how fast you can type the deploy command. What you buy with that downtime is the one guarantee no other strategy gives you — two versions of the code are never live against the same data at the same time. That matters when a schema change is not backward compatible, when the app takes an exclusive lock or licence that a second process cannot take, or when duplicated background jobs would double-process work. It is also perfectly reasonable for an internal tool nobody is using at 3am. The catch to say out loud: the rollback for a recreate is another recreate, so a bad release costs you two outage windows.
go deeper
Be able to describe the sequence plainly — all old instances stop, then new ones start, so there is a gap when nothing serves — and name one situation where that gap is acceptable.
Explain what the downtime buys: no two code versions ever touch the same data. Give a concrete case, such as an incompatible migration or an exclusive lock, where a rolling update would deadlock or corrupt state.
Show that you have measured the window rather than guessed it, and that you have thought about what happens on the way back up: retry storms, cold caches, and the fact that the rollback is a second outage. Insist migrations stay expand-only.
Own the decision framework: which services are allowed a downtime window at all, what the announced-window and communication process is, and when the correct answer is to invest in making a service tolerate mixed versions rather than accept recurring outages.
## What recreate actually does A recreate deployment (also called stop-and-start, or unkindly, big-bang) takes the whole old version down before it brings any of the new version up. The sequence is: stop routing new work to the fleet, drain or kill what is in flight, terminate every old instance, start the new instances, wait for them to report healthy, put them back behind the load balancer. Between the last old instance stopping and the first new instance passing its health check, nothing is serving. That window is not a flaw to be engineered away — it is the defining property of the strategy, and the whole point is that you are buying something with it. ## What the downtime buys Every other strategy — rolling, canary, and blue-green sharing a datastore — has a period, however short, in which two versions of your code are simultaneously live against the same downstream state. Recreate is the only strategy that guarantees this never happens. That guarantee is worth paying for when: - **The schema change is not backward compatible.** Running old and new code against the same database requires the expand/contract discipline: add the new column, write both, backfill, switch reads, only then drop the old column. That is several releases of extra work. If a change genuinely cannot be split that way, a short stop-migrate-start window is the honest alternative to pretending the mixed window is safe. - **The application holds an exclusive resource.** A leader lock, a licence seat pinned to a host, an embedded database with an exclusive file lock, a scheduler that must not double-fire. Here a rolling update does not merely risk a problem, it deadlocks: the new instance can never become ready because the old one still holds the lock, and the old one is never stopped because the new one is not ready. - **Duplicated work is destructive.** Two versions of a queue consumer or cron worker processing the same messages with different logic is a data-correctness incident, not a performance blip. - **The environment cannot host two copies.** A single VM, an appliance, a hard resource or licence ceiling. - **Nobody is watching.** An internal admin tool, a batch job between runs, a service with a business-hours SLA being deployed at night. Complexity you do not need is complexity you should not buy. ## What it costs Be specific about the cost, because "a bit of downtime" understates it: - **The window is as long as your slowest startup step.** Drain time plus process start plus migration plus cache priming plus the readiness check interval. A service with a 90-second warmup behind a four-minute migration is a six-minute outage, not a moment. - **Clients see errors, not slowness.** Connection refused and 5xx, and every client that failed retries the instant you come back — a thundering herd against a cold fleet. Backoff with jitter on the client side, and enough headroom on the way up, are part of the design. - **Downstream systems notice.** Health-check-driven service registries deregister you; circuit breakers open; upstream partners may alert. - **Scheduled work is missed** during the window and may all fire at once afterwards. ## The undo path The rollback for a recreate is another recreate: stop the new version, start the old one. So a bad release costs two outage windows, and the second happens under pressure. Two consequences follow: 1. **Keep migrations expand-only even here.** If the migration dropped a column, redeploying the previous binary does not bring the data back. A destructive migration turns a rollback into a restore. 2. **Rehearse the window.** Time it in a lower environment so the number you tell stakeholders is measured rather than guessed. ## Making it less painful Serve a static maintenance page from the edge so users see an explanation instead of a connection error. Announce the window. Run long migrations as a separate, resumable step before the switch rather than inside application startup. Shrink startup time — lazy-load what the first request does not need. And if you find yourself doing this on every deploy for an ordinary stateless service, the finding is not that recreate is bad; it is that nothing about that service required it, and a rolling update would remove the window for free.
- If recreate is so blunt, why not always use a rolling update instead?Because a rolling update guarantees a mixed-version window, and some systems cannot tolerate one. If old and new code share a database whose schema changed incompatibly, or contend for a lock only one process may hold, rolling either corrupts data or deadlocks — the new instance never becomes ready, so the rollout stalls forever. Recreate trades an announced outage for that guarantee.
- How would you shorten a recreate window without changing strategy?Attack the longest step. Move database migrations out of application startup into a separate, resumable, expand-only step run before the switch. Cut cold-start cost by lazy-loading anything the first request does not need. Tighten the readiness check interval so healthy instances are picked up promptly rather than after a slow poll. Then measure the window rather than estimating it.
- A recreate deploy of a fast-starting service is described as effectively zero downtime. What would you check?Whether anyone measured it from the client's side. Process start is only one term: drain of in-flight requests, container image pull, readiness-check interval, load-balancer re-registration and DNS or connection re-establishment all add time. Measure error rate and latency at the edge across a deploy; a service that starts in two seconds can still be unreachable for thirty.
saying these in an interview costs you the question
- Recreate is always wrong and only lazy teams use it
- If startup is fast there is no real downtime
- Rollback is free because you can just restart the old version
- Downtime only matters for customer-facing web apps
- A maintenance page means there was no outage