How does a failed schema change surface differently when the runner is embedded in application boot versus run as a separate deploy step?
answer
- who notices, and where
- boot runs on every start, not every deploy
- one red stage versus a restart loop
- a separate job still needs the lock
basics
~20 sAn embedded runner turns a bad change into unready, restarting instances and a rollout that stalls until it times out. A separate one-shot step fails once, before any instance starts, as one red stage with one owner.
solid answer
~50 sWith the runner inside the deployable, the change runs on every start, so the failure looks like an application failure: the first instance crashes or never becomes ready, the platform restarts it, and the rollout stalls until its deadline expires. The evidence is in application logs, spread over restarts, and whoever is on call for the service gets paged. Move the runner to a one-shot job or pipeline step and the same change fails once, in one place, before any new instance starts — a red stage with a single log, no partially-upgraded fleet, and a natural owner. The trade is convenience: the embedded runner needs no extra infrastructure and cannot be forgotten, while the separate step must be wired into every environment and still needs mutual exclusion, because retries and overlapping pipelines can run it twice.
go deeper
Know that the change can run from the application's own startup or from a separate step, and that in the first case a bad change looks like an app that will not start.
Contrast the two placements concretely: cost per start versus per deploy, application logs versus one job log, restart loop versus red stage. Note that the separate step still needs mutual exclusion.
Talk about the operational surface: paging, evidence gathering under restart backoff, startup allowances sized for the worst change, and keeping schema-change rights out of the serving credentials.
Treat it as where the organisation wants deploy risk to land. Standardising the placement across services is worth more than the best answer applied inconsistently, and the apply-and-exit artifact is often the compromise that keeps the change set version-locked.
## The same change, two placements The ordered set of schema changes has to be applied by *something* before code that depends on it serves traffic. Two placements dominate: - **Embedded in boot** — the deployable carries the change set and the runner, and applies it during startup, every time it starts. - **A separate step** — a one-shot job or pipeline stage runs the same change set once per deploy, before new instances start; the application itself only reads the schema. The mechanics of applying are identical. What differs is who runs it, when, how often, and what a failure looks like to the people watching. ## What failure looks like | | Embedded in boot | Separate step | |---|---|---| | First sign of trouble | Instance never becomes ready; restart loop | The stage goes red | | Where the evidence is | Application logs, repeated per restart | One job log, one run | | Fleet state | Partly old, partly crash-looping new | Unchanged — nothing new started | | Who is paged | Whoever is on call for the service | Whoever owns the deploy | | How it ends | Rollout deadline expires, platform gives up | Pipeline stops at the stage | The embedded failure is noisier and harder to read because it is *mixed in with normal startup*. An engineer paged for "instances failing to start" has to work out that this is a schema failure and not a configuration or dependency problem, while restart backoff spreads the same error over minutes. The separate step's failure is legible on sight: one stage, one attempt, one message, and nothing new deployed. ## The costs the embedded runner keeps charging 1. **It runs on every start, not every deploy.** Restarts after a crash, scale-up under load, a machine drained at 3am — each new process runs the runner. Usually it finds nothing to do, but it still connects, inspects, and contends for the lock. 2. **It puts schema-change authority in the serving artifact.** Applying a change needs permission to alter the schema; serving requests does not. An embedded runner either widens what the serving credentials may do, or needs a second set of connection settings used only during boot and dropped afterwards. 3. **It couples boot time to the change.** A slow change makes every instance in the first wave slow to start, and the platform's startup allowance has to be sized for the worst case rather than the normal one. 4. **It scales the blast radius with the fleet.** Twenty instances booting means twenty runners contending, nineteen of them waiting. ## What the separate step buys, and what it does not It buys **one place to fail and one owner**. It also lets the change run with its own identity, its own timeout, and its own retry policy, none of which have to fit inside an application's startup budget. What it does *not* buy: - **Freedom from mutual exclusion.** A retried job, two pipelines for two services sharing a database, or a manual re-run can overlap. The runner still needs its exclusive lock. - **Freedom from the boot-time shape check.** Instances should still verify their mapping against the schema they find; that check is what catches a change set that does not match the artifact being deployed. - **Immunity to being skipped.** A separate step can be disabled, mis-targeted at the wrong database, or forgotten in a new environment; an embedded runner travels with the code and cannot be left out. ## Choosing, and the middle ground The embedded runner is a reasonable default while a system is one small deployable with one database and few instances: nothing extra to build, nothing to forget. It stops paying as soon as the fleet autoscales, the change set grows long enough to matter to boot time, or the credentials question becomes real. A common middle ground keeps the runner in the artifact but does not run it on the serving path: the same image starts in a mode whose only job is to apply the change set and exit. The change set stays version-locked to the code that needs it, while the failure surfaces as one job that either succeeded or did not — and the serving instances start with no schema-change rights at all. Whichever placement you pick, the property to preserve is that a bad change is discovered **once, loudly, before users see it** — not inferred from a fleet of restarting processes.
- If the runner moves out of boot, what should the application still check at startup?That the schema it finds matches its mapping. The separate step guarantees something was applied, not that it was the change set this artifact expects — a mis-targeted job, a skipped stage or a stale image all leave a mismatch that a cheap catalogue read at boot catches immediately.
- How can a team keep the change set version-locked to the code without running it on the serving path?Ship the same artifact but start it in an apply-and-exit mode as a one-shot job before the rollout. The change set travels with the code that needs it, the failure is a single job result, and the instances that serve traffic never need permission to alter the schema.
saying these in an interview costs you the question
- Thinks a separate job removes the need for a lock
- Assumes the embedded runner only runs on deploys
- Cannot say who gets paged in each placement
- Drops the boot-time shape check once a job exists
- Ignores that serving credentials then need schema rights