When a service applies schema changes during its own boot, what must happen in what order before it serves traffic?
answer
- three steps, one direction
- schema first, then the mapping check
- readiness is the last thing
- a mismatch aborts, never warns
basics
~20 sBoot applies the pending schema changes first, then checks the mapping against the schema that now exists, and only then reports the instance ready. A failure at either step should stop the boot instead of admitting traffic.
solid answer
~50 sA boot that owns its schema change runs three steps in a fixed order. First the runner takes its exclusive lock and applies whatever is missing, in order, then releases it. Second, the layer validates its mapping against the schema that now exists — tables, columns, types, nullability, the key generators it expects — which is a cheap read of catalogue metadata. Third, and only third, the instance reports itself ready so the router starts sending it requests. The order matters because validating first would compare against the pre-change schema, and reporting ready first would put requests on an instance whose tables are half-shaped. Either failure should abort the boot with a non-zero exit and a specific message; an instance that logs a warning and serves anyway turns a deploy-time failure into a scattering of runtime errors on whichever code paths touch the changed tables.
go deeper
Remember the sequence: change the schema, check the code's mapping against it, then let requests in. Serving before those finish is what causes the odd, endpoint-specific errors right after a deploy.
Explain why the order cannot be swapped: validating first compares against the old schema, and becoming ready first exposes users to a half-shaped database. Say what the validation step actually reads and how cheap it is.
Show that you separate liveness from readiness, size the platform's startup allowance against the worst apply, and refuse to soften a mismatch into a warning. Name the failure you are buying: loud and early beats quiet and scattered.
Frame it as where failures should be paid. This ordering concentrates schema risk into a deploy-time abort; the price is slower starts and broader credentials at boot. Decide deliberately whether the fleet can afford that on every restart and scale-up.
## What "applying the change at boot" means A deployable can carry both the ordered set of schema changes it needs and the runner that applies them. When the process starts, the runner opens a connection, works out which changes the target database is missing, applies them, and hands control back to the rest of startup. The alternative placement — a separate step in the delivery pipeline that runs once before any new instance starts — changes who applies the change but not the ordering problem inside a single process, which is what this question is about. Wherever the runner lives, the process that is about to serve requests must not begin serving until the schema underneath it is the one its code was written against. ## The three steps, in order 1. **Apply.** The runner acquires an exclusive lock so that only one process applies anything, determines the pending set, applies it in order, commits, and releases the lock. 2. **Validate.** The data-access layer compares its mapping to the schema that now exists: does every mapped table exist, does every mapped column exist with a compatible type and nullability, do the key-generation structures the mapping expects exist. This is a read of catalogue metadata, typically milliseconds. 3. **Signal readiness.** Only now does the instance tell the router, load balancer or orchestrator that it can take requests. The steps are strictly sequential inside one process. Nothing in the request path — warm caches, background pollers, scheduled work — should start before step 2 finishes either, because those paths hit the same tables. ## Why validation comes after the apply, not before - The mapping describes the schema the change set is *meant to produce*. Run the comparison before the apply and it reports every change that is about to be made as a mismatch, so it can only be a noisy no-op. - Run it after and it answers a genuinely useful question: **is the change set complete with respect to the code in this artifact?** The classic miss is a developer adding a mapped field and forgetting the matching change script. Without the post-apply check that ships fine, and fails months later on the one code path that touches the column. - It is a shape check, not a correctness check. It sees what the catalogue can tell it — presence, type, nullability — and cannot tell you the data is right or that the change did what was intended. ## What each failure should do | Step | Symptom | Correct behaviour | |---|---|---| | Apply fails | A statement errors, or the lock cannot be taken in time | Abort boot, exit non-zero, do not become ready | | Validation fails | Mapping expects something the schema does not have | Abort boot with the specific mismatch named | | Either step skipped | Instance is ready while the schema is mid-change or stale | The failure mode to design out — partial, confusing errors under load | The temptation with a validation failure is to downgrade it to a warning so the deploy is not blocked. That trades one loud, cheap, deterministic failure for many quiet, expensive, request-dependent ones: reads on unchanged tables succeed, so the service looks up while a subset of endpoints throws. ## The readiness signal is a contract "The process is alive" and "the process can serve" are different statements, and the ordering above only pays off if the platform can tell them apart: - Report unready — or bind the serving port late — until steps 1 and 2 pass, so nothing is routed to an instance still applying. - Give the platform a startup allowance longer than the worst realistic apply time; otherwise it kills the process in the middle of the change and you have a half-applied schema plus a restart loop. - Keep the liveness question separate from the readiness question. An instance patiently waiting on the schema lock is alive and should not be killed for it; it is simply not ready. ## What the ordering costs - **Every start pays it**, not just deploys: a restart after a crash, a scale-up at peak, a machine being drained. The apply is usually a no-op by then, but the connect-and-inspect is not free. - **Starts serialise** while one instance holds the lock, so the first wave of a rollout is as slow as the change. - **The runner needs rights the request path does not** — permission to alter the schema — so an embedded runner either widens what the serving credentials can do or needs its own separate connection settings used only during boot. None of these argue against the ordering; they argue about *where the runner should live*. Inside one process, apply, validate, then serve is the sequence that keeps a bad deploy loud and early instead of quiet and late.
- Why can the post-apply mapping check pass and the deploy still be broken?It only compares shapes visible in the catalogue: names, types, nullability, the structures the mapping expects. It cannot see that a change did the wrong thing, that data was not backfilled, or that an index the query plans depend on is missing. It is a cheap guard against a forgotten change script, not evidence the upgrade is correct.
- Should background work and warm-up caches start before or after the schema steps?After. Schedulers, pollers and cache warmers issue the same statements the request path does, so starting them earlier just moves the failure off the request path and into a log nobody watches. Gate every component that touches the database on the same readiness point.
saying these in an interview costs you the question
- Reports the instance ready as soon as the process starts
- Validates the mapping before applying the pending changes
- Downgrades a mapping mismatch to a warning and serves anyway
- Treats the shape check as proof the upgrade is correct
- Lets background jobs run before the schema steps finish
- Gives the platform a startup deadline shorter than the apply