skip to content

Across a fleet of services, how would you decide where the schema-change runner runs, and what may it assume about which versions are serving?

level: principalimportance: nice to knowfreq 36%

answer

  1. count starts, not deploys
  2. autoscaling breaks the embedded runner
  3. the old version is still serving
  4. standardise so on-call is uniform
  5. schema rights should not outlive the change

basics

~20 s

Decide by how instances start and who should own the failure: autoscaling and unattended restarts argue for a one-shot step outside boot. The runner may never assume its own version is live; the previous one is still serving.

solid answer

~50 s

Two questions decide the placement. First, when do instances start? If the platform starts them only on deploys, an embedded runner is cheap and cannot be forgotten; if it autoscales and restarts unattended, the change step runs at arbitrary hours with nobody watching, which argues for a one-shot step before the rollout. Second, who should be paged? An embedded runner makes schema failures look like application failures to the service's on-call; a separate step gives them an owner and one log. I would standardise one answer per organisation rather than let each service choose, because the on-call story is what suffers from variety. On assumptions: the runner runs while the previous version is still serving and the new one is not ready yet, so it can never treat its own code as the live shape — and it should run with schema-change rights the serving instances do not hold.

go deeper

for a junior

Take away one fact: the change is applied while the previous version of the application is still serving requests, so it is never safe to assume the new code is the one running.

for a middle

Be able to list what pushes the runner out of boot: autoscaling, unattended restarts, many instances starting at once, a long apply, and serving credentials that should not be able to alter the schema.

for a senior

Show you would size the platform's startup allowance, decide who is paged, and keep the change set version-locked to the artifact even when the runner no longer starts on the serving path.

for a principal

Argue for one default across the fleet with a named exception path, because uniform on-call behaviour is worth more than a locally optimal choice per service. Say what the placement costs in mixed-version window and in standing privileges.

## The decision is about starts, not about deploys People compare placements by asking where the change belongs in the pipeline. The more useful question at fleet scale is **when do processes start?** An embedded runner does not run once per deploy; it runs once per process start. - A service the platform starts only during a rollout runs the runner a handful of times, all of them while somebody is watching. - A service that autoscales runs it whenever load rises, whenever a machine is drained, whenever a process crashes. The step then executes at 3am, unattended, on a schedule nobody chose. That second case is where an embedded runner stops paying. Not because the apply is dangerous — it is normally a no-op by then — but because it makes an unattended start depend on something that can block, contend for a lock, and refuse to become ready. ## Criteria worth writing down | Criterion | Points toward embedded | Points toward a separate step | |---|---|---| | How instances start | Deploys only | Autoscaling, unattended restarts | | Fleet size at rollout | A few instances | Many starting in parallel | | Length of the apply | Seconds | Minutes, or unpredictable | | Credentials | One account is acceptable | Serving must not hold schema rights | | Who should be paged | The service team, always | The deploy owner | | Environments | One or two, uniform | Many, easy to wire once | | Risk of being skipped | High, ad-hoc environments | Low, pipeline is mandatory | None of these is decisive alone. The pattern is that embedded placement suits small, simple, low-instance-count services and stops suiting anything that autoscales or has a change set long enough to be felt. ## Standardise, then allow exceptions At fleet scale the value of a consistent answer usually exceeds the value of the locally optimal one. An engineer paged at night should not have to know which of thirty services applies its schema from boot. A workable policy has three parts: 1. **One default placement** for services generated from the standard template. 2. **One mechanism** for mutual exclusion, so the failure and its remedy read the same everywhere. 3. **A named exception path** for the service that genuinely needs something else, with the reason recorded. The frequent compromise is to keep the change set inside the artifact — so it is version-locked to the code that needs it and cannot drift — while running it from a one-shot start of that same artifact rather than from the serving path. That keeps the coupling that makes embedded placement attractive and removes the property that makes it expensive. ## What the runner may assume while it runs This is the part candidates miss. Whatever the placement, at the moment the change is applied: - **The previous version is still serving.** Its instances are connected, holding pooled connections, and issuing statements against the objects being changed. - **The new version is not serving yet**, or only a first instance is. The runner therefore cannot assume the code that matches the new schema is the code taking traffic. - **Both may be live simultaneously** for the length of the rollout, and longer if the rollout pauses or reverses. The consequence for placement is direct: the change has to be one that the currently-serving version tolerates, because it will meet it before its own version does. Designing changes to satisfy that constraint is its own discipline and belongs elsewhere; what belongs here is the assumption itself, and the fact that moving the runner earlier in the deploy makes the window in which the old version meets the new schema *longer*, not shorter. ## Identity and blast radius A fleet-level policy should also fix who the runner connects as. Applying a change needs permission to alter the schema; serving requests does not. Embedded placement pushes toward one account that can do both, which means every serving process — and anything that reaches it — has the ability to alter the schema for the lifetime of the deployment. A separate step lets the schema-change identity exist only for the seconds the step runs. The detailed privilege design is a database-security question, but the placement decision determines whether that separation is even available to you. ## How to argue it in an interview State the criteria, pick a default, and be explicit about the exception path. Then show that you know the runner is not operating in a quiet system: it runs against a database that the previous version is actively using, on a schedule the platform may choose rather than the deploy, and with rights that should not outlive the change.

  • Why does moving the runner earlier in the deploy lengthen the window where the old version meets the new schema?
    Because the change lands before any new instance is ready, so the previous version serves against the changed schema for the whole rollout instead of just its tail. That is usually the right trade — it is predictable and observable — but it must be a deliberate choice, not a surprise.
  • When is keeping the embedded runner still the better answer at fleet scale?
    When services are numerous, small and uniform, and the alternative would be per-service pipeline wiring that some teams will get wrong or skip. A mechanism that always runs beats a better mechanism that is sometimes absent, especially in ad-hoc environments created outside the standard pipeline.

saying these in an interview costs you the question

  • Counts deploys instead of process starts
  • Assumes only the new version is live during the change
  • Lets every service choose its own placement
  • Leaves schema-change rights on the serving credentials
  • Thinks a separate step shortens the mixed-version window
  • Has no exception path, so teams work around the policy