What migration policy would you set for a Django team deploying continuously, so every migration stays safe while the previous release still serves traffic?
answer
- two releases share one schema
- additive now, destructive later
- classify every operation
- review the SQL, not the Python
basics
~20 sRequire every migration to work with both the new and the previous release: additive changes ship with the code, destructive steps a release later, indexes built concurrently, backfills in separate data migrations, and each migration's SQL reviewed with sqlmigrate.
solid answer
~50 sThe invariant is **N-1 compatibility**: after `migrate`, both the new release and the one still serving must work against the schema. I would classify Django operations: `CreateModel`, `AddField` that is nullable or has `db_default`, `AddIndexConcurrently` and `AddConstraintNotValid` ship with the code; `RemoveField`, a `RenameField` that renames a column, type-changing `AlterField`, `DeleteModel` and plain `AddIndex` on large tables need a multi-release plan or an agreed window. Rules: one concern per migration, backfills in their own non-atomic data migrations, removals done state-only first and dropped a release later, reverse paths written or irreversibility stated. Enforcement: `makemigrations --check` in CI, `sqlmigrate` output in the pull request, and a fresh-database replay. The trade-off is velocity: renames and drops take several deploys, so the policy needs an exception path, for example a maintenance window, that is explicit rather than improvised.
go deeper
Recall that during a deploy the old release still runs against the new schema, so some migrations must wait a release.
Classify common Django operations as safe or unsafe with the previous release running, and name the safer alternative for each unsafe one.
Design the pipeline checks: makemigrations --check, sqlmigrate in review, fresh-database replay, and tracked follow-up steps.
Balance safety against delivery speed: decide the exception process, the rollback philosophy, and how the policy adapts to the engine and team size.
## The invariant With continuous deployment, `migrate` runs while the **previous release** is still serving, and for a short time after that both versions run side by side. A policy therefore starts from one rule: **every migration must leave a schema that both the new and the previous release can use.** Everything else is a way to meet or enforce that rule in Django. ## Classifying Django operations | Operation | Safe with N-1 running? | Typical plan | |---|---|---| | `CreateModel`, `AddField(null=True)` | yes | ship with the code | | `AddField` NOT NULL with `db_default` | yes | ship with the code | | `AddField` NOT NULL with only `default` | no: the column default is dropped | use `db_default`, or nullable then tighten | | `AddIndex` on a large table | no: blocks writes during the build | `AddIndexConcurrently`, `atomic = False` | | `RemoveField`, `DeleteModel` | no: old code still selects it | nullable, then state-only removal, then drop a release later | | `RenameField` that renames a column, `RenameModel` | no | `db_column`/`db_table` to keep names, or a multi-release move | | `AlterField` changing type | usually no | new column, backfill, switch, drop | | `AlterField` to `null=False` on a large table | risky: full validation under lock | validate a check constraint first with `AddConstraintNotValid` and `ValidateConstraint`, then tighten | ## Rules worth writing down 1. One concern per migration; never mix schema changes with large data changes. 2. Backfills live in separate data migrations, batched and non-atomic when large, and safe to rerun. 3. Destructive steps land at least one release after the code stopped using the column or table. 4. Removals go through a nullable step and a state-only `SeparateDatabaseAndState` before the physical drop. 5. Index builds and drops on large tables use the concurrent operations. 6. Every migration either has a working reverse or states that it is irreversible and why. 7. Migrations never import application code; data migrations use historical models. ## Enforcement in the pipeline - **CI**: `makemigrations --check` fails when models and migrations disagree or the graph has forked. - **Review**: the pull request includes `sqlmigrate` output for each new migration, so reviewers read SQL such as `DROP DEFAULT` or `RENAME COLUMN`, not only Python. - **Replay**: CI builds a fresh database from all migrations, catching migrations that only work on the author's database. - **Deploy**: `migrate --plan` in the release log shows exactly which operations will run. - Runner placement and lock or statement timeouts on the migration connection belong in the same policy, owned by the deploy tooling. ## Trade-offs to own - **Velocity**: a rename or drop becomes a sequence of deploys and follow-up tickets; someone must track the contract steps so they are not forgotten. - **Exceptions**: some changes are cheaper in a short maintenance window; the policy should say who may approve one. - **Rollback strategy**: writing reverse code has a cost and is rarely exercised; many teams prefer forward fixes and only require reverses for data migrations. - **Engine specifics**: several rules rely on PostgreSQL behaviour, such as concurrent indexes and not-valid constraints; a team on another database needs different tools. - **History size**: multi-release changes add migrations; periodic squashing keeps fresh installs fast. There is no single correct policy; interviewers look for the invariant, a concrete classification of Django operations, and an honest account of what the rules cost.
- How do you make sure the delayed contract step, such as dropping a column, is not forgotten?Create the follow-up migration's ticket when the expand step merges, and tie it to a release. Some teams keep a checklist in the pull request template or a periodic review of fields marked for removal. The point is that the multi-release plan is tracked as work, not remembered.
- When would you accept a plain RenameField or RemoveField anyway?When no old code can run against the new schema: an agreed maintenance window, an internal tool deployed with a full stop, or a table no running code touches. The policy should name who approves that exception, so it is a decision rather than an accident.
saying these in an interview costs you the question
- If migrate succeeds, the deploy is zero-downtime
- Reviewing the Python migration file is enough without the SQL
- Dropping a column in the same release as the code change is fine
- Every migration must always have a working reverse
- The same rules apply unchanged on any database engine