In a Django data migration, why must RunPython code load models with apps.get_model() instead of importing them from models.py?
answer
- which version of the model
- the migration's point in history
- fresh install replays everything
- no custom methods or save()
basics
~20 sapps.get_model() returns the historical model rebuilt from the migrations up to that point; an imported model is today's class, so replaying old migrations on a fresh database breaks once its fields no longer match that step's schema.
solid answer
~40 sThe `apps` argument that `RunPython` passes is a registry of **historical models** built from the migration files, so `apps.get_model("crm", "Customer")` has exactly the fields the table has at that step. An imported `Customer` is whatever `models.py` says today. The migration works when written, because both agree, but months later a fresh install or CI database replays it against a table that lacks later columns, or after `full_name` was removed from the model, and it fails with a missing-column error or an unknown field. Historical models also carry no custom methods, no custom `save()`, and only managers marked `use_in_migrations = True`, so any logic the step needs must be written inside the migration itself.
go deeper
Remember the rule: in RunPython always use apps.get_model(app_label, model_name), never an import from models.py.
Explain the project state and historical registry, walk through the fresh-install failure, and list what historical models lack: custom methods, custom save(), non-migration managers.
Show how you keep migrations self-contained in review, handle cross-app dependencies, and repair an old migration that imports models without disturbing applied databases.
Treat migration code as permanent, replayed history: set a team rule that migrations never import application code, and decide when squashing to retire old references pays off.
## Two different Customer classes When `migrate` runs, Django does not use the classes in your `models.py` to understand the schema. It replays the operations in every migration file to build a **project state**: an in-memory description of every model, field and `Meta` option as of each migration. From that state it builds an **app registry of historical models**, real model classes you can query, but frozen at a specific point in history. `RunPython` hands that registry to your function as the `apps` argument. `apps.get_model("crm", "Customer")` therefore returns the `Customer` that existed right after the migration's dependencies ran. `from crm.models import Customer` returns the class as it is in the code checked out today. On the day you write the migration those two are identical, which is exactly why the bug hides. ## How the imported model breaks later Take the `full_name` split, written in `0008` with a direct import: 1. Today, `0008` runs against production. The imported `Customer` has `full_name`, `first_name` and `last_name`, matching the table, so it succeeds. 2. Next quarter, `0009` removes `full_name` and `0012` adds a `phone` column. `models.py` is updated accordingly. 3. A new developer, a CI job or a new region creates an empty database and runs `migrate`, replaying `0001` to `0012` in order. 4. When `0008` runs, the table has no `phone` yet, but the imported model selects it, so the database rejects the query for a column that does not exist. And `customer.full_name` is no longer a field on the imported class at all. The Django documentation warns about exactly this: migrations that import models directly "may work initially but will fail in the future when you try to rerun old migrations". The documented fix, even years later, is to edit the old migration to use `apps.get_model()`. ## What historical models keep and drop Migration files can only serialise declarative information, not arbitrary Python, so historical models are deliberately thin. | Aspect | Historical model from `apps.get_model()` | |---|---| | Fields and relations | yes, as of that migration | | `Meta` options | yes, also versioned | | Custom instance methods and properties | no | | Overridden `save()` or custom `__init__` | no, the base `Model` behaviour runs | | Custom managers | only those with `use_in_migrations = True` | | Methods from concrete base classes | yes, base classes are stored as references | Two references still point at live code and must be kept as long as a migration mentions them: functions used in field options such as `upload_to` or `limit_choices_to`, and **custom field classes**, which migrations import directly. Squashing is the usual way to retire old references. ## Consequences for how you write the function - Copy any business rule the step needs into the migration file; do not call `customer.display_name()` or a service module. - Do not rely on an overridden `save()` to fill derived columns; set them explicitly or use `bulk_update()` and `update()`. - If you need a custom manager method, give that manager `use_in_migrations = True` and accept that it is now frozen into history, or inline the query. - Treat the migration as frozen: its code should still be correct when the model has evolved beyond recognition. ## Models from other apps `apps` only contains what the dependency graph has built so far. If `0008` in `crm` reads a model from `billing`, add `billing`'s latest migration to `dependencies`. Otherwise, on a fresh database where the ordering is no longer accidental, `apps.get_model("billing", "Invoice")` raises `LookupError` saying there is no installed app with that label. ## Why interviewers ask it The question checks whether a candidate has understood that migrations are **replayed history**, not one-off scripts. Anyone can make a data migration pass once; the historical registry is what keeps it passing on the fiftieth fresh install, long after the model it touched has changed shape.
- You find an old migration that imports a model directly and it now fails on fresh installs; may you edit it?Yes. Django's documentation explicitly says it is fine to change such a migration to use the historical models and commit that. Databases that already applied it are unaffected, because applied migrations are not re-run; only fresh replays execute the corrected code.
- How do you make a custom manager available inside RunPython?Set `use_in_migrations = True` on the manager class. Django then serialises the manager into the migration state, so historical models get it. The trade-off is that the manager class must keep existing at that import path for as long as any migration references it.
- Does a post_save receiver registered for crm.Customer fire when RunPython saves a historical Customer?Do not rely on it. The historical class is a different class object from the one a receiver was connected to with `sender=Customer`, so sender-filtered receivers do not match it. Data migrations should do their work explicitly rather than depend on signal side effects.
The historical model is a photograph of the room taken on the day of the migration; importing from models.py is looking at the room today. Instructions written against the photograph still make sense years later, while instructions written against today's furniture fail once someone rearranges it.
saying these in an interview costs you the question
- Importing the model is fine as long as the migration passes today
- apps.get_model returns the same class as importing from models.py
- My overridden save() still runs on objects in a data migration
- Every custom manager is available on historical models
- Once applied in production, a migration never needs to run again