skip to content

Scripted Data Changes

RunPython and RunSQL change data inside a migration using historical models from apps.get_model, never imported ones. Interviewers ask why an imported model breaks old migrations.

part ofDjangooverview, primer and where to startread it →
on this pageshow

explore

questions

5

In Django, how do you write a data migration that fills new first_name and last_name columns from an existing full_name column?

level: juniorimportance: must knowfreq 55%

answer

  1. migrations do not autodetect data
  2. start from an empty migration file
  3. function receiving apps and schema_editor
  4. registered with migrations.RunPython

basics

~20 s

Create an empty migration with makemigrations --empty, write a function taking (apps, schema_editor) that loads the model via apps.get_model and fills the new columns, and add it to operations as migrations.RunPython, ideally with a reverse function.

solid answer

~40 s

Django only autodetects schema changes, so a data change starts from `python manage.py makemigrations crm --empty --name split_full_name`, which creates a file already depending on the app's latest migration. Above the `Migration` class I write a forwards function with the signature `(apps, schema_editor)`: it gets the model with `apps.get_model("crm", "Customer")` instead of importing it, iterates the rows, splits `full_name`, and saves the two new fields in batches with `bulk_update()`. I register it as `migrations.RunPython(split_full_name, join_full_name)`, where the second function undoes the change so the migration can be rolled back. After that, `migrate` runs it in dependency order like any other migration and records it in the `django_migrations` table.

code

bash · 1 line
bash
python manage.py makemigrations crm --empty --name split_full_name

go deeper

for a junior

Recall the recipe: makemigrations --empty, a function with apps and schema_editor, apps.get_model, and RunPython in operations. Mention that a reverse function makes it undoable.

for a middle

Explain why the data step sits in its own migration between the AddField and RemoveField migrations, and why bulk_update or update() beats per-row save() here.

for a senior

Show you think about the target database alias, large-table streaming with iterator(), and keeping the function independent of application code that will change.

for a principal

Frame data migrations as permanent history: they replay on every fresh install, so review them for self-containment as strictly as schema changes.

## What a data migration is A Django **migration** is a Python file in an app's `migrations/` package holding a `Migration` class with two lists: `dependencies` (which migrations must run first) and `operations` (what to do). Most operations are **schema** operations such as `AddField` or `RemoveField`, and `makemigrations` writes them for you by comparing your models with the state recorded in earlier migrations. A **data migration** changes rows rather than structure: filling a new column, merging duplicates, normalising values. `makemigrations` never generates one, because a change to rows leaves the models identical and there is nothing to detect. You write it by hand, and the main tool is the **`RunPython`** operation (with `RunSQL` as the raw-SQL alternative). ## Step by step for the full_name split Suppose a `Customer` model in an app labelled `crm` has a `full_name` field and you want separate `first_name` and `last_name` fields. 1. Add `first_name` and `last_name` to the model and run `makemigrations`; Django writes a schema migration (say `0007`) with two `AddField` operations. 2. Run `python manage.py makemigrations crm --empty --name split_full_name`. Django creates `0008_split_full_name.py` with `dependencies` pointing at `0007` and an empty `operations` list. 3. Write a **forwards function** above the `Migration` class. Django calls it with two arguments: `apps`, a registry of **historical models** as they exist at this point in the migration history, and `schema_editor`, the object that issues DDL and exposes the current connection. 4. Inside it, fetch the model with `apps.get_model("crm", "Customer")`, loop over the rows, split the name, and write the new values back, preferably in batches. 5. Write a **reverse function** that rebuilds `full_name` from the parts, so `migrate crm 0007` can undo the step. 6. Add `migrations.RunPython(forwards, reverse)` to `operations` and run `python manage.py migrate`. 7. Later, once nothing reads `full_name`, remove the field in a third migration. ## How the three migrations line up | Migration | Operation | What it does | |---|---|---| | `0007` | `AddField` x2 | adds the empty `first_name` and `last_name` columns | | `0008` | `RunPython` | copies data out of `full_name` into the new columns | | `0009` | `RemoveField` | drops `full_name` once the application no longer uses it | Keeping the data step in its own migration is the pattern Django's documentation recommends: data migrations are "best written as separate migrations, sitting alongside your schema migrations". It keeps each file reviewable and lets you roll back the data step independently. ## What goes inside the function - Use **`apps.get_model()`**, never `from crm.models import Customer`; the imported class reflects today's code, not the schema at migration `0008`. - Route queries through **`schema_editor.connection.alias`** with `.using(alias)` so the migration writes to the database that `migrate` is targeting, not always `default`. - Prefer **`bulk_update(objs, fields, batch_size=...)`** or a queryset `update()` over calling `save()` per row; historical models do not have your custom `save()` anyway. - Stream large tables with **`.iterator(chunk_size=...)`** rather than loading every row into memory at once. - Keep the function a **module-level function** in the migration file, not a lambda and not a method imported from application code that may change or disappear. ## Common beginner mistakes - Expecting `makemigrations` to notice that data needs changing. - Importing the model from `models.py`, which works today and breaks a fresh install months later. - Omitting the reverse function, which makes the migration impossible to unapply. - Putting the `AddField`, the data copy and the `RemoveField` in one file, so a failure or rollback cannot separate them. - Calling application services or signals from the migration, coupling a frozen historical step to code that will keep evolving. Once written, the data migration is an ordinary node in the migration graph: `showmigrations` lists it, `migrate` applies it exactly once per database, and the `django_migrations` table records that it ran.

  • Why pass schema_editor.connection.alias to .using() inside the function?
    `RunPython` does not re-point model managers at the database being migrated; plain `Customer.objects` goes to the default database. When `migrate --database=reports` runs the migration against another alias, only `.using(schema_editor.connection.alias)` makes the function read and write that database instead of silently changing `default`.
  • The function needs a model from another app; what must the migration declare?
    Add that app's latest migration to `dependencies`. Otherwise the historical registry may not contain the model yet and `apps.get_model()` fails with a `LookupError` about no installed app with that label, typically on a fresh database where ordering is not accidental.

saying these in an interview costs you the question

  • makemigrations will detect that the existing rows need splitting
  • Import Customer from crm.models inside the migration function
  • A data migration is just a management command run by hand
  • Put AddField, the copy and RemoveField in one migration for simplicity
  • The reverse function is optional and has no effect on rollback
open as a page

In a Django data migration, why must RunPython code load models with apps.get_model() instead of importing them from models.py?

level: middleimportance: must knowfreq 58%

basics

~20 s

apps.get_model() returns the historical model rebuilt from the migrations up to that point; an imported model is today's class, so replaying old migrations on a fresh database breaks once its fields no longer match that step's schema.

open as a page

In Django, what happens when you unapply a migration whose RunPython has no reverse_code, and when is RunPython.noop the right reverse?

level: middleimportance: should knowfreq 42%

basics

~20 s

Django raises IrreversibleError before undoing any of that migration's operations, so the rollback stops there. RunPython.noop is right when leaving the forward change in place is harmless on the way back; otherwise write a real reverse function.

open as a page

In a Django migration, when would you use RunSQL instead of RunPython, and what do its reverse_sql and state_operations arguments do?

level: middleimportance: should knowfreq 35%

basics

~10 s

RunSQL runs raw SQL: use it for set-based updates or database features Django does not model. reverse_sql is the SQL run on unapply; state_operations tells Django's project state what schema the SQL created.

open as a page

On PostgreSQL, one Django migration adds first_name and last_name and then backfills 20 million customers with RunPython; why is that risky, and how do Migration.atomic and RunPython's atomic argument help?

level: seniorimportance: should knowfreq 38%

basics

~20 s

On PostgreSQL the whole migration is one transaction, so the table stays locked by the ALTER while the backfill runs and everything rolls back together. Split schema and data; for huge tables set Migration.atomic = False and commit batches yourself.

open as a page