skip to content

Fixtures & Serializers

dumpdata and loaddata move rows as JSON, XML or YAML fixtures, with natural keys replacing raw ids. Interviewers ask why signal receivers must check raw=True and when a fixture beats a data migration.

part ofDjangooverview, primer and where to startread it →
on this pageshow

explore

questions

5

In Django, what do the dumpdata and loaddata management commands do, and where does loaddata look for fixture files?

level: juniorimportance: must knowfreq 55%

answer

  1. rows out, rows back in
  2. app_label or app_label.ModelName arguments
  3. a directory inside each app
  4. one extra setting, then the working directory

basics

~20 s

dumpdata writes database rows out as a fixture (JSON by default; XML, JSONL or YAML on request) and loaddata reads fixtures back in one transaction, searching each installed app's fixtures directory, then FIXTURE_DIRS, then the current directory.

solid answer

~40 s

`manage.py dumpdata` serializes the rows of the apps or models you name (`app_label` or `app_label.ModelName`) into a **fixture**: a list of objects each carrying `model`, `pk` and `fields`. It defaults to JSON on one line; `--format`, `--indent`, `--exclude`, `--natural-foreign` and `-o` shape the output. `manage.py loaddata <label>` finds a file by that label in every installed app's `fixtures/` directory, in the directories listed in `FIXTURE_DIRS`, and in the current directory (an absolute path skips the search), then saves every object inside one `transaction.atomic` block. Rows whose primary key already exists are overwritten, and afterwards the database sequences of the loaded models are reset so the next normal insert does not collide.

code

bash · 3 lines
bash
python manage.py dumpdata geo.Currency geo.Country --indent 2 --natural-foreign -o geo/fixtures/geo/reference.json
python manage.py loaddata geo/reference
python manage.py loaddata --database replica geo/reference.json

go deeper

for a junior

Know that dumpdata exports rows as a fixture and loaddata imports one, that JSON is the default format, and that fixtures live in an app's fixtures directory.

for a middle

Explain the search order (app fixtures directories, FIXTURE_DIRS, current directory), that loading is one transaction, and that existing primary keys are overwritten, not duplicated.

for a senior

Show you have been burned by dumps that include contenttypes or permissions, by a filtering default manager, or by sequences, and know which option prevents each.

for a principal

Be ready to say where fixtures belong in a team's workflow: demo and test data, one-off copies between databases, never the path production reference data takes.

## What a fixture is A **fixture** is a file of serialized model rows that Django knows how to import. Each entry names the model as `app_label.modelname`, gives the primary key as `pk`, and lists the other column values under `fields`. Fixtures can be written by hand, but the usual way to create one is to export existing rows with `dumpdata`. The format is Django's own serialization format (see `django.core.serializers`), not an arbitrary JSON shape, so a fixture produced by any other tool will not load unless it follows that structure. ## dumpdata: exporting rows `manage.py dumpdata` with no arguments dumps every installed app. Positional arguments narrow it to an app (`geo`) or a model (`geo.Currency`). The options worth knowing: | Option | What it changes | |---|---| | `--format` | Output format; the default is `json`. Built-in public formats are `json`, `jsonl`, `xml` and `yaml` (YAML needs PyYAML installed). | | `--indent` | Pretty-prints; without it the output is a single line. | | `-e` / `--exclude` | Leaves out an app or a model; repeat it for several. | | `--natural-foreign`, `--natural-primary` | Write natural keys instead of raw ids where models define them. | | `-a` / `--all` | Uses the base manager, so rows hidden by a custom default manager are dumped too. | | `--pks` | Dumps only the listed primary keys; valid only when one model is named. | | `-o` / `--output` | Writes to a file; a `.gz`, `.bz2`, `.xz` or `.lzma` extension compresses it. | A detail that surprises people: `dumpdata` selects rows through the model's **default manager**. If that manager filters (for example, it hides soft-deleted rows), those rows silently disappear from the dump unless you pass `--all`. ## loaddata: finding the file `manage.py loaddata countries` does not need an extension. Django looks for a file named `countries` with any known serialization format and any supported compression (`gz`, `zip`, `bz2`, `lzma`, `xz`) in: 1. the `fixtures/` subdirectory of every installed app; 2. each directory listed in the `FIXTURE_DIRS` setting (an empty list by default); 3. the current working directory. An absolute path turns the search off. Because the label is matched in every app, two apps shipping `fixtures/initial.json` become ambiguous, which is why the docs recommend **namespacing** fixtures in a subdirectory named after the app (`loaddata geo/countries`). A single `-` reads from standard input, and then `--format` is mandatory. `--app` limits the per-app part of the search to one app, `-i` / `--ignorenonexistent` skips fields and models that no longer exist, and `--database` picks the target alias. ## loaddata: what it does to the database - **One transaction.** All the named fixtures load inside `transaction.atomic` on the target database, so a failure part-way leaves nothing behind. - **Raw saves.** Each object is written with `Model.save_base(raw=True)`: the model's own `save()` override is not called, and `pre_save`/`post_save` receivers are told `raw=True`. - **Integrity checked at the end.** Where the backend supports it, foreign-key checking is suspended while rows go in, so the order of objects inside the file matters less; the loaded tables are checked before the command finishes. - **Sequences reset.** Fixtures usually carry explicit primary keys, so after loading at least one object Django resets the sequences of the loaded models; otherwise the next ordinary insert could reuse an id. - **Overwrite, not merge.** An object whose `pk` already exists is saved over the existing row. Edit a row that came from a fixture, run `loaddata` again, and your edit is gone. ## Mistakes interviewers probe - Dumping `contenttypes` and `auth.permission` with raw ids and loading them into another database, where `migrate` already created those rows with different ids. Exclude them, or use `--natural-foreign`. - Expecting fixtures to load on their own. Only a test case's `fixtures` attribute loads them automatically; everywhere else someone runs `loaddata`. - Keeping an old dump after models change: `loaddata` deserializes against the **current** models, so a renamed field breaks it. ## Copying rows between databases The two commands compose. Because `loaddata` accepts `-` for standard input, you can stream a model from one configured database alias to another without an intermediate file: dump from the `test` alias with `--format=json` and pipe into `loaddata --format=json --database=prod -`. For anything beyond a handful of lookup rows, add `--natural-foreign` so references survive differing ids, and consider the `jsonl` format, which writes one object per line and suits large dumps. Remember the limits: every object is saved individually, so this is a convenience for modest volumes, not a bulk-migration tool, and the target must already have the schema, because `loaddata` never creates tables.

  • Why can a fixture dumped from one Django database fail to load into a freshly migrated one?
    Rows created by `migrate` itself, such as `ContentType` and `Permission`, get ids that differ between databases. A fixture that references them by raw id, or that re-inserts them, collides with or points at the wrong rows. Exclude `contenttypes` and `auth.permission` from the dump, or dump with `--natural-foreign` so references are written as natural keys and resolved on load.
  • Running loaddata twice on the same Django fixture: does it duplicate rows?
    No, not when the fixture carries primary keys. Each object is saved with its `pk`, so the second run updates the existing rows in place and overwrites any manual edits. Objects without a `pk` behave differently: they are inserted as new rows unless the model supports natural keys, in which case `get_by_natural_key()` finds the existing row first.

saying these in an interview costs you the question

  • loaddata inserts new rows every run, so running it twice doubles the data
  • loaddata only looks in the project root, never inside the apps
  • dumpdata always exports every row, whatever the default manager filters
  • A fixture can be any JSON document that happens to match the column names
  • Fixtures are applied automatically by migrate, like migrations
open as a page

In Django serialization, what are natural_key() and get_by_natural_key(), and why would a fixture use them instead of primary keys?

level: middleimportance: should knowfreq 42%

basics

~20 s

natural_key() on a model returns a tuple of unique business values; get_by_natural_key() on its default manager finds the row from that tuple. Fixtures then reference rows by stable values like an ISO code instead of database-specific ids.

open as a page

When Django's loaddata saves fixture objects, why should post_save receivers check raw=True, and what else does a raw save skip?

level: middleimportance: should knowfreq 38%

basics

~20 s

loaddata saves each object with raw=True: the model's save() override is bypassed and pre_save/post_save receivers get raw=True, because related rows may not be loaded yet and derived rows are usually already in the fixture. Receivers should return early.

open as a page

Your Django project must ship country and currency reference rows to every environment, production included: should they come from a fixture or a data migration, and why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Use a data migration: migrate applies it automatically, once, in order, against the historical model state, including test databases. A fixture only loads when someone runs loaddata, uses today's models, and overwrites rows by primary key.

open as a page

What does Django's django.core.serializers module provide, and how does it differ from a Django REST Framework serializer?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

django.core.serializers turns model instances into Django's fixture format (JSON, JSONL, XML, YAML) and back, powering dumpdata and loaddata. It has a fixed model/pk/fields shape and no validation, unlike DRF serializers, which define API representations and validate input.

open as a page