skip to content

In pytest-django, why does seeding shared rows once per session behave differently from per-test fixtures, and what breaks when transactional tests run?

level: seniorimportance: should knowfreq 28%

answer

  1. outside any test transaction
  2. django_db_blocker to unblock
  3. flush wipes everything
  4. kept databases keep the rows

basics

~20 s

Rows seeded in a session-level django_db_setup override are committed outside any test transaction, so rollbacks never remove them. The first transactional test's table flush deletes them, and with --reuse-db they persist into the next run.

solid answer

~40 s

The `db` fixture is function-scoped, so a session- or module-scoped fixture cannot request it. The documented pattern is to override the session-scoped `django_db_setup` and write inside `with django_db_blocker.unblock():`. Those writes are committed straight into the test database, outside the per-test transactions, so every rollback-based test sees them and none removes them. Two things break that picture. A test with `transaction=True`, `transactional_db` or `live_server` flushes every table afterwards, so the seed data is gone for everything after it; pytest-django runs transactional tests after the others partly for this reason, and those tests must reseed their own data. And with `--reuse-db` the kept database still holds last run's seed rows, so non-idempotent seeding creates duplicates or `IntegrityError`s. Seed idempotently, keep seed data read-only in tests, and prefer function-scoped factories for anything a test mutates.

code

python · 13 lines
python
import pytest

from catalog.models import Currency


@pytest.mark.django_db
def test_eur_is_seeded():
    assert Currency.objects.filter(code="EUR").exists()


@pytest.mark.django_db(transaction=True)
def test_import_commits():
    ...  # afterwards every table is flushed, seed rows included

go deeper

for a junior

Know that data created in a session-level setup is shared by all tests and is not rolled back like per-test data.

for a middle

Explain why db cannot be used from a session fixture and how django_db_setup plus django_db_blocker seed the database.

for a senior

Predict the failure modes: flushes from transactional tests, duplicates under --reuse-db, and per-worker seeding, and design idempotent, read-only seed data.

for a principal

Weigh shared seeded data against per-test factories for a large suite, trading speed for isolation and debuggability.

## Why people seed at session scope Porting a Django suite that used `setUpTestData` or large class-level fixtures to plain pytest functions often slows it down: the `db` fixture rolls back after every test, so any data a test needs must be created again for each test. The natural reaction is to create shared reference data **once per session**. pytest-django supports that, but the data then lives by different rules. ## How it is done The `db` fixture is **function-scoped**, and pytest refuses to let a broader-scoped fixture depend on a narrower one, so a session fixture cannot simply request `db`. Instead, pytest-django documents overriding the session-scoped **`django_db_setup`** and unblocking the database by hand with **`django_db_blocker`**: ```python # conftest.py import pytest from catalog.models import Currency @pytest.fixture(scope="session") def django_db_setup(django_db_setup, django_db_blocker): with django_db_blocker.unblock(): for code in ("EUR", "USD", "GBP"): Currency.objects.get_or_create(code=code) ``` Requesting the original `django_db_setup` in the signature makes sure the test database already exists before the seeding runs. ## The rules seeded data follows 1. **It is committed.** The writes happen in autocommit, outside any test's transaction. 2. **Rollbacks do not remove it.** Every test using the plain `django_db` marker or `db` starts and ends with the seed rows in place. 3. **Changes to it are rolled back, but only in rollback mode.** A default-mode test that edits a seeded `Currency` has its edit undone at the end. 4. **A transactional test destroys it.** Tests with `transaction=True`, `transactional_db` or `live_server` flush every table afterwards, and the seed rows go with them. 5. **A kept database keeps it.** With `--reuse-db`, the database is not destroyed at the end, so the rows are still there on the next run. ## What breaks, and how it shows | Situation | Symptom | |---|---| | A transactional test runs, then a test that expects seed data | lookups for the seeded rows fail after the flush | | Seeding uses `create()` and the run uses `--reuse-db` | `IntegrityError` on unique fields, or duplicated rows | | A transactional test itself expects seed data | missing rows, because the previous transactional test flushed them | | Parallel runs with pytest-xdist | each worker has its own test database, so each worker process runs the seeding itself | pytest-django orders the suite as rollback-based database tests first, then transactional tests, then database-free tests. That ordering protects the rollback-based tests from an earlier flush, but it does nothing for transactional tests that follow one another. ## Making it robust - **Seed idempotently** with `get_or_create()` or `update_or_create()` so a reused database does not duplicate rows. - **Treat seed data as read-only reference data.** Anything a test changes should come from a function-scoped fixture or factory, so tests never depend on each other's edits. - **Reseed for transactional tests.** pytest-django's docs show a function-scoped override of `django_db_setup` for suites that rely on transactional tests, reseeding before each one; alternatively give those tests their own fixtures. - **Use `serialized_rollback=True` sparingly**; it restores contents after a flush but is slow. - **Remember `--create-db`** when seed definitions change, so a stale kept database does not hide the change. ## Where this lands in an interview The strong answer names the mechanism, not just the fix: session seeding bypasses the per-test transaction, so its lifetime is governed by commits, flushes and database reuse rather than by pytest's fixture teardown. Knowing that makes the failure modes above predictable instead of mysterious. ## Alternatives worth weighing - **Function-scoped factories** for everything a test touches: slowest in raw terms, but every test is self-contained and order-independent. - **Data migrations** for genuinely static reference data (currencies, plan tiers): the rows then exist in every test database because migrations run during setup. Note that a transactional test's flush removes them too unless `serialized_rollback=True` is used. - **Keeping a few Django `TestCase` classes** for groups of tests that share heavy data: `setUpTestData` still works there, with its once-per-class transaction, even in a suite that is otherwise plain pytest functions. The right mix is usually factories by default, a small idempotent session seed for read-only lookups, and transactional tests that build their own data.

  • Why can't a module-scoped fixture in pytest-django simply request the db fixture to create shared rows?
    `db` is function-scoped, and pytest rejects a fixture that depends on one with a narrower scope, reporting a scope mismatch. That is by design: `db` wraps a single test in a transaction. Shared rows need an explicit `django_db_blocker.unblock()` in a broader fixture, with the commit semantics that implies.
  • What changes if a seeded pytest-django suite is also run with --reuse-db?
    The test database survives between runs, and so do the committed seed rows. Seeding that uses `create()` then hits unique constraints or doubles the data on the second run, so seeding must be idempotent, and `--create-db` is needed after the seed definitions or the schema change.

saying these in an interview costs you the question

  • Believing session-seeded rows are rolled back after each test
  • Thinking transactional tests leave session seed data in place
  • Requesting db from a session-scoped fixture and expecting it to work
  • Seeding with create() while running with --reuse-db
  • Letting tests modify shared seed rows and depend on the edits