skip to content

Create, Update & Delete Calls

Writing rows with create(), get_or_create, update_or_create, update() with F(), bulk_create, bulk_update and delete(). Interviewers probe races and which calls skip save() and signals.

part ofDjangooverview, primer and where to startread it →
on this pageshow

explore

questions

5

In Django's ORM, how do get_or_create() and update_or_create() differ, and how does each behave when two requests race to create the same row?

level: middleimportance: must knowfreq 62%

answer

  1. both return a two-item tuple
  2. defaults are not part of the lookup
  3. atomic insert, catch IntegrityError, get again
  4. only as safe as the unique constraint
  5. create_defaults since 5.0

basics

~20 s

get_or_create() fetches a row by the lookup kwargs or inserts one; update_or_create() also writes defaults onto a found row. Both return (object, created) and survive a race only when a database unique constraint covers the lookup fields.

solid answer

~40 s

`get_or_create(defaults=None, **kwargs)` runs a `get()` with the kwargs; if the row is missing it inserts one built from the exact-match kwargs plus `defaults`, and returns `(obj, created)`. `update_or_create()` does the same lookup but, when it finds a row, applies `defaults` to it and saves with `update_fields`; since Django 5.0 a separate `create_defaults` dict supplies values used only on insert. Under concurrency, `get_or_create()` does the insert inside `transaction.atomic()`, and if that raises `IntegrityError` it runs the `get()` again and returns the winner's row. That recovery only works when the database enforces uniqueness on the lookup fields; without a `UniqueConstraint` both requests insert, and later calls raise `MultipleObjectsReturned`. `update_or_create()` additionally locks an existing row with `select_for_update()` so two updates do not interleave.

code

python · 29 lines
python
from django.db import models


class Registration(models.Model):
    workshop = models.ForeignKey("Workshop", on_delete=models.CASCADE)
    email = models.EmailField()
    name = models.CharField(max_length=200)
    source = models.CharField(max_length=20, default="web")

    class Meta:
        constraints = [
            models.UniqueConstraint(
                fields=["workshop", "email"], name="uniq_registration"
            ),
        ]


# Existing row is returned untouched; defaults apply only on insert.
reg, created = Registration.objects.get_or_create(
    workshop=workshop, email=email, defaults={"name": name}
)

# Existing row gets defaults written; create_defaults apply only on insert.
reg, created = Registration.objects.update_or_create(
    workshop=workshop,
    email=email,
    defaults={"name": name},
    create_defaults={"name": name, "source": "waitlist"},
)

go deeper

for a junior

Recall the (obj, created) return tuple and that defaults are used only when a row is created by get_or_create().

for a middle

Explain the get, atomic insert, IntegrityError, get-again sequence, and what update_or_create() adds: a row lock and a save with update_fields.

for a senior

Show that race safety comes from the database constraint, not the method; diagnose MultipleObjectsReturned as a missing UniqueConstraint and fix it with a migration.

for a principal

Judge where idempotent upserts belong: ORM helpers with constraints, a native INSERT ... ON CONFLICT through bulk_create(update_conflicts=True), or a queue that serialises writers.

## The two methods and their return value Both are **QuerySet methods** in Django's ORM, usually called through a manager such as `Registration.objects`. Both return a **two-item tuple** `(obj, created)`: the model instance and a boolean that is `True` only when a new row was inserted. - **`get_or_create(defaults=None, **kwargs)`**: look up a row matching `kwargs`; if there is none, create one. An existing row is returned **untouched**. - **`update_or_create(defaults=None, create_defaults=None, **kwargs)`**: look up a row matching `kwargs`; if there is one, set the fields in `defaults` on it and save; if not, create one. The keyword arguments split into two roles: | Argument | Used for lookup | Used on insert | Used on update | |---|---|---|---| | plain `kwargs` (e.g. `workshop=w, email=e`) | yes | yes, if they are exact matches (no `__`) | no | | `defaults` | no | yes (`get_or_create`; `update_or_create` without `create_defaults`) | yes (`update_or_create` only) | | `create_defaults` (Django 5.0+) | no | yes (`update_or_create` only) | no | Callables inside `defaults` and `create_defaults` are called at write time, so `defaults={"joined_at": timezone.now}` is evaluated only when needed. If your model has a field literally named `defaults` or `create_defaults`, filter on it as `defaults__exact` or `create_defaults__exact`. ## How `get_or_create()` handles a race The Django 6.1 implementation is short and worth knowing by heart: 1. Mark the query for the write database and try `self.get(**kwargs)`. If it succeeds, return `(obj, False)`. 2. On `DoesNotExist`, build the parameters and, **inside `transaction.atomic()`**, call `create()`. 3. If that insert raises **`IntegrityError`**, try `self.get(**kwargs)` once more. If a row is now there (another request inserted it first), return `(that_row, False)`. If it is still missing, re-raise the error. The `atomic()` block matters when you are already inside a transaction: it becomes a savepoint, so the failed insert can be rolled back without poisoning the outer transaction, and the second `get()` can run. The weak point is in step 3: an `IntegrityError` only happens if the **database** rejects the duplicate. The documentation says it plainly: the method is atomic *assuming the database enforces uniqueness* of the lookup fields. With no constraint: - both racing requests see `DoesNotExist`; - both inserts succeed, leaving two rows; - every later `get_or_create()` with the same kwargs raises **`MultipleObjectsReturned`**, because its `get()` now matches two rows. The fix is a constraint in the model's `Meta`, for example `UniqueConstraint(fields=["workshop", "email"], name="uniq_registration")`, plus a migration. ## How `update_or_create()` differs `update_or_create()` wraps everything in `transaction.atomic()` and calls `self.select_for_update().get_or_create(create_defaults, **kwargs)`: - If the row **exists**, the `SELECT ... FOR UPDATE` locks it until the block ends, so a second `update_or_create()` on the same row waits instead of interleaving its write. Then `defaults` are applied with `setattr` and the object is saved with **`update_fields`** limited to those fields plus fields that compute a value on save, such as `auto_now` timestamps (the `update_fields` behaviour dates from Django 4.2). - If the row **does not exist**, there is nothing to lock, so the insert path is exactly `get_or_create()`'s, including its dependence on a unique constraint. It returns `(obj, created)` like `get_or_create()`, with `obj` reflecting the update. ## Choosing between them - Use **`get_or_create()`** when an existing row is the right answer as it stands: an idempotent sign-up, a tag looked up by name, a settings row created on first use. - Use **`update_or_create()`** when the caller's data should win: syncing a profile from an external source, recording the latest status of a job. - Use **`create_defaults`** when a new row needs values an update must never overwrite, such as a `source` or `created_by` field. - Reach for **`bulk_create(update_conflicts=True, ...)`** instead when you upsert many rows at once; calling `update_or_create()` in a loop costs a locked read and a write per row. ## Common mistakes - Putting a changing value in the lookup kwargs instead of `defaults`, e.g. `get_or_create(email=e, name=n)`; a user who fixes a typo in their name gets a second row. - Expecting `get_or_create()` to refresh an existing row with `defaults`; it never does, that is `update_or_create()`'s job. - Calling either method from a `GET` view; the docs recommend them only for requests that are allowed to change data. - Using `get_or_create()` through a related manager (`workshop.registration_set.get_or_create(...)`) and forgetting that the lookup is scoped to that relation, so a row that exists but belongs to another parent triggers an insert that can fail with `IntegrityError` on a unique field.

  • Why does Django's get_or_create() wrap the insert in transaction.atomic() instead of just catching IntegrityError?
    Inside an outer transaction, a failed INSERT leaves the database transaction unusable on PostgreSQL until it is rolled back. The inner `atomic()` becomes a savepoint, so only the failed insert is rolled back and the follow-up `get()` that finds the winner's row can still run.
  • In Django, what happens if update_or_create() is called with a field in defaults that is not a concrete column?
    It cannot pass that name in `update_fields`, which only accepts concrete fields, so it falls back to a full `save()` that writes every column instead of the restricted update. The object is still updated; the write is just wider.
  • In Django, can get_or_create() be combined with filter() and Q objects?
    Yes. `Registration.objects.filter(Q(email=a) | Q(email=b)).get_or_create(workshop=w, defaults={...})` looks up within the filtered set. Only exact-match keyword arguments of the `get_or_create()` call itself feed the new row, so the Q conditions never become field values.

saying these in an interview costs you the question

  • get_or_create() updates the found row with the values in defaults
  • get_or_create() locks the table, so it never creates duplicates
  • update_or_create() returns only the object, not a created flag
  • Fields passed in defaults are part of the lookup
  • A unique constraint is unnecessary because Django checks for duplicates first
open as a page

When two Django requests register for the last seat of a workshop at the same moment and both succeed, why does it happen and how do update() and F() fix it?

level: seniorimportance: must knowfreq 50%

basics

~10 s

Both requests read seats_taken before either writes, so both pass the capacity check in Python. A conditional filter(seats_taken__lt=F("capacity")).update(seats_taken=F("seats_taken") + 1) moves check and increment into one UPDATE; a return of 0 means full.

open as a page

In Django's ORM, what does Model.objects.create() do, and how does it differ from building an instance and calling save()?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Model.objects.create(**kwargs) builds the instance and saves it with force_insert=True in one call, returning the saved object. It always issues an INSERT, while save() on an instance with a hand-set primary key may issue an UPDATE instead.

open as a page

In Django's ORM, which write calls bypass a model's save() method and its pre_save and post_save signals, and what does QuerySet.delete() still run?

level: middleimportance: should knowfreq 55%

basics

~10 s

QuerySet.update(), bulk_create() and bulk_update() skip save() and the pre_save/post_save signals. QuerySet.delete() skips each instance's delete() method but still sends pre_delete and post_delete for every row it deletes, including Python-emulated cascades.

open as a page

When importing 200,000 attendee rows with Django's bulk_create() and later correcting them with bulk_update(), which caveats and batch_size choices matter?

level: seniorimportance: should knowfreq 42%

basics

~20 s

bulk_create() inserts in batches without save() or save signals, sets primary keys only on PostgreSQL, MariaDB and SQLite, and loses them with ignore_conflicts. bulk_update() writes CASE WHEN updates per batch; batch_size bounds statement size and memory.

open as a page