skip to content

In Django REST Framework, what does passing many=True to a serializer actually create, and how do you support bulk create and bulk update?

level: seniorimportance: should knowfreq 36%

answer

  1. __new__ returns another class
  2. child serializer inside a list
  3. Meta.list_serializer_class
  4. update has no default

basics

~20 s

many=True makes the serializer's new return a ListSerializer wrapping one instance of your class as child. Its create() calls the child's create() per item; update() raises NotImplementedError, so bulk writes need a custom ListSerializer set via Meta.list_serializer_class.

solid answer

~40 s

`ProductSerializer(data=items, many=True)` never builds a `ProductSerializer`: `__new__` pops `many` and calls `many_init()`, which returns a `ListSerializer` (or `Meta.list_serializer_class`) whose `child` is a `ProductSerializer`. List-only arguments such as `allow_empty`, `min_length` and `max_length` go to the list, most others to both. Validation requires a list, validates each item with the child, and in DRF 3.18 reports errors as a dict keyed by item index. `save()` calls the list's `create()`, which by default calls the child's `create()` once per item, so N inserts; overriding it with `bulk_create()` gives one statement but skips `save()`, signals and many-to-many writes. `update()` raises `NotImplementedError` because DRF cannot know how to match, insert or delete items; a custom `ListSerializer.update()` must match instances by id and decide those rules.

code

python · 36 lines
python
from rest_framework import serializers

from catalog.models import Product


class ProductBulkSerializer(serializers.ListSerializer):
    def run_child_validation(self, data):
        # validate each item against the row it will update (instance is a queryset)
        self.child.instance = None
        if self.instance is not None and data.get("id") is not None:
            self.child.instance = self.instance.filter(pk=data["id"]).first()
        self.child.initial_data = data
        return super().run_child_validation(data)

    def create(self, validated_data):
        products = [Product(**attrs) for attrs in validated_data]
        return Product.objects.bulk_create(products)  # no save(), no signals, no M2M

    def update(self, instance, validated_data):
        by_id = {product.id: product for product in instance}
        updated = []
        for attrs in validated_data:
            product = by_id.get(attrs.get("id"))
            if product is None:
                raise serializers.ValidationError({"id": f"Unknown product {attrs.get('id')}."})
            updated.append(self.child.update(product, attrs))
        return updated  # items missing from the payload are left untouched


class ProductSerializer(serializers.ModelSerializer):
    id = serializers.IntegerField(required=False)  # writable so bulk updates can match rows

    class Meta:
        model = Product
        fields = ["id", "sku", "name", "price", "stock"]
        list_serializer_class = ProductBulkSerializer

go deeper

for a junior

Know that many=True lets one serializer handle a list of objects, for output and for input.

for a middle

Explain that many=True builds a ListSerializer around a child, which arguments go where, and what the list validation errors look like in DRF 3.18.

for a senior

Implement bulk create and update with list_serializer_class, and account for what bulk_create() skips and which matching, insert and delete rules an update needs.

for a principal

Decide whether a bulk endpoint belongs in the API at all, and what size limits, partial-failure semantics and side-effect guarantees it must promise.

## many=True returns a different class `serializers.Serializer.__new__` intercepts the `many` keyword. When it is true, the constructor does not return an instance of your serializer at all; it calls the class method **`many_init()`**, which: 1. creates one instance of your serializer as the **`child`**; 2. reads `Meta.list_serializer_class`, defaulting to **`ListSerializer`**; 3. returns that list serializer, wrapping the child. So `isinstance(ProductSerializer(products, many=True), ProductSerializer)` is `False`. Every item, on output and on input, goes through the same child instance. ## How arguments are split - **`allow_empty`**, **`min_length`** and **`max_length`** belong to the list only; they are removed before the child is built. - Most other arguments, including `instance`, `data`, `context`, `partial`, `read_only` and `required`, are passed to the child and, where the list serializer understands them, to the list as well. - To control the split yourself, override the `many_init()` class method; DRF's docstring shows a minimal version that builds the child and returns a custom list serializer. ## Validating a list - The input must be a **list**; anything else fails with `Expected a list of items but got type "dict".` under the non-field errors key. - `allow_empty` defaults to `True`; with `allow_empty=False` an empty list fails with `This list may not be empty.` - `max_length` and `min_length` bound the number of items, which is a cheap guard for a bulk endpoint. - Each item is validated by the child through `run_child_validation()`; `validated_data` is a list of dicts. - In **DRF 3.18**, errors are a **dict keyed by the index** of each invalid item, and valid items are omitted. The older format, a list with an empty dict per valid item, is available through `LIST_SERIALIZER_ERRORS_AS_DICT = False`, is deprecated, and is scheduled for removal in 3.20. | Operation | `ListSerializer` default | Consequence | |---|---|---| | output (`.data`) | calls the child's `to_representation()` per item | per-item cost multiplies on large lists | | `create()` | calls the child's `create()` per item | one insert per product | | `update()` | raises `NotImplementedError` | bulk update needs your own rules | ## save() on a list - `save(**kwargs)` merges the keyword arguments into **every** item's validated dict, so `serializer.save(supplier=supplier)` stamps each product in the batch. - It calls `update()` when an instance was passed to the serializer and `create()` otherwise, and asserts that the method returned something. - `save(commit=False)` is rejected by an assertion that points you at `validated_data` instead. ## Bulk create The default `create()` is correct but issues one insert per item. A custom list serializer can replace it with a single `bulk_create()`: - set `list_serializer_class = ProductBulkSerializer` in the child's `Meta`; - implement `create(self, validated_data)` to build unsaved `Product(**attrs)` objects and pass them to `Product.objects.bulk_create(...)`, returning the list. Django's `bulk_create()` has limits that now apply to your endpoint: the model's **`save()` is not called**, **`pre_save` and `post_save` signals are not sent**, **many-to-many relations are not supported**, and primary keys of the new rows are set only on some databases, PostgreSQL among them. If a catalog import relies on a `save()` override to compute a slug, that logic has to move. ## Bulk update DRF refuses to guess because a list update raises questions only your domain can answer: 1. **Matching**: which incoming item updates which existing row? Usually by `id` or `sku`. 2. **Insertions**: does an item with no match create a product or fail? 3. **Deletions**: does a product missing from the payload get deleted, deactivated or left alone? A custom `ListSerializer.update(self, instance, validated_data)` receives the queryset passed as `instance` and the list of validated dicts; it builds a lookup by id, calls `self.child.update()` for matches and applies your insert and delete rules. Two details make it work: the item's `id` must be present in `validated_data`, which means declaring it as a writable field because the generated primary key is read-only; and per-item validation must know which existing row each item belongs to. With `many=True` the child's `instance` is the whole queryset, so a uniqueness validator on, say, `sku` raises a `RuntimeError` asking you to override `ListSerializer.run_child_validation()` and set `self.child.instance` per item, which is exactly what DRF's docstring for that method shows. Wrapping the whole update in a transaction is a separate concern.

  • What does validation error output look like for a DRF ListSerializer in 3.18?
    A dict keyed by the index of each invalid item, each value being that item's field errors; valid items are omitted. Before 3.18 it was a list with an empty dict for every valid item. Setting `LIST_SERIALIZER_ERRORS_AS_DICT` to `False` restores that list temporarily; it is deprecated and due for removal in 3.20.
  • Why must the id be declared as a writable field for a DRF bulk update?
    A `ModelSerializer` makes the `AutoField` primary key read-only, so any `id` in the input is ignored and never reaches `validated_data`. Without it, the list serializer's `update()` cannot match items to existing rows. Declaring `id = serializers.IntegerField(required=False)` keeps it in the validated data.

saying these in an interview costs you the question

  • many=True returns an instance of the same serializer class in list mode.
  • ListSerializer.update() matches items to existing rows by primary key automatically.
  • The default bulk create issues a single INSERT for the whole list.
  • Swapping in bulk_create() keeps save() overrides and post_save signals working.
  • In DRF 3.18 list errors are a list with an empty dict per valid item.