skip to content

Django REST Framework takes seconds to serialize 5,000 articles although the query count is constant; where does the time go, and what helps?

level: seniorimportance: nice to knowfreq 28%

answer

  1. queries fixed, Python still per row
  2. field calls times rows
  3. one serializer per row is the trap
  4. smaller pages, slimmer list serializer

basics

~20 s

The time goes to Python: DRF calls get_attribute() and to_representation() for every field of every row, and nested serializers, method fields and hyperlinks multiply that. Paginate, use a slim read-only list serializer, and never build a serializer per row.

solid answer

~40 s

Once queries are fixed, the cost is DRF's per-field work: for each object, `Serializer.to_representation()` loops over the readable fields, calling `get_attribute()` and the field's `to_representation()`. 5,000 rows times 25 fields is 125,000 field calls, plus nested serializers per child, `SerializerMethodField` methods and a `reverse()` per `HyperlinkedIdentityField`. With `many=True`, DRF builds one child serializer and its fields once, so construction is not the problem — unless code builds a serializer per row, in a loop or inside a method field, which deep-copies the fields every time. What helps: pagination (off by default: `DEFAULT_PAGINATION_CLASS` is `None`), a slim read-only serializer for the list action chosen in `get_serializer_class()`, dropping per-row hyperlinks and method fields, and for exports skipping serializers entirely.

code

python · 30 lines
python
from rest_framework import serializers, viewsets
from rest_framework.pagination import PageNumberPagination

from blog.models import Article
from blog.serializers import ArticleDetailSerializer


class ArticleListSerializer(serializers.ModelSerializer):
    author_name = serializers.CharField(source="author.name", read_only=True)

    class Meta:
        model = Article
        fields = ["id", "title", "author_name", "published_at"]
        read_only_fields = fields


class ArticlePagination(PageNumberPagination):
    page_size = 50


class ArticleViewSet(viewsets.ModelViewSet):
    pagination_class = ArticlePagination

    def get_queryset(self):
        return Article.objects.select_related("author")

    def get_serializer_class(self):
        if self.action == "list":
            return ArticleListSerializer
        return ArticleDetailSerializer

go deeper

for a junior

Know that DRF does not paginate by default and that serializing thousands of rows in one response is slow even with few queries.

for a middle

Explain the per-row field loop in to_representation(), why many=True builds one child serializer, and why building serializers in a loop repeats construction.

for a senior

Profile before optimising, then paginate, give the list action a slim read-only serializer, and move exports off serializers entirely.

for a principal

Weigh API shape against cost: list and detail representations can differ deliberately, and bulk data may deserve its own export path rather than a bigger JSON page.

## Why a constant query count can still be slow Fixing N+1 queries removes database round trips, but Django REST Framework (DRF) still does **Python work for every field of every row**. For each object, `Serializer.to_representation()` iterates `_readable_fields` and, per field: 1. calls `field.get_attribute(instance)` to fetch the value (following dotted `source=` paths); 2. checks for `None`; 3. calls `field.to_representation(value)` to turn it into a primitive — formatting a datetime, converting a decimal, rendering a nested serializer, reversing a URL. For 5,000 articles with 25 fields each, that is 125,000 field calls before the renderer even starts producing JSON — and it runs in one request, on one worker. ## What multiplies the cost | Field or pattern | Extra work per row | |---|---| | nested serializer (`AuthorSerializer`) | its own field loop for each child | | `many=True` nested list (`TagSerializer`) | a field loop per child element | | `SerializerMethodField` | a Python method call, often more | | `HyperlinkedIdentityField` / `HyperlinkedRelatedField` | a `reverse()` URL build per value | | `DateTimeField`, `DecimalField` | formatting and conversion | ## Construction cost: when it matters Building a serializer is not free. `Serializer.fields` is computed lazily: `get_fields()` **deep-copies** the declared fields, and `ModelSerializer` also introspects the model to build its generated fields. With `many=True`, DRF's `many_init()` creates a `ListSerializer` wrapping **one** child instance, whose fields are built once and reused for every row, so for a normal list this cost is paid once per request. It becomes a per-row cost when code builds serializers per row: - `[ArticleSerializer(a).data for a in articles]` instead of `ArticleSerializer(articles, many=True).data`; - a `SerializerMethodField` returning `AuthorSerializer(obj.author).data` instead of a declared nested field. Each builds fields again, including the model introspection for a `ModelSerializer`, 5,000 times. ## What helps, in order 1. **Paginate.** DRF ships with pagination off: `DEFAULT_PAGINATION_CLASS` is `None` and `PAGE_SIZE` is `None`, so a plain `ListAPIView` serializes the whole queryset. Set a pagination class and page size; nobody reads 5,000 rows in one response. 2. **Use a slim read-only serializer for the list action.** Override `get_serializer_class()` to return, for `self.action == "list"`, a serializer with only the columns a list shows, all read-only — no nested author object, just `author_name` via an annotation, no per-row hyperlinks, no method fields. Keep the full serializer for retrieve and writes. 3. **Replace method fields with annotated attributes**, so each becomes a plain attribute read. 4. **Drop hyperlinks in lists** if clients do not follow them; an `id` is cheaper than a reversed URL. 5. **For bulk exports, bypass serializers**: stream rows fetched with `values()` or `iterator()` directly. Those ORM tools belong to the large-results topic; the DRF point is that a serializer is the wrong tool for a 100,000-row export. 6. **Cache** the rendered list when it is shared and changes rarely. ## A worked estimate For the article list, suppose each row renders 12 plain fields, a nested author (5 fields), 3 tags (3 fields each), one method field and one hyperlink: - 12 + 5 + 9 = 26 field conversions, plus one method call and one `reverse()`, per article; - for 5,000 articles: about 130,000 field conversions, 5,000 method calls and 5,000 URL reversals — all in one request. A page of 50 turns that into about 1,300 conversions; a list serializer with 5 flat fields and no hyperlink turns it into 250. ## Measuring before optimising - Profile one request: if most time is in `to_representation()` and field methods, you are CPU-bound in serialization; if it is in query execution, go back to the loading plan. - Compare a page of 50 against the full list: time should scale with rows; if it does not, something per-request is slow instead. - Watch for one field dominating — frequently a method field or hyperlink. ## What does not help - `read_only=True` on fields does not make output faster; it only removes them from input. - Switching `ModelSerializer` to a hand-written `Serializer` saves construction time, which `many=True` already pays only once; it does not remove the per-row field loop.

  • Why is [ArticleSerializer(a).data for a in qs] slower than ArticleSerializer(qs, many=True).data with identical output?
    The loop builds a new serializer per row, so each one deep-copies its declared fields and, for a `ModelSerializer`, introspects the model again. With `many=True`, a `ListSerializer` wraps one child serializer whose fields are built once and reused for every row.
  • Does DRF paginate list endpoints by default?
    No. `DEFAULT_PAGINATION_CLASS` and `PAGE_SIZE` both default to `None`, so `list()` serializes the whole filtered queryset unless you configure a pagination class globally or on the view.

saying these in an interview costs you the question

  • Once the query count is constant, serialization time cannot be the bottleneck.
  • DRF paginates list endpoints to 100 rows by default.
  • many=True builds a new serializer instance for every row.
  • Marking every field read_only makes rendering noticeably faster.
  • Returning a nested serializer's .data from a method field is as cheap as a nested field.