skip to content

A Django view that exports warehouse inventory as CSV runs out of memory and times out on large warehouses; how would you stream it with StreamingHttpResponse, and what do you give up?

level: seniorimportance: should knowfreq 50%

answer

  1. a generator, not a string
  2. HttpResponse drains iterators eagerly
  3. csv.writer over an echo buffer
  4. QuerySet.iterator() in chunks
  5. no Content-Length, no ETag

basics

~20 s

Return a StreamingHttpResponse over a generator that yields CSV rows built with csv.writer and QuerySet.iterator(), so memory stays flat and bytes flow early. You lose Content-Length, ETags, caching and the chance to change the status mid-stream.

solid answer

~40 s

The fix is to yield the file instead of building it. Read rows with `Item.objects.values_list(...).iterator(chunk_size=2000)` so the queryset is not cached, write each row with `csv.writer` over a tiny object whose `write()` just returns the string, batch rows, and return `StreamingHttpResponse(gen(), content_type="text/csv", headers={"Content-Disposition": 'attachment; filename="inventory.csv"'})`. Passing the same generator to a plain `HttpResponse` does not help: it joins the iterator in its constructor. The costs: no `Content-Length` or automatic `ETag`, the cache middleware skips it, the status is fixed at 200 once bytes leave, the generator runs after the view returns so it is outside `ATOMIC_REQUESTS`, and under WSGI it holds a worker for the whole download. Under ASGI yield from an async generator. For very large exports, build the file in the background and serve it instead.

code

python · 33 lines
python
import csv
from itertools import batched

from django.http import StreamingHttpResponse

from .models import StockItem


class Echo:
    def write(self, value):
        return value


def export_inventory(request, warehouse_id):
    # Validate and authorise here: once streaming starts the status is fixed.
    rows = (
        StockItem.objects.filter(warehouse_id=warehouse_id)
        .order_by("sku")
        .values_list("sku", "location", "qty")
        .iterator(chunk_size=2000)
    )
    writer = csv.writer(Echo())

    def stream():
        yield writer.writerow(["sku", "location", "qty"])
        for batch in batched(rows, 500):
            yield "".join(writer.writerow(row) for row in batch)

    return StreamingHttpResponse(
        stream(),
        content_type="text/csv",
        headers={"Content-Disposition": 'attachment; filename="inventory.csv"'},
    )

go deeper

for a junior

Know that StreamingHttpResponse takes an iterator and sends it in pieces, while HttpResponse builds the whole body first.

for a middle

Show the generator, the echo pseudo-buffer for csv.writer and QuerySet.iterator(), and explain why a generator in HttpResponse does not stream.

for a senior

Name the costs: no Content-Length or ETag, no caching, a fixed 200, generation outside ATOMIC_REQUESTS, held WSGI workers, and matching async iterators under ASGI.

for a principal

Decide between streaming and background generation per export size and latency, and set limits so a few huge exports cannot starve the worker pool.

## Why the naive export fails The usual first version builds the whole CSV in memory: it evaluates `StockItem.objects.all()` (which caches every model instance on the `QuerySet`), writes each row into an `HttpResponse` or a `StringIO`, and returns it. For a warehouse with millions of stock lines that means every row object and the full CSV string are held at once, and the client sees nothing until the last row is written, so a proxy or load balancer may close an idle connection first. A common half-fix is to pass a generator to `HttpResponse`. It does not stream: the `HttpResponse.content` setter consumes any iterator immediately with `b"".join(...)`, so the whole body is still built before the view returns. ## Streaming it properly `django.http.StreamingHttpResponse` takes an iterator as `streaming_content` and hands chunks to the server as they are produced. A working export has four parts: 1. **Stream from the database.** `values_list("sku", "location", "qty").iterator(chunk_size=2000)` fetches in chunks (2000 is the default when `chunk_size` is omitted, and it becomes mandatory after `prefetch_related()`) and does not fill the queryset's result cache. On backends that support them it uses server-side cursors. 2. **Format rows without a buffer.** `csv.writer` needs a file-like object. The documented trick is an `Echo` class whose `write(value)` returns `value`; `writer.writerow(row)` then returns the formatted line instead of storing it. 3. **Batch.** Yielding one row at a time adds per-chunk overhead; joining 100-1000 rows per yield (for example with `itertools.batched`, Python 3.12+) is cheaper and compresses better with `GZipMiddleware`. 4. **Set the headers up front.** `content_type="text/csv"` and a `Content-Disposition: attachment; filename="inventory.csv"` header, passed through `headers=`. ## What you give up | Capability | Buffered `HttpResponse` | `StreamingHttpResponse` | |---|---|---| | `Content-Length` | added by `CommonMiddleware` | absent (length unknown) | | Automatic `ETag` from the body | added by `ConditionalGetMiddleware` | skipped | | Page cache middleware | can store it | skips streaming responses | | `response.content` in code and tests | available | `AttributeError`; use `streaming_content` | | Changing status after an error | possible until returned | impossible once bytes are sent | Further consequences a senior engineer should name: - **Errors mid-stream.** If the generator raises halfway, the client already has a 200 and a truncated file. Validate inputs and permissions in the view, before returning the response. - **Transactions.** The generator runs after the view has returned, so with `ATOMIC_REQUESTS` it executes outside the view's transaction; do not write to the database from it. - **Worker occupancy under WSGI.** The documentation warns that Django under WSGI is designed for short-lived requests; a streamed response ties up a worker for the whole download. - **Sync versus async iterators.** Under ASGI, give it an async iterator (for example one using `QuerySet.aiterator()`). A mismatched iterator type still works, but Django warns and consumes it fully first, which defeats streaming. The same happens in reverse under WSGI. - **Middleware that reads the body** must use `streaming_content`; `GZipMiddleware` does and drops `Content-Length`. ## When not to stream Streaming fixes memory and first-byte latency, not total work. If an export takes minutes, the documentation's own advice is to do expensive work outside the request-response cycle: generate the file in a background job, store it, notify the user and serve it with `FileResponse` or from storage. Streaming fits the middle ground: large but fast-to-generate data where the user expects the download to start immediately. ## A pre-flight checklist for a streamed export 1. Authorise the user and validate filters in the view body, where you can still return a 403 or 400. 2. Order the queryset (for example by `sku`) so the file is deterministic and diffable. 3. Select only the needed columns with `values_list()` and read with `iterator()` or, under ASGI, `aiterator()`. 4. Set `content_type` and `Content-Disposition` before returning; nothing can be added once the body starts. 5. Keep the generator read-only and free of side effects. 6. Decide a size threshold above which the export moves to a background job. ## Testing it In a test, `b"".join(response.streaming_content)` (or `response.getvalue()`) gives the full body. Consuming it drains the iterator, so read it once.

  • Why does HttpResponse(generator()) not stream the CSV?
    The `HttpResponse.content` setter consumes any non-string iterable at once with `b"".join(...)` so the content can be read repeatedly. The whole body is therefore built in memory inside the constructor, before the view even returns. Only `StreamingHttpResponse` keeps the iterator lazy.
  • The same streaming view is moved to an ASGI deployment and memory climbs again. What is going on?
    Under ASGI Django iterates the response asynchronously. A sync generator is still accepted, but Django warns and first consumes it fully through `sync_to_async(list)`, so nothing streams. Supply an async generator, for example one looping over `qs.aiterator()`, so chunks are yielded as they are read.
  • How do you assert the CSV contents in a Django test?
    Read the iterator once: `body = b"".join(response.streaming_content)` or `response.getvalue()`, then decode and parse it with `csv.reader`. Accessing `response.content` raises `AttributeError` on a streaming response, and iterating a second time yields nothing.

A buffered export is a truck that is loaded completely before it leaves the dock; streaming is a conveyor belt that starts delivering boxes at once but cannot tell the customer in advance how many boxes are coming.

saying these in an interview costs you the question

  • Passing a generator to HttpResponse streams the body
  • StreamingHttpResponse lets you send a 500 if the generator fails midway
  • Django adds Content-Length to streaming responses automatically
  • Streaming makes a slow export finish faster
  • The generator's queries run inside the ATOMIC_REQUESTS transaction
  • Any iterator type streams equally well under WSGI and ASGI