skip to content

In Django, why should a view write an uploaded file with UploadedFile.chunks() rather than read(), and how do chunks behave for in-memory uploads?

level: middleimportance: should knowfreq 42%

answer

  1. read() loads everything at once
  2. 64 KiB pieces by default
  3. multiple_chunks() tells you
  4. in memory means one chunk

basics

~20 s

UploadedFile.read() loads the whole file into memory, while chunks() yields it in 64 KiB pieces, so writing a large upload chunk by chunk keeps memory flat. For an InMemoryUploadedFile, chunks() yields the whole content once, since it is already in memory.

solid answer

~40 s

`read()` returns the entire file as one `bytes` object, so a 20 MB scanned invoice that Django carefully spooled to a temp file is pulled back into worker memory. `chunks(chunk_size=None)` is a generator that seeks to the start and yields pieces of `File.DEFAULT_CHUNK_SIZE`, 64 KiB, which you write to the destination one by one. `multiple_chunks()` returns `True` when the file is larger than the chunk size, so it tells you whether streaming matters. For an `InMemoryUploadedFile` the data is already in RAM, so `chunks()` yields everything in a single chunk and `multiple_chunks()` is always `False`. Because `chunks()` rewinds first, you can peek at the header with `read(8)` for a sanity check and still stream the whole file afterwards. In practice, saving through a storage backend or a `FileField` does the chunked copy for you.

code

python · 17 lines
python
import hashlib
from pathlib import Path

from django.core.exceptions import ValidationError

INBOX = Path("/srv/invoices/inbox")


def store_invoice(uploaded, target_name):
    if uploaded.read(5) != b"%PDF-":  # peek; chunks() rewinds afterwards
        raise ValidationError("Not a PDF")
    digest = hashlib.sha256()
    with open(INBOX / target_name, "wb") as out:
        for chunk in uploaded.chunks():  # 64 KiB at a time for disk-backed files
            digest.update(chunk)
            out.write(chunk)
    return digest.hexdigest()

go deeper

for a junior

Use a for loop over uploaded.chunks() to write an upload to disk instead of uploaded.read().

for a middle

Explain the 64 KiB default, multiple_chunks(), the single-chunk behaviour of InMemoryUploadedFile and the rewind before chunking.

for a senior

Keep per-request memory bounded under concurrent large uploads, and do hashing or validation during the chunked copy rather than after a full read.

for a principal

Decide where upload processing runs, in the request, in a background job reading a stored copy, or outside Django, based on file sizes and volume.

## The API on every uploaded file Every value in `request.FILES` is an `UploadedFile`, a subclass of Django's `File` wrapper. For reading the content it offers: | Method | Returns | Memory cost | |---|---|---| | `read()` | the whole content as `bytes` | the full file size | | `chunks(chunk_size=None)` | a generator of `bytes` pieces | one chunk at a time | | `multiple_chunks(chunk_size=None)` | `True` if the file is larger than one chunk | none | | iteration (`for line in f`) | lines split on newlines | unpredictable for binary data | The default chunk size is `File.DEFAULT_CHUNK_SIZE`, which is `64 * 2**10` bytes (64 KiB). Pass `chunk_size` to change it. ## Why read() is the wrong default Django's upload handlers already went to some trouble to keep large files out of memory: anything in a request larger than 2.5 MB is streamed into a temporary file. Calling `read()` undoes that. With scanned invoices of up to 20 MB and several uploads in flight, a worker can briefly hold hundreds of megabytes just to copy files it never needed to see whole. Iterating the file with `for line in uploaded` is not a fix for binary data: a PDF or a TIFF scan can contain very long stretches without a newline byte, so a "line" can be most of the file. `chunks()` gives a predictable bound: at most one chunk is in memory at a time, however large the upload. ## How chunks() behaves for each file class - **`TemporaryUploadedFile`** (disk-backed): `chunks()` uses the base `File.chunks()`. It seeks to position 0, then reads and yields `chunk_size` bytes until the file is exhausted. `multiple_chunks()` returns `self.size > chunk_size`. - **`InMemoryUploadedFile`** (memory-backed): the class overrides both methods. `chunks()` seeks to 0 and yields the whole content in **one** chunk, and `multiple_chunks()` always returns `False`. The rationale in the source: there is no good reason to read from memory in pieces. So code written with `chunks()` is correct for both, and efficient for the one that matters. ## The rewind detail Both `chunks()` implementations seek to the start first. That allows a quick sanity check before the copy: 1. `header = uploaded.read(5)` to check that a claimed PDF starts with `%PDF-`; 2. then `for chunk in uploaded.chunks(): ...` streams the entire file, including those five bytes. Without the rewind, the first bytes would be lost. `uploaded.content_type` comes from the client and should not be trusted, which is why a check like this is worth doing. ## Doing something useful per chunk Chunked reading also lets you compute things while copying: - a SHA-256 digest with `hashlib.sha256().update(chunk)` to spot duplicate invoice scans; - a running byte count to enforce a business limit on what you keep; - writing to a destination file opened with `"wb"`. ## Two different chunk sizes Two 64 KiB values appear in the upload path and are easy to confuse: - `FileUploadHandler.chunk_size` controls how the **multipart parser reads the request** and feeds handlers; the parser uses the smallest `chunk_size` among the installed handlers; - `File.DEFAULT_CHUNK_SIZE` controls how **`chunks()` reads the finished file** when your code copies it. Changing `chunk_size` in a `chunks()` call only affects your copy loop. Larger pieces mean fewer write calls and more memory per iteration; for disk-backed uploads a few hundred kilobytes is still modest. ## When you do not write the loop yourself Most projects never copy bytes by hand. Assigning the upload to a model `FileField` and saving, or calling a storage backend's `save()`, streams the file through the storage layer, which reads it in chunks internally. Knowing `chunks()` still matters for custom processing: hashing, virus-scan hand-off, OCR on a temporary copy, or writing to a location the storage layer does not manage.

  • Why does multiple_chunks() always return False for an InMemoryUploadedFile?
    The whole file is already in a `BytesIO`, so reading it in pieces saves nothing. `InMemoryUploadedFile` overrides `chunks()` to yield all content at once and `multiple_chunks()` to return `False`; the base `File` docstring says in-memory representations should always report `False`.
  • You call uploaded.read(5) to check the header and then loop over uploaded.chunks(). Are the first five bytes lost?
    No. Both `File.chunks()` and `InMemoryUploadedFile.chunks()` seek to position 0 before yielding, so the loop starts from the beginning of the file and writes it completely, header included.

saying these in an interview costs you the question

  • read() is fine because Django already streamed the file to disk
  • chunks() always yields 64 KiB pieces, even for in-memory uploads
  • Iterating an uploaded PDF line by line bounds memory use
  • After a partial read(), chunks() continues from the current position
  • multiple_chunks() returns True for any file over 2.5 MB held in memory