skip to content

How does Django's page cache build the key for a stored response, and why does it keep a separate header list for each URL?

level: middleimportance: nice to knowfreq 25%

answer

  1. two keys per URL
  2. learn on store, look up on fetch
  3. absolute URI hashed
  4. language and time zone suffixes

basics

~20 s

Django hashes the full URL plus the request's values for the headers named in the response's Vary. Since Vary exists only on responses, it saves that header list under a URL-only key and reads it back on later requests.

solid answer

~40 s

The key has two layers. When storing, `learn_cache_key()` reads the response's `Vary` header, turns each name into its `request.META` form (`Cookie` becomes `HTTP_COOKIE`), stores that sorted list under a **header key** built from the URL, then builds the **page key**: `views.decorators.cache.cache_page.<prefix>.<method>.<md5 of absolute URI>.<md5 of those header values>`. When fetching, `get_cache_key()` has no response yet, so it reads the header list back first; if it is missing, the request is a miss. With `USE_I18N` the active language is appended and redundant `Accept-Language` is dropped; with `USE_TZ` the current time zone is appended. The alias's key function then adds `KEY_PREFIX` and `VERSION`. Django 6.1 changed how varied header values are hashed, so such pages miss once after upgrading.

go deeper

for a junior

Recall that page keys include the full URL and the headers the response varies on.

for a middle

Explain the header-list key learned on store and read on fetch, and why the fetch side cannot read Vary itself.

for a senior

Account for language, time zone, host and query string in the key when predicting hit rates or leaks.

for a principal

Judge whether per-URL, header-varied page caching can meet invalidation needs, or whether content should be cached below the page.

## The problem the design solves A cached page must be keyed by everything that changes its content. The URL is known on every request. The other inputs, the request headers the page **varies on**, are only declared by the response's `Vary` header. On the fetch side, before the view runs, there is no response to read. Django's solution is to **remember** the header list per URL. ## Two keys per URL Functions in `django.utils.cache` do the work: | Key | Format | Holds | |---|---|---| | Header key | `views.decorators.cache.cache_header.<prefix>.<md5(absolute URI)>` | The sorted list of `request.META` names to vary on | | Page key | `views.decorators.cache.cache_page.<prefix>.<method>.<md5(absolute URI)>.<md5(header values)>` | The stored response | Both then pass through the alias's key function, which adds `KEY_PREFIX` and `VERSION` from `CACHES`. The `<prefix>` is `CACHE_MIDDLEWARE_KEY_PREFIX` or `cache_page`'s `key_prefix`. ## Storing: learn_cache_key() When `UpdateCacheMiddleware` (or `cache_page`) stores a response: 1. It reads `Vary`, for example `Cookie, Accept-Encoding`. 2. Each name becomes its WSGI form: `HTTP_COOKIE`, `HTTP_ACCEPT_ENCODING`. 3. If `USE_I18N` is on, `Accept-Language` is dropped from the list, because the active language is appended to the key anyway and keeping both would store duplicates. 4. The sorted list is saved under the header key with the page's timeout. 5. The page key is built from the method, the URL hash and a hash of this request's values for those headers, and the response is stored under it. ## Fetching: get_cache_key() When `FetchFromCacheMiddleware` handles a GET or HEAD: 1. It reads the header list from the header key. 2. If there is none, it returns `None`: a miss, and the view runs. 3. Otherwise it hashes this request's values for the listed headers, builds the page key and looks it up. A HEAD request tries the GET entry first. If the header key expires or is evicted before the page, the next request is simply a miss; the page is rebuilt and the list relearned. ## What goes into the URL part - `request.build_absolute_uri()` covers **scheme, host, path and query string**, so `http` and `https`, different hosts, and `?month=10` versus `?month=11` are separate pages. - The **method** is part of the page key, which is how HEAD and GET entries are kept apart. ## Suffixes Django appends - With **`USE_I18N`**, the language code from `request.LANGUAGE_CODE` (set by `LocaleMiddleware`) or the active language. - With **`USE_TZ`**, the current time zone name. That is how one URL can safely hold English and German versions without the view declaring anything. ## Version note Django 6.1 changed the hashing of varied header values (each value is now length-delimited), so page keys for responses that vary on headers differ from older versions; the release notes warn that the first request to such a page after upgrading is a cache miss. Nothing breaks, but expect a short warm-up. ## Debugging a page that never hits When a page is supposed to be cached but every request runs the view: 1. Check the response's `Vary`: a `Cookie` entry keys the page per visitor, and `*` stops storage entirely. 2. Check whether the response sets a cookie, for example a CSRF or session cookie; combined with `Vary: Cookie` it is never stored. 3. Compare the full URL of consecutive requests, including scheme, host and query string, since tracking parameters create new keys. 4. Confirm both keys land in the same alias the fetch half reads, `CACHE_MIDDLEWARE_ALIAS` or `cache_page`'s `cache`. ## Consequences in practice - Invalidating one page early means rebuilding both keys yourself; most projects rely on the timeout. - A page varying on `Cookie` gets one entry per distinct cookie header, which can fill a cache quickly. - Two hosts pointing at one Django project cache separately because the host is in the URI.

  • Why is Accept-Language removed from the learned header list when USE_I18N is True?
    With `USE_I18N` on, Django appends the active language code to every page key anyway. Hashing the raw `Accept-Language` header as well would store the same German page separately for every browser's slightly different header value, wasting cache space without changing the content.
  • What happens if the header-list entry is evicted but the page entry is still in the cache?
    `get_cache_key()` finds no header list and returns `None`, so the request is treated as a miss and the view runs. On the way out the list is learned again and the page is stored again under the same page key, so the only cost is one extra render.

saying these in an interview costs you the question

  • Django's page cache key is just the URL path without the query string.
  • The fetch step reads Vary from the incoming request.
  • HTTP and HTTPS versions of a page share one cache entry.
  • The page cache ignores the active language unless the view varies on it.
  • CACHES KEY_PREFIX is not applied to page cache keys.