Why can a Django QuerySet stored at module level, on a class, or pickled into a cache serve stale rows, and how do you avoid it?
answer
- cache outlives the request
- same object, same rows
- pickling forces evaluation
- clone with .all() per use
- pickle the query instead
basics
~20 sA QuerySet's result cache lives as long as the object. One stored at module or class level fills once and replays those rows for the process's life; pickling evaluates it and freezes the rows. Clone per use, or pickle qs.query.
solid answer
~40 sThe result cache belongs to the QuerySet object, and nothing in Django empties it. A module-level `OPEN = Order.objects.filter(status='open')` is lazy until the first request iterates it; from then on every request in that worker process iterating `OPEN` gets the same rows, so new orders never appear. Class attributes behave the same, which is why `ListView.get_queryset()` and `ModelChoiceField` call `.all()` to work on a fresh clone. Pickling, including `cache.set(key, qs)`, calls `__getstate__`, which evaluates the whole QuerySet and stores its rows as they were; the unpickled object never re-queries, and a Django version mismatch triggers a `RuntimeWarning`. To reuse a query without its rows, pickle `qs.query` and assign it to a fresh QuerySet, or store a function that builds the QuerySet.
code
python · 19 linesfrom django.views.generic import ListView
from orders.models import Order
OPEN = Order.objects.filter(status='open') # lazy at import time
def stale_dashboard(request):
rows = list(OPEN) # first request fills OPEN's cache; later ones reuse it
...
def fresh_dashboard(request):
rows = list(OPEN.all()) # .all() clones: new object, empty cache, fresh SELECT
...
class OpenOrderList(ListView):
queryset = Order.objects.filter(status='open') # safe: get_queryset() calls .all()go deeper
Know that a QuerySet stored outside a view keeps whatever rows it loaded first, so build QuerySets inside the code that uses them.
Explain why ListView and ModelChoiceField call .all() on stored QuerySets, and that pickling forces a full evaluation.
Hunt snapshot bugs in module globals, class attributes, cache.set calls and long-running loops, and choose between caching rows and caching qs.query deliberately.
Define what may be cached where, with explicit staleness budgets, so accidental QuerySet snapshots do not become invisible data-freshness incidents across workers.
## The cache has the object's lifetime A Django `QuerySet` fills its **result cache** on its first full evaluation, and keeps it for as long as the object exists. There is no expiry, no request-end hook and no invalidation when rows change. Inside a view that is what makes reuse cheap. Outside a request, it turns a query into a **snapshot**. ## Module-level and class-level QuerySets 1. At import time, `OPEN = Order.objects.filter(status='open')` builds a lazy QuerySet; no query runs, which is why this pattern looks harmless. 2. The first request that iterates `OPEN` itself runs the SELECT and fills the cache. 3. Every later request in the same worker process iterating `OPEN` reads the cached rows. New, closed or edited orders never show up until the process restarts, and different workers show different snapshots. Deriving a new QuerySet from it, as in `OPEN.filter(...)` or `OPEN.all()`, creates a fresh object with an empty cache, so the bug appears only where the exact stored object is evaluated. Django's own components defend against it: - `MultipleObjectMixin.get_queryset()`, used by `ListView`, calls `.all()` on a class-level `queryset` attribute before each use. - `ModelChoiceField` re-clones its `queryset` with `.all()` when the form field is deep-copied for each form instance. The rule for your own code: store the **recipe**, not the evaluated object. Keep a function or a manager method that returns a new QuerySet, or call `.all()` on a shared QuerySet at the point of use. ## Pickling evaluates `QuerySet.__getstate__` calls the internal fetch first, so pickling a QuerySet: - runs the full query, however many rows it matches; - stores the model instances inside the pickle, along with the Django version that produced it; - produces, on unpickling, a QuerySet whose cache holds those rows; it **never** re-queries while that cache is set. That is what the documentation intends: pickling is usually a prelude to caching, and a cached QuerySet should not hit the database. It becomes a bug when someone expects to cache the *question*: `cache.set('open_orders', Order.objects.filter(status='open'))` stores the answer as of now, possibly megabytes of rows. Options, by intent: | Intent | Store | On read | |---|---|---| | cache the rows for a while | the evaluated QuerySet, or better `list(qs)` or plain values | accept staleness up to the timeout | | cache the query definition | `pickle.dumps(qs.query)` | `qs = Order.objects.all(); qs.query = pickle.loads(data)` | | never go stale | nothing; build the QuerySet each time | one query per use | Pickles are also tied to the Django version: unpickling one produced by a different version emits a `RuntimeWarning`, and the docs warn that compatibility across versions is not guaranteed, which matters for a shared cache during a rolling upgrade. ## repr() in logs `repr(qs)` evaluates a limited copy, up to 21 rows, without filling the cache. A log line like `logger.debug(f'orders={qs!r}')` builds its f-string eagerly and therefore runs that query on every call, even when debug logging is off. Logging's own `%r` argument style defers the formatting until a record is actually emitted. ## Why it hides in development The snapshot bug is easy to miss locally: - The development server's autoreloader restarts the process on every code change, which resets module-level QuerySets and makes the data look fresh. - Tests often create rows before the first request touches the stored QuerySet, so the snapshot already contains them. - In production, a worker process lives for hours or days and serves thousands of requests from the first snapshot, and each worker holds a different one, so users see results that depend on which worker answered. When a bug report reads "new items appear only after a deploy", a stored, evaluated QuerySet is a prime suspect. ## Checklist for review - A QuerySet assigned at module level or as a class attribute: is the same object ever iterated directly? - Anything passing a QuerySet to `cache.set()`, a session or a pickle: is the intent rows or recipe? - An f-string or `str()` of a QuerySet in a log or exception message on a hot path. - Long-running processes, such as management commands or workers, that reuse one QuerySet across iterations of a loop and expect fresh data each time.
- Is it ever right to cache an evaluated QuerySet in Django's cache framework?Yes, when the rows are expensive and a bounded staleness is acceptable, such as a leaderboard refreshed every minute. Then cache deliberately: store `list(qs)` or plain values with a timeout, keep the payload small, and invalidate on writes that matter. What goes wrong is caching a QuerySet by accident while believing you cached the query.
- Why does a long-running management command that loops forever see no new rows?If it builds one QuerySet before the loop and iterates that same object each time, the first pass fills the cache and every later pass replays it. Build the QuerySet inside the loop, or call `.all()` each iteration, so each pass runs a new SELECT against current data.
saying these in an interview costs you the question
- Django clears QuerySet caches at the end of each request
- A module-level QuerySet re-queries on every use because it is lazy
- Pickling a QuerySet stores only its SQL
- An unpickled QuerySet refreshes its rows when iterated
- f-string logging of a QuerySet costs nothing when debug is off