skip to content

In Django's ORM, what does it mean that a QuerySet is lazy, and which operations make it actually query the database?

level: juniorimportance: must knowfreq 78%

answer

  1. building is not running
  2. every chain returns a copy
  3. rows needed, SQL sent
  4. len, bool, list, iteration, repr

basics

~20 s

Building or chaining a QuerySet only composes a query and returns a new QuerySet. SQL runs when Python needs rows: iteration, list(), len(), bool() or an if test, repr(), a slice with a step, or pickling.

solid answer

~40 s

`Order.objects.filter(status='open')` sends no SQL; it returns a `QuerySet` object describing a query, and each chained `.exclude()`, `.order_by()` or `.filter()` returns a *new* QuerySet with one more clause. The database is hit only when rows are needed: a `for` loop or a template `{% for %}`, `list()`, `len()`, `bool()` or an `if`, `repr()` in the shell, a slice with a step, or pickling. Some methods run at once because they return something other than a QuerySet: `get()`, `count()`, `exists()`, `first()`, `aggregate()`, `update()` and `delete()`. After a full evaluation the QuerySet keeps its rows in a result cache, so iterating the same object again runs no query.

code

python · 11 lines
python
from orders.models import Order

qs = Order.objects.filter(status='open')        # no query
qs = qs.exclude(total=0).order_by('-created_at')  # still no query
print(qs.query)                                  # renders SQL, executes nothing

for order in qs:                                 # SELECT runs here, fills the cache
    print(order.pk)

len(qs)                                          # served from the cache
qs.filter(total__gt=100).exists()                # new QuerySet: new query

go deeper

for a junior

Recall that building and chaining a QuerySet runs no SQL, and list the operations that do: iteration, list, len, bool, repr, step slices and pickling.

for a middle

Explain the result cache: which operations fill it, why a derived QuerySet queries again, and why repr() and indexing leave it empty.

for a senior

Use laziness deliberately: predict where a query will run in templates, serializers or logs, and catch evaluation happening in the wrong layer during review.

for a principal

Consider how much hidden query timing a team can tolerate, and when to make evaluation explicit at a service boundary instead of passing lazy QuerySets around.

## What a QuerySet is A Django `QuerySet` is a Python object that **describes** a database query: which model, which filters, which ordering, which limits. Calling a manager method such as `Order.objects.filter(status='open')` builds that description and returns it. No SQL is sent. The compiled query can be inspected with `print(qs.query)`, which renders the SQL without executing it. This deferral is called **laziness**: the query runs as late as possible, when the program actually needs the rows. ## Chaining returns new QuerySets - Every chainable method (`filter()`, `exclude()`, `order_by()`, `annotate()`, `values()`, `all()`, a slice without a step) returns a **new** QuerySet; the original is not modified. - Because of that, a base QuerySet can be refined in several directions: `open_orders.filter(total__gt=100)` and `open_orders.order_by('-created_at')` are independent objects. - Building is cheap, but not entirely inert: an unknown field name in `filter()` raises `FieldError` immediately, because lookups are resolved against the model's metadata as the query is built. Errors that only the database can report, such as a missing table, appear at evaluation time, possibly while a template is rendering. ## What evaluates a QuerySet | Operation | Runs SQL? | Fills the result cache? | |---|---|---| | iteration (`for`, template `{% for %}`) | yes | yes | | `list(qs)` | yes | yes | | `len(qs)` | yes, full SELECT | yes | | `bool(qs)`, `if qs:` | yes, full SELECT | yes | | `obj in qs` | yes | yes | | `repr(qs)`, shell echo | yes, at most 21 rows | no | | `qs[5:10]` | no, returns a QuerySet | no | | `qs[0:10:2]` (step) | yes, returns a list | no | | `qs[3]` | yes, `LIMIT 1` | no | | `pickle.dumps(qs)` | yes | yes | The `repr()` row explains a common surprise: typing a QuerySet in the shell prints results, so it must have queried, but it fetched only the first 21 rows on a copy and left the original's cache empty. ## Methods that run immediately Some methods are not chainable because they return a value rather than a QuerySet, so they execute on the spot: 1. **Single objects**: `get()`, `first()`, `last()`, `latest()`, `earliest()`. 2. **Scalars**: `count()`, `exists()`, `contains(obj)`, `aggregate()`. 3. **Writes**: `create()`, `update()`, `delete()`, `bulk_create()`. `count()`, `exists()` and `contains()` first look at the result cache, so on an already evaluated QuerySet they answer without a query. ## The result cache The first full evaluation stores the model instances in the QuerySet's **result cache**. Later iteration, `len()`, `bool()`, indexing and slicing of that **same object** read from the cache. A new QuerySet derived from it with `filter()` or `.all()` starts with an empty cache and queries again. This is the source of most "why does this run twice" bugs. ## Evaluation inside templates Views usually hand QuerySets to templates unevaluated, so the template engine is often where the SQL actually runs: - `{% for order in orders %}` iterates and fills the cache. - `{% if orders %}` tests truthiness, which is a full evaluation, not an `exists()`. - `{{ orders|length }}` calls `len()`, and `{{ orders.count }}` calls the `count()` method. - `{{ orders|slice:":5" }}` slices, which on an unevaluated QuerySet stays a limited query. The same object used in several of these tags shares one cache; a fresh lookup such as `user.orders.all` in each tag does not. Reading a template with the evaluation table in mind is the fastest way to predict its query count. ## Why laziness is useful - Views can build a QuerySet, hand it to helpers that add filters, and pass it to a template; one query runs, when the template loops. - A paginator can take an unevaluated QuerySet and turn it into a counted, sliced query instead of loading every row. - Default ordering, related-object loading and limits can all be added before anything is fetched. The cost of laziness is that the query's **timing** is hidden: it may run in a template, in a serializer or in a log statement, far from the line that built it. Knowing the evaluation list above is what lets you predict where the SQL will actually happen.

  • If a filter() uses a field name that does not exist, when does the error appear?
    Immediately, as a `FieldError` from `filter()`, because Django resolves lookups against model metadata while building the query. Laziness only postpones sending SQL: problems the database itself detects, such as a missing table or column after a failed migration, surface when the QuerySet is evaluated, which may be inside template rendering.
  • Why does typing a QuerySet at the shell prompt show rows, yet a later loop over it still queries?
    The shell calls `repr()`, which slices the QuerySet to its first 21 rows and fetches them on a copy, so it prints at most 20 plus a truncation marker. That slice does not fill the original's result cache, so the first full iteration afterwards runs its own SELECT.

A QuerySet is a shopping list: writing items or crossing some out buys nothing. The trip to the shop happens when someone reads the list to fill a basket, and the basket then stays on the table; writing a new list means a new trip.

saying these in an interview costs you the question

  • filter() runs a SELECT as soon as it is called
  • Chaining filter() modifies the original QuerySet in place
  • Printing a QuerySet in the shell never touches the database
  • count() and exists() are lazy like filter()
  • Slicing a QuerySet loads the whole table and slices in Python