Why should a contextvars.ContextVar be created at module level rather than inside a function?
answer
- The label is not the key
- Every call builds a stranger
- The previous value is never found
- Something keeps them alive
- The docs forbid closures for this
basics
~20 sThe ContextVar object is the key, not its name, so one built inside a function is a fresh variable each call and never sees what an earlier call set. Contexts also keep every variable set in them alive.
solid answer
~40 sLookup is by object identity: the `contextvars.ContextVar` instance is the key, and the `name` argument is only a label for `repr` and tracebacks. Build one inside a function and each call creates a fresh, unrelated variable — a value stored by one call is invisible to the next, and `get()` keeps returning the default even though a `set()` clearly happened. There is a lifetime cost too: a context holds a strong reference to every variable that has a value in it, so per-call variables and their values pile up and are never collected. The documentation states the rule directly — create context variables at the top module level, never in closures. Module level also gives you one declaration site where the name, the type and the default are stated once for every reader.
code
python · 18 linesimport contextvars
def archive(chat_id):
# Wrong: a brand-new variable on every call.
current = contextvars.ContextVar("current_chat", default="none")
print("before set:", current.get())
current.set(chat_id)
print("after set: ", current.get())
archive("chat-1")
archive("chat-2")
# before set: none
# after set: chat-1
# before set: none <- the previous call's value is unreachable
# after set: chat-2go deeper
Just learn the habit: declare each context variable once at module level, like a constant, and import it where it is needed. Never build one inside the function that uses it.
Explain the mechanism — the variable object is the key and the name is only a label — and predict the symptom: a value set in one call is simply not there in the next, with no exception raised.
Add the lifetime argument. A context keeps a strong reference to every variable set in it, so per-call variables leak in proportion to traffic; recognise that shape when memory climbs with request rate.
Set the policy: the set of ambient variables a system carries should be small, enumerable and declared in one place, with libraries exposing helpers rather than raw variables so no consumer invents dynamic ones.
This is a small rule with two independent justifications, one about semantics and one about lifetime. Candidates usually know that the rule exists; the interesting part is being able to say *why*, because each reason predicts a different symptom. ## Identity, not name A `contextvars.ContextVar` is the key under which a value is stored in a context. **The key is the object, compared by identity.** The `name` string passed to the constructor is documentation: it shows up in `repr(var)` and in the message of the `LookupError` raised by a failed `get()`, and nothing looks a variable up by it. Two `ContextVar` objects constructed with the same name are completely unrelated — writing to one is invisible through the other. That makes the **per-call construction bug** quiet rather than loud. The code reads perfectly: build the variable, set it, read it back, and within one call everything behaves. The next call constructs a different object, finds no entry for it, and returns the default. Nothing raises. What you observe is an ambient value that is never there when a *different* function reads it — which is the exact opposite of the point of the mechanism. The same reasoning covers a subtler variant: a `ContextVar` created per instance in `__init__`. - If a handful of long-lived objects each want their own ambient slot, that is defensible. - If the objects are created per request, per job or per row, it is the closure bug wearing a different hat. ## Lifetime The second reason is the one the standard-library documentation calls out explicitly: context variables should be created at the top module level and never in closures, because **`Context` objects hold strong references to context variables**, which prevents the variables from being garbage collected. Every time a per-call variable is written to, the current context gains an entry keyed by that object. The entries survive the call — nothing removes them unless something resets the token — so a long-lived context slowly accumulates one entry per call, each pinning a variable object and the value it holds. You have built a leak whose growth rate is your request rate. This is one of the few places in Python where an object being *unreachable from your code* is not enough for it to go away, because the context itself is the last remaining reference. It is also why the leak is hard to read in a heap dump: the retained objects look like ordinary small objects held by a mapping you never wrote. ## The interview shape This rarely gates an offer, which is why it is a curiosity rather than a screening question. It comes up in two ways: - **As a code-review question** — "what is wrong with this helper?" over a function that constructs its own variable — where the expected answer is the identity argument. - **And as a diagnosis question**, in the form of a service whose memory climbs in proportion to traffic and whose retained objects are held by the context, where the expected answer is the documented lifetime rule. ## The habits that follow - Declare each variable **once at module level**, next to the code that owns the concept. - Give it a name string matching the Python identifier, so tracebacks are readable. - Give it a `default=` if reads happen on paths that must not raise. - Keep the number of variables small and fixed — the set of ambient values a system carries should be enumerable in one file, not generated dynamically. - If a library exposes ambient state, keep the variable private and export a helper or a context manager that pairs the write with its restore, so no consumer can construct variables in a loop or forget to unwind one. - And if you genuinely need a *dynamic* set of ambient keys, put a single module-level variable holding an immutable mapping, and replace the whole mapping with `set()` — one key, one entry, no growth. ## A diagnostic tip One diagnostic tip closes the loop between the two reasons. Because the name is only a label, the fastest confirmation of the closure bug is to **print the variable itself rather than its value**: the `repr` includes the object's address, so two calls printing two different addresses proves the keys differ, while the identical name in both lines is exactly what made the bug invisible in the source.
- If lookup is by identity, what is the name argument to contextvars.ContextVar for?Debugging only. It appears in the variable's `repr` and in the message of the `LookupError` a failed `get()` raises, which is why the string should match the Python identifier you bind it to. It is never used to find a value, and two variables sharing a name share nothing else.
- Is a ContextVar created per object instance also a problem?It depends on how many instances exist and how long they live. A few long-lived services each owning an ambient slot is fine. One per request, job or row reproduces the closure problem exactly: unrelated keys, values invisible across instances, and entries the context keeps alive.
- How would you expose ambient state from a library without letting consumers misuse it?Keep the variable module-level and private, and export a helper — ideally a context manager — that performs the write and guarantees the restore. Consumers then cannot construct variables dynamically, cannot forget the unwind, and you keep the freedom to change the representation later.
saying these in an interview costs you the question
- Thinks two ContextVars with the same name share a value
- Constructs a ContextVar inside the request handler
- Believes the name string is the lookup key
- Assumes unused context entries are collected automatically
- Generates one ContextVar per dynamic key