skip to content

A shared collections.ChainMap holds an invoice renderer's config — why do per-invoice overrides leak into later invoices?

level: seniorimportance: should knowfreq 30%

answer

  1. reads scope, writes do not
  2. one shared front mapping, many callers
  3. the override outlives the invoice
  4. push a fresh layer per unit of work
  5. new_child returns a new chain, mutates nothing

basics

~20 s

Every ChainMap write lands in its first mapping, and a shared instance shares that mapping, so one invoice's override is still there for the next. Give each invoice its own layer with new_child(), which returns a new ChainMap.

solid answer

~50 s

`collections.ChainMap` routes every write — assignment, `pop`, `setdefault`, `update` — into `maps[0]`, and a process-wide instance has exactly one of those. So a per-invoice override such as a watermark flag is written into the shared front layer and simply stays there, and the next invoice through the renderer inherits it and stamps the watermark a second time. Nothing resets it, because a ChainMap has no notion of scope on its own. The fix is `new_child(overrides)`: it returns a **new** ChainMap layering a fresh dict over the same shared layers, so the per-invoice mutations live in an object that dies with the request. Also remember that layers are live references, so a background config reload can change values under a render in flight — flatten with `dict(cfg)` at the start of a unit of work when you need a stable snapshot.

code

python · 13 lines
python
from collections import ChainMap

base = ChainMap({"dpi": 150, "watermark": False})   # shared, built once

def render(invoice, **overrides):
    cfg = base.new_child(dict(overrides))   # a NEW ChainMap; base is untouched
    cfg["rendered"] = invoice              # the write lands in the child layer
    return cfg

first = render("INV-1", watermark=True)
second = render("INV-2")
print(first["watermark"], second["watermark"], base.get("rendered"))
# True False None

go deeper

for a junior

Take away the one rule that causes this bug: a ChainMap reads through every layer but writes only into the first one. A shared instance therefore shares everything anyone writes to it.

for a middle

Explain the fix mechanically: new_child() returns a new ChainMap with a fresh front dict and references to the same deeper layers, so per-request writes die with the request and nothing is copied.

for a senior

Diagnose it: state that persists across units of work, writes to a process-wide object, and del failing with KeyError on a deeper key. Then decide where to snapshot with dict(cfg) so a live reload cannot change values mid-render.

for a principal

Set the policy — the shared chain is mutated only at startup, every request gets its own layer, and configuration crosses the boundary into workers as a snapshot so behaviour is reproducible from a logged input.

## The failure, precisely The renderer builds one `ChainMap` at import time — file defaults at the bottom, environment above them — and the seventeen services in the render dependency graph each reach for that module-level object to read `dpi`, `locale`, `watermark` and the rest. Somewhere in the fan-out, one service applies a per-invoice override by *writing* to it: `cfg["watermark"] = True`. `ChainMap.__setitem__` writes into `maps[0]`. That mapping is shared with everything else in the process. The invoice that asked for a watermark gets one; so does the next invoice, and the one after that, because nothing ever removes the entry. The visible symptom is a duplicated side effect — the same stamp drawn on documents that never requested it — and it is intermittent in a way that makes it look like a concurrency bug, because whether an invoice is affected depends only on whether some earlier invoice happened to set the flag. The deeper mistake is treating a ChainMap as if it scoped writes the way it scopes reads. It does not. Reads traverse all layers; writes are pinned to one. ## Why `del` does not save you The instinctive patch is to write the override, render, then delete it. Two problems. First, if the key exists in a deeper layer, `del cfg[key]` raises `KeyError` — deletion only ever consults the first mapping, and a ChainMap cannot mask a deeper key. Second, even where it works, the write/delete pair is not exception-safe and not concurrency-safe: an exception between them leaves the flag set, and two invoices rendering concurrently in threads interleave on the same front dict. A `try/finally` narrows the first hole and does nothing about the second. ## The fix: a child layer per unit of work `new_child(m=None)` returns a **new** `ChainMap` whose `maps` list is `[m or {}] + self.maps`. The receiver is not modified; the shared layers are shared by reference, so no data is copied. Every per-invoice write lands in that fresh front dict, and when the render finishes and the object is dropped, the overrides go with it. Reads still fall through to the shared defaults exactly as before, so the seventeen consumers need no change beyond receiving the child instead of reaching for the module global. That last clause is the real remedy: **stop letting call sites reach for the shared object**. Pass the per-invoice ChainMap down explicitly, or bind it to a `contextvars.ContextVar` so concurrent renders each see their own. `parents` gives any layer a read-only view of everything it is shadowing if a service genuinely needs the un-overridden value. ## Liveness is the second hazard A ChainMap holds references, not copies. If a config-reload path swaps values into the defaults dict mid-render, a render that read `dpi` at step 3 and again at step 11 can see two different numbers, producing a document that is internally inconsistent. When a unit of work must be self-consistent, take a snapshot at its boundary: `dict(cfg)` flattens the distinct keys into an ordinary dict, or `cfg.copy()` gives a new ChainMap with a fresh copy of the front layer and shared references to the rest — note that `copy()` is *not* a deep snapshot of the deeper layers. Which you want depends on whether you are isolating writes or isolating reads. ## Diagnosing it in the wild The tell is state that persists across units of work in an object nobody thinks of as mutable. Three moves find it quickly: * **Assert the front layer is empty** at the end of each render in a debug build — the shared instance's front mapping should never accumulate keys. * **Log the layer that supplied a value** when behaviour looks wrong: walk the mappings and report the index of the first one containing the key. Because layers stay separate objects, provenance is recoverable, which is one of the reasons to use a ChainMap at all. * **Grep for writes to the shared name.** Reads of a config object are everywhere; writes should be countable on one hand, and every one of them is a candidate. ## The rule to carry away A `ChainMap` is a stack of read-only context plus one writable scratch layer. Whoever owns the scratch layer owns everything written through the chain. Make that layer belong to the narrowest scope that does the writing — per request, per invoice, per frame — and mutate a process-wide chain only during startup, before anything concurrent can observe it.

  • Why is writing the override and deleting it in a finally block still the wrong fix?
    It keeps the mutation on the shared front mapping, so two concurrent renders still interleave on it, and `del` raises `KeyError` whenever the key exists only in a deeper layer — you cannot unset an override that shadows a default, only overwrite it. A per-render child layer removes both problems by never touching the shared object.
  • How would you give each concurrent render its own layer without threading the object through 17 call sites?
    Bind the per-render ChainMap to a `contextvars.ContextVar` set at the entry point; each thread or task reads its own value, and the shared base stays read-only. Explicit passing is still clearer where the call graph allows it — the context variable is the pragmatic option when the graph is deep and pre-existing.
  • What is the difference between dict(cfg) and cfg.copy() for a ChainMap?
    `dict(cfg)` flattens the chain into a plain dict of the distinct keys — a snapshot that is no longer layered and no longer live. `cfg.copy()` returns a new ChainMap with a copy of the first mapping and references to the rest, so writes are isolated but deeper values still change underneath you. Snapshot for read stability, copy for write isolation.

saying these in an interview costs you the question

  • Blames threading before checking where writes land
  • Thinks a ChainMap write updates the layer it read from
  • Believes new_child mutates the receiver
  • Deletes an override that lives in a deeper layer
  • Treats the shared chain as an immutable snapshot
  • Copies all layers when only scoping was needed

context