Why does an @functools.lru_cache-wrapped formatter in an invoice-PDF renderer keep returning the previous currency format after a global setting changes, and how do you fix it?
answer
- The cache promises a pure function
- One input never reaches the key
- Same key means same answer, forever
- Read cache_info hits against wrong output
- Widen the signature before clearing globally
basics
~20 sThe cache key is built from the arguments only. The currency setting is ambient state the key never sees, so a call with the same amount hits the entry stored before the change. Fix it by making the setting an explicit parameter, or clear the cache when it changes.
solid answer
~60 s`functools.lru_cache` keys strictly on the arguments it is passed, so any input the function reads but does not receive is invisible to the cache. A formatter that takes an amount and reads the currency or locale from a module global has two inputs and one key: after the global changes, the same amount hits the entry computed under the old setting and the renderer emits the old format. Confirm it with `cache_info()` — hits climbing while output is wrong is the signature, as opposed to zero hits and a rising `currsize`, which is a keying-too-narrow problem instead. The right fix is to widen the key: pass the currency as a parameter so `render(12.50, "USD")` and `render(12.50, "EUR")` are different entries. `cache_clear()` at the point of change is the blunt fallback — it drops every caller's entries and races with threads that may repopulate from the old state. In a 340-case regression pack the same module-level cache is shared across cases, so clear it in per-test teardown or the failures become order-dependent.
code
python · 14 linesimport functools
CURRENCY = "USD"
@functools.lru_cache
def render(amount):
return f"{amount:.2f} {CURRENCY}"
print(render(12.5))
CURRENCY = "EUR"
print(render(12.5))
print(render.cache_info())
render.cache_clear()
print(render(12.5))go deeper
Recall the rule this scenario rests on: a cached function must depend only on its arguments. If the body reads a global, the cache cannot see it change.
Explain the mechanism and the two fixes: put the setting in the signature so the key covers it, or call cache_clear() when it changes, and be able to say why the first is better than the second.
Show the diagnosis path — cache_info() readings, calling wrapped to bypass the cache, reviewing the body for non-parameter reads — and talk about cache lifetime versus setting lifetime, plus the order-dependent test failures this produces.
Own the standard: what a team is allowed to decorate, who reviews the purity claim, and why an unbounded process-lifetime memo under configurable behaviour is a correctness risk rather than an optimisation. Say what you would ban outright.
## The bug class: an input that never reaches the key Memoization is only sound when the function is a function *of its arguments*. The moment the body reads something the caller did not pass — a module global, an environment variable, a process-wide setting, a file, a clock, a row in a database — the cache is keyed on a strict subset of the real inputs. Every hit after that hidden input changes is a wrong answer, delivered fast. ```python import functools CURRENCY = "USD" @functools.lru_cache def render(amount): return f"{amount:.2f} {CURRENCY}" print(render(12.5)) # 12.50 USD CURRENCY = "EUR" print(render(12.5)) # 12.50 USD <- stale hit render.cache_clear() print(render(12.5)) # 12.50 EUR ``` Nothing here is malfunctioning. The wrapper did exactly what it promised: same key, same answer. The defect is the design that put a second input outside the signature. ## Diagnosis The symptom in an invoice renderer is oddly shaped: most output is right, and a subset — usually the rows rendered right after a settings change, or the rows produced by whichever worker warmed the cache first — carries the old format. Three moves narrow it quickly. First, read `cache_info()` around the failing operation. High `hits` with wrong output points at staleness. The opposite reading — `hits` near zero with `currsize` climbing — is a different problem entirely: the key contains something that varies per call, so nothing is being reused and the cache is only costing you memory. Second, call `render.__wrapped__(amount)` and compare. If the undecorated function is right and the wrapper is wrong, the cache is the story and you can stop looking at the formatting code. Third, read the body for every name it uses that is not a parameter. Globals, `self` attributes mutated elsewhere, module-level configuration objects and anything imported at module scope are all candidates. This is a mechanical review step and it finds the bug more reliably than reasoning about it does. ## Fixes, in the order you should prefer them **Widen the key.** Make the hidden input a parameter. It is the only fix that restores the invariant the cache depends on, it makes the cache correct under concurrency for free, and it usually improves the hit rate because entries for different settings coexist instead of thrashing: ```python import functools @functools.lru_cache(maxsize=340) def render(amount, currency): return f"{amount:.2f} {currency}" ``` The caller now passes what it already knew. If the setting is threaded through several layers, that is a signal about the design, not an argument for keeping the global. **Bind the cache to the right lifetime.** If the setting belongs to a render job rather than to the process, a module-level cache is the wrong scope: the cache outlives the thing it is valid for. Move the memo onto an object whose lifetime matches the setting, so a new job starts with an empty one and nothing needs invalidating. **Clear on change.** `cache_clear()` when the setting flips is the blunt instrument. It works in a single-threaded batch. It is unsatisfying because it is global — every caller of that function loses its entries — and it is genuinely unsafe under threads, since another thread can miss, read the old global and repopulate the cache in the window around the change. Treat it as a stopgap you can explain, not a design. **Do not cache.** If the function is cheap and impure, the cache was never buying enough to justify the invariant it needs. ## The test-suite variant A 340-case regression pack that runs in one process shares module-level caches across every case. A case that renders under one currency warms an entry; a later case under a different currency reads it. The failures are order-dependent, they vanish when the failing case is run alone, and they will move around as cases are added — the most expensive kind of flake to chase. The remedies are the same in priority order: parameterise so the entries cannot collide; failing that, clear the cache in per-test teardown so each case starts cold; failing that, do not put a process-lifetime cache under code the suite is meant to vary. ## What to say about it in review A cache is a promise that a function's answer depends only on what you hand it. When you approve a `@functools.lru_cache` on a function, you are approving that promise. The two questions worth asking every time are *what else does this body read?* and *what would make an entry wrong rather than merely old?* — because there is no TTL to bail you out and no size setting that bounds staleness. A stale-cache bug is a correctness bug wearing a performance decorator, and it is discovered in the output, days after the change that caused it.
- What would cache_info() tell you if the cache were hurting rather than lying?A hit count near zero with `misses` and `currsize` both climbing means every call keys differently — some argument varies per call, so nothing is reused and you are paying for storage and hashing with no return. That is a keying or scope problem. Stale results look the opposite: plenty of hits, and output that disagrees with the current state.
- Why is calling cache_clear() when the setting changes an unsatisfying fix?It is global and racy. It throws away every entry for every caller of that function, including entries that were still valid, and under threads another caller can miss, read the pre-change state and repopulate the cache in the window around the clear. It also leaves the unsound invariant in place, so the next hidden input reintroduces the bug.
- What goes wrong when a 340-case regression pack runs in one process against a module-level cached function?Every case shares the same cache, so a case that warms an entry under one setting changes what a later case sees. Failures become order-dependent and disappear when the case is run alone. Parameterise so entries cannot collide; otherwise clear the cache in per-test teardown so each case starts cold.
- When would you keep the cache but move it somewhere else entirely?When the setting belongs to a job rather than to the process. A module-level cache has process lifetime, which outlives a per-job configuration; holding the memo on an object created per render job means a new job starts empty, nothing needs invalidating, and the cached values cannot outlive the state that made them valid.
It is a receptionist who writes down the answer to each question asked. Someone changes the office hours on the wall; the receptionist is still reading from her notes, and she is fast and confidently wrong.
saying these in an interview costs you the question
- Blaming stale output on eviction rather than on the key
- Reaching for a TTL argument that lru_cache does not have
- Assuming cache_clear() removes only the outdated entries
- Thinking a smaller maxsize bounds how stale a result can be
- Expecting the cache to notice that a global changed
- Treating a wrong cached answer as a performance issue