When is functools.lru_cache the wrong caching tool for a long-lived Python service?
answer
- One table per process, never shared
- No expiry, no per-key invalidation
- Entries counted, bytes not budgeted
- Invisible at the call site
- Explicit cache object buys the missing levers
basics
~20 sIt is wrong wherever you need expiry, a byte budget, targeted invalidation, per-tenant scoping or a view shared across processes. It offers none of those: one process-local table, entries counted not sized, and cache_clear() as the only eviction you control.
solid answer
~50 s`functools.lru_cache` is a memoization decorator, not a service cache, and the gap shows in five places. It is **process-local**, so a pool of workers holds one warm table each and they cannot agree or invalidate one another. It has **no expiry**, so anything derived from changing state goes stale silently. `maxsize` counts **entries, not bytes**, so a budget in megabytes has to be reasoned about by hand. Invalidation is all-or-nothing via `cache_clear()`; there is no way to drop one key. And it is **invisible at the call site**, which is exactly what makes it pleasant to add and hard to operate - process-global mutable state that also leaks between tests unless something clears it. It fits pure functions over a small, stable key space. Anything with a staleness horizon, a tenant dimension or a memory budget wants an explicit cache object or an out-of-process one.
code
python · 13 linesimport functools
FLAGS = {"thumbnails"}
@functools.cache
def feature_enabled(name):
return name in FLAGS
print(feature_enabled("thumbnails")) # True
FLAGS.discard("thumbnails")
print(feature_enabled("thumbnails")) # True - stale, served from the table
feature_enabled.cache_clear()
print(feature_enabled("thumbnails")) # False, and every other key was dropped toogo deeper
You will not make this call yet, but know that a memoized result can be wrong once the underlying data changes, and that the stored values live inside one process only.
Be able to list what the decorator does not provide: no expiry, no byte budget, no per-key invalidation, no view shared between processes.
Decide per function - purity, key-space size, staleness horizon - and make sure production has a way to bound and clear the cache without a restart.
Own the policy: where memoization is permitted, how it is bounded and observed, how tests reset it, and the threshold at which caching moves into an explicit component or out of the process.
## The decision, stated as questions A memoization decorator is the cheapest cache in Python and the least operable one. Deciding whether it belongs is five questions, and each maps to a property the decorator does not have. **Is the function pure with respect to its keyed arguments?** Memoization answers from the past, so anything reading a clock, a random source, mutable module state, a file or a database can be served an answer that was true once. There is no version stamp, no dependency tracking and no invalidation signal. If the answer can change while the arguments do not, the decorator is already the wrong tool, and this is the failure people discover last because it produces wrong output rather than a crash. **Is the key space bounded, and how large is an entry?** With `maxsize=None` nothing evicts; with a cap, the cap counts *entries*, so a budget in bytes is your arithmetic to do. A cache of large results is a memory commitment that no counter reports in bytes and no allocator pressure ever releases. **What is the staleness horizon?** There is no per-entry lifetime. If the acceptable staleness is "until the process restarts", the decorator fits. If it is "thirty seconds" or "until the config changes", it does not, and bolting a timer onto `cache_clear()` gets you a sawtooth of cold-cache latency plus a scheduling dependency invisible from the call site. **How many processes are there?** Each process imports the module and owns its own table. Ten workers means ten warm-ups, ten copies of the memory, and no way to invalidate the other nine. That is fine for a small derived lookup and untenable for anything a deploy or an operator must be able to purge. **Can you see it and clear it?** `cache_info()` is the entire observability surface, and it only exists if something in your process exports it. Nothing tells you which keys are hot or how many bytes are held. ## The organizational cost A decorator on a module-level function is process-global mutable state created at import time. That has consequences beyond production. Tests become order-dependent: one test's call populates a table the next test reads, so a suite passes in one order and fails in another, and the usual patch is a setup hook that calls `cache_clear()` on every memoized function - a list somebody must maintain. Because the cache is invisible at the call site, a reader debugging a stale value has no local evidence that caching is involved at all. Teams that use it heavily end up with a rule: memoization is permitted on pure functions over small key spaces, must be capped unless the space is provably tiny, and is forbidden on anything that touches I/O. ## What to reach for instead Staying inside the language, the alternatives ladder up. `functools.cached_property` covers per-object derived values, with the value dying with the object and no shared table. Precomputing at startup - building the table once and holding it in a module-level mapping - trades memory determinism for warm-up time and makes the contents inspectable. An explicit cache object passed as a dependency costs a few lines and buys everything the decorator withholds: expiry, per-key invalidation, metrics, per-tenant scoping, and substitutability in tests. And once several processes must agree on a value, or the working set exceeds what you want resident per worker, the cache belongs outside the process entirely, with its own eviction and invalidation story. ## Free-threading, briefly 3.14 makes the free-threaded build officially supported (PEP 779). It does not make a shared cache free: one hot memoized function is a single shared structure that every thread touches, so its bookkeeping is a serialization point, and a very cheap computation can be worth repeating rather than coordinating over. The honest answer at this level is that the shape to prefer - one shared cache, sharded caches, thread-local caches, or recomputation - is a measurement on the real workload, not a rule. ## Saying it well The strong answer resists the framing that there is a single right cache. Name the properties the workload needs - purity, key-space size, staleness horizon, process count, observability - and show `functools.lru_cache` scoring well on the first two and badly on the rest. Then say where the line sits in your codebase, because at this level the question is really about the policy, not the decorator.
- What would you require before approving @functools.cache on a function in a shared codebase?Evidence that the function is pure in its keyed arguments, a bounded key space or an explicit maxsize with the memory arithmetic written down, a way for tests to reset it so the suite does not become order-dependent, and a name or docstring that admits the result may be stale. If any of those cannot be shown, the caching wants to be an explicit object with expiry and per-key invalidation instead.
- When would you keep the memoization but move it out of the decorator?As soon as you need any lever the decorator withholds: expiry, invalidating one key, per-tenant scoping, metrics beyond four counters, or the ability to swap the cache in tests. An explicit cache object passed as a dependency costs a few lines and makes all of that ordinary code. The decorator's virtue - being invisible - is precisely what makes it impossible to operate once the requirements grow.
- How does the free-threaded build change how you would size or shape these caches?3.14 officially supports free-threading, so threads really do run Python in parallel and a single hot memoized function becomes one shared structure they all touch; its bookkeeping is then a serialization point rather than free speed. Depending on the workload the answer may be sharded or thread-local caches, or simply recomputing a cheap value instead of coordinating. Treat it as something to measure on the real traffic rather than a rule to apply.
It is a sticky note on one desk in one office. Fine for a fact that never changes; useless when every desk must be told the fact was revised.
saying these in an interview costs you the question
- Treats lru_cache as a general-purpose service cache
- Assumes one worker's cache is visible to the others
- Thinks maxsize is a memory budget in bytes
- Memoizes a function that reads mutable global state
- Leaves caches warm between tests and blames flakiness elsewhere
- Claims free-threading makes a shared cache automatically faster