skip to content

@lru_cache on a method makes a fraud-scoring service grow in memory for days. Why?

level: seniorimportance: should knowfreq 45%

answer

  1. The decorator wrapped the class attribute
  2. One cache, every instance shares it
  3. The first argument is part of the key
  4. Strong reference, so nothing is collected
  5. currsize climbs while hits stay flat

basics

~20 s

The cache lives on the class, not the instance, and self is part of every cache key. So the cache holds a strong reference to every object it has ever been called on, and with an unbounded cache neither those objects nor their results are ever released.

solid answer

~50 s

Decorating a method wraps the *function on the class*, so there is exactly one cache shared by all instances, and `self` is the first element of every key. The cache therefore holds a strong reference to each instance it has seen: those objects can never be collected while the entry lives, and with `functools.cache` or `maxsize=None` no entry is ever evicted, so the process grows for as long as it creates scorer objects. The second half of the leak is key cardinality — caching on something like an account identifier means the number of distinct keys grows with traffic, not with the code. Diagnose it by reading `cache_info()`: `currsize` climbing without a matching hit rate is conclusive, and `tracemalloc` snapshot diffs or `gc.get_referrers` will point at the cache tuple. The fixes are to move the cache to a module-level function that takes only the values it needs, use `functools.cached_property` for per-instance derived values, or keep the method cache but give it a finite `maxsize`.

code

python · 18 lines
python
import functools
import gc
import weakref

class Scorer:
    @functools.lru_cache(maxsize=None)
    def score(self, account_id):
        return len(account_id)

s = Scorer()
ref = weakref.ref(s)
s.score("acct-1")
del s
gc.collect()
print("still alive:", ref() is not None)
Scorer.score.cache_clear()
gc.collect()
print("released:", ref() is None)

go deeper

for a junior

Know that a decorator on a method applies to the shared function on the class, and that whatever a cache stores stays in memory until something evicts or clears it. That much already explains why memory grows.

for a middle

Explain that self becomes part of the cache key, so the class-level cache holds a strong reference to every instance, and that maxsize=None never evicts. Be able to name the alternative of caching a module-level function over plain values.

for a senior

Show the diagnosis path end to end: read cache_info() for currsize against hits, diff tracemalloc snapshots, and prove retention with a weak reference that only dies after cache_clear(). Then choose the fix that matches the object's lifetime.

for a principal

Own the policy: which caches a service is allowed to hold, how each is bounded against a measured key-cardinality budget, how they are cleared on configuration or model reload, and how their hit rates reach the same dashboards as the rest of the service.

## The shape of the bug ```python class Scorer: @functools.cache def score(self, account_id): ... ``` This reads as "cache this scorer's results". It is not what happens. Decorators run in the class body on the plain function, so the cache is created once, attached to the class attribute, and shared by every instance forever. To distinguish one instance's results from another's, the wrapper puts `self` into the key like any other argument — which means the cache dictionary holds a strong reference to that instance. Two consequences compound: 1. **Instances never die.** A scorer built per request, per model version, or per configuration reload stays reachable through the cache. Everything it holds — loaded rules, feature tables, buffers — stays with it. 2. **Nothing evicts.** With `functools.cache` or `maxsize=None` there is no eviction at all, so the dictionary only grows. Even a bounded `maxsize=N` pins up to N instances, which matters when each is large. And because the instance is part of the key, the hit rate is worse than expected too: a fresh scorer object shares nothing with the previous one, so a new instance starts cold while the old one's entries linger, paying memory for results that can never be hit again. ## The second leak: unbounded key space Even on a module-level function, `maxsize=None` is only safe when the set of distinct arguments is small and bounded by the code rather than by traffic. Caching on an account identifier, a request id, a timestamp or a user-supplied string is unbounded by construction: every new caller adds an entry that will never be evicted and, in a fraud-scoring service, will almost never be looked up again. The cache then behaves as a log of everything the service has ever seen, which is a leak with a friendly name. ## Diagnosing it Start with the cache's own instrumentation, because it is free and unambiguous. Expose `cache_info()` on an admin endpoint or log it periodically: `currsize` growing monotonically while `hits` stays flat is the whole story. A cache whose `currsize` tracks process RSS is the cause, not a symptom. When you do not yet know which cache, take `tracemalloc` snapshots minutes apart and diff them; memoized results and their keys show up under the allocation site of the cached function. To confirm object retention specifically, take a `weakref.ref` to an instance you expect to be freed, drop your own reference, run `gc.collect()`, and see whether the weak reference is still live — then call `cache_clear()` and check again. `gc.get_referrers` on a surviving instance will show the key tuple. Size the remaining cache against a real number rather than a guess: measure the distinct keys seen per hour and set `maxsize` so the 92nd-percentile hour fits, then watch the hit rate to confirm the bound is not throwing away the working set. ## The fixes, in the order to consider them **Move the cache off the method.** Most methods that look cacheable depend on only a few hashable values. Extract a module-level function that takes exactly those values, cache *that*, and let the method call it. No instance enters the key, the cache is a pure function of its inputs, and the bound is about data rather than object lifetime. **Use `functools.cached_property` for per-instance derived values.** Its cache is the instance's own `__dict__`, so it dies with the object and cannot retain anything. **Keep the method cache but bound it.** Sometimes the instance really is long-lived and singular — one scorer for the process lifetime. Then a finite `maxsize` is honest, and the object retention is irrelevant because the object was never going to be collected. **Clear at lifecycle boundaries.** If the cached values are derived from configuration or a model version, call `cache_clear()` on reload. This is also the fix for the subtler correctness bug that rides along with this leak: a cache keyed on the old instance keeps serving results computed from the old rules. What is *not* a fix is switching from `functools.cache` to `lru_cache(maxsize=100_000)` and calling it bounded. That is still hundreds of thousands of retained keys and results, and it converts a fast leak into a slow one.

  • Why does the hit rate also suffer when a per-request object is cached this way?
    Because the instance is part of the key, a new object shares no entries with the previous one — every request starts cold. The old entries stay resident but are unreachable by any future call, so you pay full memory for a cache that can only ever miss. That combination, growing `currsize` with a near-zero hit rate, is the clearest signal that `self` is in the key.
  • If the object is a process-lifetime singleton, is caching its method now safe?
    The retention concern disappears, since the singleton was never going to be collected. What remains is key cardinality and staleness: bound `maxsize` against the distinct arguments you actually see rather than leaving it unbounded, and call `cache_clear()` whenever the configuration or model the results derive from is reloaded.
  • How do you prove a cache is what is holding an object rather than something else?
    Take a `weakref.ref` to the object, drop every reference you control, call `gc.collect()`, and check whether the weak reference is still live. If it is, call the cache's `cache_clear()`, collect again, and re-check. If it dies at that point the cache was the holder; `gc.get_referrers` on the object before clearing will show the key tuple directly.

saying these in an interview costs you the question

  • Thinks each instance gets its own method cache
  • Says self is excluded from the cache key
  • Believes a large maxsize counts as bounded
  • Blames the garbage collector for not collecting reachable objects
  • Adds a cache without any way to read its hit rate
  • Proposes weakening the key by hashing self by identity only

context