functools.cache on an image-thumbnail worker's sizing step keeps growing its memory - how do you diagnose and fix that?
answer
- Nothing ever leaves an uncapped table
- Keys and values are held strongly
- Check hits against currsize over time
- A key that never repeats buys nothing
- Re-key on the parameters, not the payload
basics
~20 sRead cache_info() first: currsize tracking the call count while hits stay near zero means the key space is unbounded, so the table only grows. Key on the stable parameters instead of per-image identity, or cap it with functools.lru_cache(maxsize=N).
solid answer
~50 s`functools.cache` never evicts and holds strong references to both key and result, so it grows for as long as new argument values keep arriving. On a worker that memoizes something keyed by image identity, every image is a new key and no key ever repeats: the table grows one entry per job forever. The diagnosis starts with `cache_info()` sampled from the live process - `currsize` rising in step with `misses` while `hits` stays near zero is proof the cache is pure cost. Contrast the same worker's cached resolution over its 17-service dependency graph: a few dozen distinct keys, a hit ratio near one, bounded and worth keeping. The fix is to re-key on the small stable inputs the result actually depends on - the target dimensions and format, not the image - and to cap anything open-ended with `lru_cache(maxsize=N)`. Periodic `cache_clear()` is a blunt instrument, not a design.
code
python · 11 linesimport functools
@functools.cache
def thumbnail_size(image_id, width):
return (width, width * 3 // 4)
for image_id in range(50_000):
thumbnail_size(image_id, 320)
info = thumbnail_size.cache_info()
print(info.hits, info.misses, info.currsize) # 0 50000 50000go deeper
Know that functools.cache never removes anything, so a function called with a new argument every time keeps adding entries and the memory only goes up.
Be ready to read cache_info() and interpret it: hits against misses tells you whether the cache earns its keep, and currsize tells you what it is holding.
Show the diagnosis in order - counters first, then key design, then a deliberate cap - and mention that entries are counted, not sized, so large values need a byte budget of their own.
Own the operating rule: caches that must be observable and bounded by declaration, and the point at which per-process memoization gives way to precomputation or a shared cache.
## Why an unbounded cache grows without bound `functools.cache` is `functools.lru_cache(maxsize=None)`. With no cap there is no eviction, and the only way an entry leaves is `cache_clear()`. The dictionary holds strong references to the key and the value, which means every argument that ever appeared and every result ever computed stays reachable and therefore uncollectable. So the memory cost is not a function of how expensive the work was - it is a function of **how many distinct argument values the process has seen**. That is the crux of the thumbnail worker. If the memoized function is keyed by something per-image - an identifier, a path, a digest - then every job produces a key that will never be seen again. The cache records a miss, stores a new entry, and gets no hit for it ever. Memory climbs monotonically, roughly linearly with jobs processed, and the process looks like a slow leak that restarts "fix". It is not a leak in the reference-cycle sense: every byte is legitimately reachable from a dictionary that was told to keep everything. ## Diagnosing it Start with the decorator's own instrumentation rather than a profiler, because it answers the question directly. `cache_info()` returns hits, misses, maxsize and currsize; expose it on an admin endpoint or log it periodically, and read the shape over time: - `currsize` climbing in lockstep with `misses`, `hits` near zero: an unbounded key space. The cache is pure memory cost and should be removed or re-keyed. - `currsize` flat at a small number, `hits` far above `misses`: the cache is doing its job and is not your growth. - `currsize` pinned at `maxsize` with a mediocre hit ratio: the working set exceeds the cap and the cache is thrashing; raising the cap may pay, or the key may be too specific. The same worker usually has both kinds. A cached resolver over a 17-service dependency graph sees at most a few dozen distinct keys and settles instantly at a high hit ratio; the per-image function next to it never repeats a key. Comparing the two counters side by side is the fastest way to point at the culprit without attaching anything to the process. Size matters as much as count. `maxsize` counts *entries*, not bytes, so a cache of a hundred entries can hold hundreds of megabytes if the values are large. When the results are byte payloads or decoded buffers, an entry count that looks harmless is not. ## Fixing it **Re-key on what the result depends on.** The strongest fix is usually not a cap but a better key. If the derived value depends only on the requested dimensions and the output format, key on those - a handful of combinations, an enormous hit ratio, and a table that stops growing on the first few jobs. Keying on the image was keying on the *input to the job*, not on the *inputs to the computation*. **Cap what is genuinely open-ended.** Where the key space is legitimately large but repeats within a window, swap `cache` for `lru_cache(maxsize=N)` and choose `N` from the observed working set times the size of an entry. Then the memory is a bounded, declared budget instead of an emergent property of traffic. **Do not key on payloads or per-job objects.** A key holds its object alive. Keying on a request object, an open resource wrapper or a decoded buffer pins that object for the entry's lifetime, which converts a cache into an object retainer. **Treat cache_clear() as a fallback.** A timer that clears the cache does bound memory, but it throws away the hot entries with the cold ones, so the cost returns immediately after each sweep, and it introduces a scheduling dependency nobody sees at the call site. It is a mitigation while you fix the key, not the fix. **Remember the worker count.** Each worker process imports the module and gets its own cache, so the footprint multiplies by process count while the hit ratio is divided among them. A cap that looks fine in one process is that cap times the pool size on the host, and it is one of the reasons a per-process memoization cache stops scaling before an out-of-process cache does. ## What to say in the interview Name the mechanism (no eviction, strong references), name the evidence (`cache_info()` counters over time), then give the ordered fix: better key first, explicit cap second, clearing last, and a note that the budget multiplies across worker processes.
- How would you choose a maxsize rather than guessing one?Estimate the working set: the number of distinct keys seen in the window over which repeats actually happen, then multiply by the measured size of one entry to get a memory budget you can defend. Validate it from cache_info() in production - currsize sitting at maxsize with a poor hit ratio means the cap is below the working set and you are thrashing, while a currsize that never approaches the cap means the key space, not the cap, is bounding you.
- Why is a timer that calls cache_clear() a mitigation rather than a fix?It drops the whole table, hot entries included, so every sweep is followed by a burst of recomputation and a latency spike, and the memory sawtooths rather than staying bounded. It also adds a scheduling dependency invisible at the call site: whoever removes the timer later has no way to know it was load-bearing. Bounding with maxsize or narrowing the key are structural fixes that need no coordinator.
- How does running the worker as a process pool change the memory budget?Each process imports the module independently and therefore owns a separate cache, so the memory is the per-process budget times the number of workers, while each cache sees only its own share of traffic and warms more slowly. That is the point at which a per-process memoization cache stops being the cheap option and a shared out-of-process cache, or precomputing the table at startup, becomes worth its complexity.
It is a filing cabinet with instructions to keep every document and throw none away. If each job files a form nobody ever asks for again, the cabinet is not a cache, it is a warehouse.
saying these in an interview costs you the question
- Says the cache evicts when memory gets tight
- Adds a timer calling cache_clear() and calls it fixed
- Keys the cache on the image payload itself
- Expects the garbage collector to reclaim cached entries
- Assumes maxsize bounds megabytes rather than entries
- Ignores that every worker process holds its own cache