skip to content

How do you decide whether a functools cache is the right fix for a hot function, and how do you bound it?

level: principalimportance: should knowfreq 35%

answer

  1. Three axes before the decorator goes on
  2. Purity, payoff, price
  3. Key cardinality decides the hit rate
  4. Multiply the memory by worker count
  5. No expiry means no TTL requirement allowed

basics

~20 s

Judge it on three axes: correctness (is the function pure and its result safe to share), payoff (does the real key distribution produce hits), and cost (memory per entry times entries times processes). A cache that fails any one of them is not a fix.

solid answer

~50 s

Start with correctness, because a cache that returns stale or shared-mutable results is worse than a slow function: the function must be deterministic in its arguments, free of I/O and clock reads, and must return something callers cannot mutate. Then check payoff against real key cardinality — memoization only pays when a small set of arguments recurs, so estimate distinct keys per unit of traffic and validate with `cache_info()` under production-shaped load rather than a microbenchmark. Then price it: an in-process cache is per-process, so the memory is entry size times `maxsize` times worker count, and I size `maxsize` so the 92nd-percentile hour of distinct keys fits rather than picking a round number. Finally decide the lifecycle: what clears it on a configuration or model reload, and whether the hit rate is exported like any other metric. If the answer needs per-key expiry, coordination between processes, or sharing across hosts, `functools` is the wrong layer and the requirement belongs to a real cache tier.

code

python · 20 lines
python
import functools

@functools.lru_cache(maxsize=8192)
def risk_band(merchant_category, region):
    return f"{region}:{merchant_category}"

def cache_metrics():
    info = risk_band.cache_info()
    total = info.hits + info.misses
    return {
        "hit_ratio": info.hits / total if total else 0.0,
        "fill_ratio": info.currsize / info.maxsize,
    }

risk_band("grocery", "eu")
risk_band("grocery", "eu")
print(cache_metrics())

def on_config_reload():
    risk_band.cache_clear()

go deeper

for a junior

Know that a cache is only correct for a function that always returns the same result for the same arguments, and that it costs memory. Asking whether the arguments actually repeat is already the right instinct.

for a middle

Be able to reason about key cardinality and to pick a finite maxsize rather than the unbounded shorthand. Explain how you would confirm a cache is helping using its own hit and miss counters.

for a senior

Demonstrate that you validate under production-shaped traffic, price the memory across every worker process, wire clearing to configuration reload, and export the counters so that a future memory investigation starts with evidence.

for a principal

Own the boundary: state clearly when a process-local memoization dictionary is the right layer and when the requirement — expiry, cross-process invalidation, survival across restarts — belongs elsewhere, and make that a standard the team applies without you.

## Why this is a judgement call rather than a rule Adding `@functools.cache` is a one-line change with no obvious blast radius, which is exactly why it gets added carelessly. The decorator makes four commitments on the team's behalf: that the function is pure forever, that its result is safe to share by reference, that the memory is available in every process, and that nobody will need to invalidate it. Each of those is a maintenance obligation, and none of them is visible at the call site. ## Axis one: correctness, and it is not negotiable The function must be a mathematical mapping from its arguments to its result. That rules out anything touching a file, a socket, a database, a clock, a random source, or module-level mutable state. It also rules out functions whose result depends on data the arguments merely *reference* — passing an object whose fields change later gives you a key that no longer describes what was computed. The subtler half is the return value. The cache stores the object, so every hit hands out the same instance; a single caller mutating it corrupts all future hits. Cached functions should return immutable results, and enforcing that is a code-review rule rather than something the runtime will check. When a function is *nearly* pure — correct for a while, then stale — the honest question is what "a while" is. `functools` has no expiry mechanism at all. If the answer is "until the model reloads", you can express that with `cache_clear()` at the reload boundary. If the answer is "about thirty seconds", you have a TTL requirement and this is the wrong tool. ## Axis two: payoff, estimated then measured Memoization pays only when arguments repeat. The number that decides it is key cardinality against call volume: if a fraud-scoring path is called with a handful of merchant categories, the cache is nearly free money; if it is called with an account identifier, most calls are unique and the cache is a growing dictionary that never returns anything. Estimate first from the domain — what *is* the key, and is its range set by the code or by traffic? — then measure with `cache_info()` on a shadow or canary carrying real traffic shape. A microbenchmark that calls the function in a loop with the same arguments proves only that dictionary lookups are fast. Watch for `currsize` pinned at `maxsize` with a mediocre hit rate, which means the working set does not fit and you are paying eviction churn for little benefit. Also ask whether caching is the right *kind* of fix. Frequently the honest answer is to hoist the work out of a loop, precompute a table at import time, restructure the caller so the value is computed once and passed down, or make the function cheap enough that the question disappears. A cache is often a way of not fixing an algorithm. ## Axis three: cost, priced across the fleet An in-process cache costs (average entry size) times (`maxsize`) times (worker processes per host) times (hosts). The per-process multiplication is what people forget: sixteen workers means sixteen independent caches, sixteen cold starts after a deploy, and sixteen times the memory. It also means the hit rate you measured in one process is what each process gets, not what the fleet gets — traffic split across workers dilutes repetition. Size the bound against measured distinct keys rather than a round number: take the distinct keys observed per hour, take the 92nd-percentile hour, and set `maxsize` so that fits with headroom. Then treat memory as a budget the cache draws from, alongside everything else the process holds. And prefer bounded to unbounded by default: `functools.cache` is appropriate when the key space is closed by the code — an enum, a config table, a small recursion — and a liability the moment a key can come from a request. ## Axis four: lifecycle and visibility Decide up front what clears the cache and who can see it. A cache derived from configuration or a deployed model needs `cache_clear()` wired to the reload path, or the process must be recycled on change. Its `hits`, `misses` and `currsize` should be exported to the same place as every other metric, so that an eventual memory investigation starts with data instead of a code search — and so that a cache that stopped paying can be deleted with evidence. ## Knowing when to leave the layer The boundary is clean: `functools` gives you a process-local, in-memory, non-expiring dictionary with LRU eviction. The moment the requirement includes per-key time-to-live, invalidation broadcast to other processes or hosts, a cache that survives a restart, or capacity larger than one process should hold, this is not the tool, and stretching it — a hand-rolled timestamp check inside the cached value, a cache cleared on a timer thread — produces something with all the failure modes of a real cache tier and none of its guarantees. Say so, and move the requirement to the layer that owns it.

  • A team wants a five-minute TTL on a memoized function. What do you tell them?
    That `functools` has no expiry, and the two ways people fake it — storing a timestamp inside the cached value and re-checking, or clearing the whole cache on a timer — give you either a cache that reruns its own validation on every hit or one that drops the entire working set on a schedule. If the requirement is genuinely time-based invalidation, it belongs in a cache layer built for it, and this decorator should not be stretched to imitate one.
  • How does running sixteen worker processes change the decision?
    It multiplies the memory by sixteen and dilutes the hit rate, because each process sees only its share of the traffic and each starts cold after a deploy. A cache that looked worthwhile at one process may not clear the bar at sixteen. It also means there is no single place to clear: a configuration change must reach every process, which usually makes process recycling the more honest invalidation mechanism.
  • What would make you remove a cache that is already in production?
    A hit ratio that does not justify its resident size, a `currsize` that grows with traffic rather than settling, or the discovery that the underlying function is no longer pure — a new branch that reads configuration, for example. Removal is cheap because the decorator is one line, and I would rather delete it and re-measure than keep a cache nobody can explain.

saying these in an interview costs you the question

  • Adds a cache without estimating key cardinality
  • Treats functools.cache as a safe default everywhere
  • Forgets the cache is duplicated per worker process
  • Fakes a TTL by clearing the whole cache on a timer
  • Measures the win with a same-argument microbenchmark
  • Has no plan for invalidation on configuration reload

context