skip to content

Why would a WeakValueDictionary cache in a billing run show a near-zero hit rate?

level: seniorimportance: should knowfreq 24%

answer

  1. The container promised never to own anything
  2. Ask who else is still holding the value
  3. The loop variable is the whole lifetime
  4. Insert count high, length near zero
  5. Add an owner, not a bigger cache

basics

~20 s

Because nothing else holds the cached values. A WeakValueDictionary only retains an entry while some other strong reference exists, so if each record is used and dropped inside one loop iteration the entry is gone before the next lookup.

solid answer

~50 s

The cache is doing exactly what it promises: it never owns its values. In a batch loop that loads a customer, processes it and moves on, the only strong reference is the loop variable, so the entry is dropped the moment that name is rebound — and the next lookup misses. Confirm it cheaply by logging `len(cache)` against the number of insertions, or by watching the size stay near zero while insert counts climb. The fix is an ownership decision, not a weakref tweak: keep a small strong-reference tier for the hot working set — a bounded structure the run itself owns — and let the weak mapping serve as the identity index behind it, so an object already in use is shared rather than duplicated. If entries must survive without a strong owner, a weak-valued mapping is simply the wrong container and a plain dict with explicit eviction is the honest choice.

code

python · 21 lines
python
import weakref


class Customer:
    def __init__(self, cid):
        self.cid = cid


cache = weakref.WeakValueDictionary()
hits = 0

for _ in range(2):
    for cid in ("c1", "c2"):
        found = cache.get(cid)
        if found is None:
            found = Customer(cid)
            cache[cid] = found
        else:
            hits += 1

print(hits, len(cache))   # 0 0

go deeper

for a junior

Focus on the rule behind the symptom: a weak-valued mapping keeps an entry only while something else holds that value, so a value created and dropped in one loop iteration is never there next time.

for a middle

Explain the lifetime step by step — loop variable rebound, refcount to zero, entry removed — and show the two-counter instrumentation that distinguishes eviction from a cache that was never populated.

for a senior

Drive to the ownership fix rather than a container tweak: an explicit bounded working set owned by the run, with the weak mapping kept as the identity index, and the judgement to delete the cache outright when the workload has no reuse.

for a principal

Own the policy question across services: which caches are allowed to retain, what bounds them, and how the team reviews for the mirror-image failure where an unnoticed strong holder turns a self-cleaning index into unbounded growth.

This is the most common production disappointment with weak references, and it usually surfaces in a batch job. Picture a subscription-billing run whose loader memoises customer records in a mapping shared across calls as a default argument: ```python import weakref _CACHE = weakref.WeakValueDictionary() def load_customer(cid, cache=_CACHE): hit = cache.get(cid) if hit is not None: return hit record = fetch_customer(cid) cache[cid] = record return record ``` It reviews cleanly, it shipped in a three-week release train, and the hit-rate metric added afterwards reads about one percent. ### Why it misses A `weakref.WeakValueDictionary` holds its values weakly by definition: an entry exists only while some *other* part of the process holds a strong reference to that value. The caller's loop is: ```python for cid in due_today: customer = load_customer(cid) invoice(customer) ``` The single strong reference is the local name `customer`. On the next iteration that name is rebound, the refcount drops to zero, the record is deallocated and the mapping quietly removes the entry. By the time the same customer id comes round — on a retry, on a second pass, in the next batch — the entry is long gone. The cache is not broken; it is a cache with no retention policy of its own, and the program supplied no owner. Note the mutable default argument is a red herring here in the usual sense — sharing one mapping across calls is exactly what a memo cache wants — but it is worth naming in review, because a shared mutable default is also how a *strong* dict in the same position becomes an unbounded leak that nobody notices until RSS climbs over the run. ### Confirming it rather than guessing The diagnosis is cheap and does not need a profiler. Count insertions and compare with `len(cache)` at intervals: an insert count in the millions against a length that hovers in single digits is conclusive. A second confirmation is to bind one known record to a module-level name and watch the hit rate for that one id go to a hundred percent while every other id still misses. If you want to see the eviction as it happens, register a callback on a `weakref.ref` to one record and log when it fires — you will see it fire inside the loop, not at the end of the run. ### The fix is an ownership decision There are three honest resolutions, and choosing between them is the actual interview content. **Give the working set a strong owner.** Keep a small, explicitly bounded structure — held by the run itself — containing the records currently in play, and let the weak mapping remain as the identity index. Now a lookup hits while the record is in the working set, and the weak mapping guarantees that two parts of the run handling the same customer share one object rather than building two. This is the arrangement weak values were designed for: the index never extends anyone's life, the owner decides lifetime. **Restructure so the object is genuinely still in use.** If the run processes each customer exactly once, there is nothing to cache and the cache should be deleted. A hit rate near zero sometimes means the workload has no reuse, and the honest response is to stop paying for a lookup that never pays back. **Switch containers.** If entries must survive with no strong owner, you want retention, and retention is a plain `dict` plus a deliberate eviction rule. A weak-valued mapping cannot express "keep this for a while" because it has no notion of a while. ### The mirror-image failure Worth naming, because interviewers often follow up with it: the same design fails in the opposite direction when the values are not as unowned as you think. Add a callback list, an audit trail, or an in-flight index that holds each record strongly, and the weak mapping now never drops anything — the hit rate goes to a hundred percent and memory grows without bound for the whole run. Both failure modes have the same root cause: nobody wrote down who owns the object. A weak mapping does not answer that question; it makes the absence of an answer visible.

  • How would you prove the cache is evicting rather than never being populated?
    Instrument both sides: count insertions and sample `len(cache)` on the same interval. Insertions climbing into the millions while the length stays in single digits proves entries are entering and leaving. As a control, pin one record to a module-level name and watch that single id start hitting every time while the rest still miss.
  • If you add a strong working set alongside the weak mapping, what has the weak mapping bought you?
    Identity. The weak mapping guarantees that any part of the run reaching for a customer already in play gets the same object rather than constructing a second one, which keeps mutations and per-object state consistent — and it never extends anyone's lifetime, so when the owner releases a record the index cleans itself up. The owner sets the policy; the index enforces uniqueness.
  • What is the opposite failure mode for the same cache, and what causes it?
    Unbounded growth. If something else in the process — an audit list, a callback registry, an in-flight index — holds every record strongly, the weak mapping never drops an entry, the hit rate goes to a hundred percent and memory climbs for the whole run. Both symptoms trace to the same root cause: no explicit owner was ever chosen for the objects.

saying these in an interview costs you the question

  • Blames the garbage collector and calls gc.collect to fix it
  • Thinks a weak mapping retains values until memory pressure appears
  • Suggests raising a cache size limit that does not exist here
  • Cannot name which side of the mapping is held weakly
  • Treats a hundred percent hit rate on a weak cache as a success
  • Adds a strong reference everywhere without deciding on an owner

context