A `__getattr__` proxy in a 6-hour nightly inventory sync grows memory without bound — how do you find and fix it?
answer
- Measure before you redesign anything
- The hook only fires on a miss
- The speed trick is the retention
- Instance dict grows one entry per name
- Slots, bounded fields, weak registry, per-batch scope
basics
~20 sSuspect memoisation: because getattr fires only on a miss, the usual speed trick stores each resolved value on the instance, turning every proxy into a permanent cache that pins the fetched data. Confirm with heap snapshots and live-instance counts, then bound it.
solid answer
~50 sFirst confirm what is growing rather than guessing: take `tracemalloc` snapshots at intervals across the run and diff them, and count live proxy instances by type with `gc.get_objects()` or a counter in `__init__`. Two causes dominate. One is the caching idiom — `__getattr__` computes a value and stores it in the instance `__dict__` so later reads skip the hook, which makes each proxy accumulate every name it has ever been asked for and hold the fetched payload alive. The other is retention: a memo dict or registry of proxies keyed by row id, kept for the whole 6-hour run instead of per batch. The fixes follow: give the proxy `__slots__` so it has no instance dictionary and cannot silently cache, restrict any caching to a small fixed field set, keep registries in a `weakref.WeakValueDictionary`, and scope proxies to a batch so they die with it.
code
python · 17 linesclass CachingProxy:
def __init__(self, source):
object.__setattr__(self, "_source", source)
def __getattr__(self, name):
try:
value = object.__getattribute__(self, "_source")[name]
except KeyError:
raise AttributeError(name) from None
object.__setattr__(self, name, value) # hook never fires again
return value
p = CachingProxy({"sku": "A-1", "qty": 5})
p.sku
p.qty
print(sorted(p.__dict__)) # ['_source', 'qty', 'sku'] - the cache is the leakgo deeper
Know the shape of the trick being used: the fallback runs only on a miss, so writing the value onto the instance makes later reads skip it. That stored value stays for the object's whole life.
Explain the mechanics that turn the optimisation into growth: one instance-dictionary entry per name asked for, each keeping its fetched value alive, plus any container that keeps the proxies themselves.
Demonstrate the diagnosis order — reproduce on a slice, diff snapshots, count live instances, find the referrers — and then justify a bounded design such as slots, a fixed cached field set, a weak registry, or per-batch scoping.
Own the wider call: whether dynamic proxies belong in a long-running batch pipeline at all, given that they hide cost, defeat static checking, and make memory behaviour a property of usage patterns rather than of the code.
## Why a `__getattr__` proxy is a memory design, not just an attribute trick A proxy built on `__getattr__` works because the hook fires only when the ordinary lookup misses: the proxy class stays almost empty, so nearly every name misses and gets forwarded to the wrapped object. The hook is a slow path — a Python-level call after a failed search — so the standard optimisation is to memoise: compute or fetch the value once, write it into the instance `__dict__`, and let every later read find it by the fast ordinary lookup, never re-entering the hook. That optimisation is exactly the leak. The instance dictionary is unbounded and lives as long as the proxy does. Each proxy accumulates one entry per distinct name it has been asked for, and each entry keeps whatever was fetched alive. Add a proxy per record in a batch job that walks a large catalogue, and steady-state memory becomes "every field of every row ever touched". ## Establish what is growing before changing anything Resist the urge to rewrite the proxy on suspicion. In a run this long, the sequence that pays is: 1. **Reproduce on a slice.** Feed a bounded portion of the inventory feed and watch whether RSS grows monotonically. Growth that plateaus is allocator behaviour, not retention. 2. **Snapshot and diff.** Take `tracemalloc` snapshots at fixed intervals and compare them; the allocation sites that keep growing between snapshots name the guilty code, not just the guilty type. 3. **Count live objects.** Increment a class-level counter in the proxy's `__init__` and log it per batch, or walk `gc.get_objects()` filtering by type. A count that rises without bound proves the proxies themselves are retained; a flat count with rising memory points at their contents. 4. **Find the referrer.** If the proxies are retained, `gc.get_referrers()` on a sample shows what holds them — usually a module-level memo dict, a per-run registry keyed by record id, a closure captured by a callback, or an exception object whose traceback pins a frame full of locals. 5. **Weigh the payload.** `sys.getsizeof` on a proxy's instance dictionary, plus a look at what the cached values actually are, tells you whether to bound the cache or to stop wrapping whole payloads when only three fields are used. ## The fixes, and what each trades **Give the proxy `__slots__`.** With no instance dictionary, nothing can silently cache into it — the memoisation idiom becomes an error rather than a leak, and every access goes through the hook. That is the strongest guardrail available, and the trade is speed: the forwarding cost is paid on every read instead of once. For a batch job walking each record a handful of times, that is usually the right trade. **Bound what you cache.** If the hook must memoise, memoise a fixed, named set of fields, not whatever the caller asks for. An unbounded key space plus a per-instance cache is a leak by construction. **Weaken the registry.** A registry that exists to return the same proxy for the same record should be a `weakref.WeakValueDictionary`, so an entry disappears once nothing else refers to the proxy. A plain dict keyed by record id retains every record the run has ever seen. **Scope by batch.** Build proxies for one chunk of the sync, use them, drop the container, and let the next chunk start clean. Long-running processes leak through long-lived containers far more often than through cycles. **Do not mint a new wrapper per access.** A `__getattr__` that returns `SomeProxy(value)` for every attribute read multiplies objects invisibly; return the plain value, or cache the wrapper for the handful of names that need one. ## Two correctness issues that ride along Caching into the instance dictionary does not only cost memory, it also **freezes the value**. Once the name is present, the hook never runs again for it, so a refreshed record in the source system is never picked up until someone does `del proxy.name`. If the data can change during the run, the cache is a correctness bug before it is a memory bug. And the hook must still `raise AttributeError` for names it cannot supply. A proxy that answers every probe makes `hasattr` universally true and misleads code that probes for optional hooks with `getattr`, including `copy.deepcopy` and the pickling machinery — which, on a proxy, can even trigger a fetch. Keep the fallback narrow, and if the proxy should expose the wrapped object's dynamic names to `dir()` and interactive completion, implement `__dir__` returning `object.__dir__(self)` plus those names. One structural caveat to state in the interview: special methods are looked up on the type, not on the instance, so a proxy does not forward them through `__getattr__` — they have to be declared on the proxy class explicitly.
- Why does caching a resolved value into the instance dictionary also make it impossible to refresh?Once the name is present, the ordinary lookup succeeds and the fallback is never consulted for it again, so the object keeps serving the first value it ever fetched. Refreshing requires deleting the cached attribute, which means the cache needs an invalidation story. If the source data changes during a long run, that is a correctness bug independent of the memory one.
- How do you keep `dir()` and interactive completion useful on such a proxy?Implement `__dir__` on the proxy class, returning `object.__dir__(self)` for the real attributes plus the dynamic names the fallback can serve. `dir()` sorts and de-duplicates whatever you return, so a plain list is fine. Without it the object looks empty to introspection, because the dynamic names exist only as a lookup failure the hook happens to handle.
- The proxy count is flat but memory still climbs. What now?Then the proxies are not retained; their contents or something else is. Compare `tracemalloc` snapshots to find the growing allocation site, and check whether each proxy's cache keeps expanding for a long-lived set of proxies, whether fetched payloads are being appended to a module-level list or log buffer, and whether an exception with a traceback is holding a frame full of batch locals.
The hook is a courier who is only called when a parcel is not already on the shelf, so the obvious speed-up is to leave every parcel on the shelf after the first delivery — and the shelf is the warehouse you eventually run out of.
saying these in an interview costs you the question
- Blames the garbage collector without measuring anything
- Calls gc.collect() in a loop and calls it fixed
- Thinks caching into the instance dict is free
- Assumes the fallback fires again after the value is cached
- Uses a plain dict registry keyed by record id forever
- Expects special methods to forward through the hook