Capturing frame locals in error reports made a flight-schedule differ leak memory - what is going wrong?
answer
- Follow what the report keeps a handle on
- An exception is not a small object
- Traceback to frame to locals
- Count-bounded is not memory-bounded
- Locals were never a curated set
basics
~20 sThe reports keep exception objects alive. An exception holds its traceback, the traceback holds every frame, and each frame holds its locals, so a count-bounded buffer pins whole schedules. The captured locals also leak secrets.
solid answer
~50 sThere are two distinct costs and they need separate fixes. **Retention**: an exception object references its traceback via `__traceback__`, the traceback references each frame, and a frame references its locals. Storing exceptions - in a recent-errors ring, a retry queue, a dedupe cache - therefore pins every object those frames touched, which for a differ means two entire schedules per failure. Python deletes the `except ... as exc` name at the end of the block precisely to break that chain; stashing the object opts back in. **Disclosure**: rendering locals calls `repr()` on everything in scope, so tokens, connection objects and passenger records land in the report and in whatever store reads it. Fix retention by keeping formatted strings, never exception objects, and by bounding the buffer. Fix disclosure by capturing a small, explicitly chosen context dictionary instead of whole frames.
code
python · 9 linesdef load_schedule(rows):
payload = rows * 1000
raise ValueError("unparsable leg")
try:
load_schedule(["JFK-LHR"])
except ValueError as exc:
frame = exc.__traceback__.tb_next.tb_frame
print(len(frame.f_locals["payload"]))go deeper
Know that an exception object is not just a message: it carries its traceback, which carries the frames, which carry their local variables. Do not store exceptions in long-lived containers.
Explain the retention chain and why Python deletes the except ... as name binding at the end of the block. Be able to say what changes when you render the report to a string instead of keeping the object.
Diagnose it: growth that tracks error count rather than traffic, referrer chains running through frame objects, and a fix that separates retention from disclosure. Argue for chosen context fields over whole-frame capture in deployed environments.
Own the policy that error reports are a data surface with a wider audience than the process. Decide once what a report may contain, bound both its cardinality and its size, and make the development-time convenience explicitly non-default in deployment.
A flight-schedule differ loads two schedules, walks their legs and reports mismatches. When a leg fails to parse it builds a rich error report that attaches each frame's local variables, so an engineer can see the offending leg without reproducing. Resident memory then climbs run after run and never comes back down. Both of that report's problems are worth separating, because they have different fixes. ### Problem one: the report retains the whole world In CPython an exception instance references its traceback through `__traceback__`. A traceback object references, frame by frame, the actual frame objects that were executing. A frame object references its local variables. So holding one exception object holds - transitively - every object that was a local anywhere on the stack at the moment of the failure. For this differ, that stack includes the function holding both parsed schedules. A single retained exception therefore pins two schedules. A "last 500 errors" buffer that stores exception objects is not bounded in memory at all; it is bounded in *count*, which is a very different guarantee. Memory grows with the size of the data each failure happened to be near. This is also why Python deletes the name in `except ValueError as exc:` at the end of the block. That deletion exists to break the exception - traceback - frame cycle promptly rather than leave it to the cyclic collector. The moment you assign the exception somewhere that outlives the handler, you have opted back into the retention and you own it. ```python try: diff(left, right) except ValueError as exc: recent.append(exc) # pins both schedules, indefinitely recent.append(render(exc)) # pins a bounded string ``` ### Problem two: the report contains things the report should not Rendering frame locals means calling `repr()` on every name in scope. That is not a curated set. It includes credentials read from the environment, API tokens held for the schedule feed, session or client objects whose `repr()` embeds a URL with a password, and - in a flight domain - passenger identifiers. Whatever your error store's audience is, it is wider than the process that produced the report, and this is the mechanism by which secrets get there. The standard library will do exactly this for you: `traceback.TracebackException` accepts a flag that captures frame locals, and rendering the result prints each local's value. Note the subtlety - that path **formats eagerly to strings**, so it *fixes the retention problem*: once the locals are text, the frames can be collected. It does nothing about disclosure; it bakes the secret into the string permanently. Two problems, two mechanisms. A third, smaller cost belongs with disclosure: `repr()` of a large object can be enormous and slow. A parsed schedule's repr can be megabytes, so the reports themselves become the memory growth even when the frames are freed, and an unbounded field can push a log shipper over its per-event limit and lose the event entirely. ### What to do instead 1. **Capture context you chose, not context you happened to have.** A small dictionary - run identifier, source file name, leg index, counts - is what an engineer actually needs. It is explicit, small, and reviewable. 2. **Store text, not exceptions.** Render at the boundary where the error is handled and keep only the rendered string. This alone converts unbounded retention into bounded strings. 3. **Bound both dimensions.** Cap the number of retained reports *and* the size of each. Truncate any captured value to a fixed length. 4. **Rate-limit and group.** A differ that fails on every leg of a bad file produces thousands of near-identical reports; deduplicate by exception type and code location before you store. 5. **Keep whole-frame capture for development.** It is a genuinely good debugging tool locally, where the audience is one engineer and the process is short-lived. It is the deployment default that is wrong. ### Confirming the diagnosis Before rewriting anything, prove it. Run the differ with error reporting disabled and see whether the growth disappears; if it does, sample the heap and look at what is reachable from retained traceback and frame objects. The signature is distinctive: object counts that grow in step with the *error* count rather than the request count, and retained objects whose referrer chain runs through a frame rather than through your own caches.
- If exceptions retain frames, why does an ordinary `except` block not leak?Because Python deletes the name bound by `except ... as exc` when the block ends, dropping the last reference to the exception and with it the traceback and frames. That deletion exists specifically to break the cycle promptly. The leak appears only when you copy the exception somewhere with a longer lifetime - a list, a queue, a cache, a closure.
- Does eagerly formatting the locals into strings solve the problem?It solves half of it. Once the locals are rendered to text the frames can be collected, so the retention chain is broken and memory stops growing with the size of the data. The disclosure problem gets worse, not better: the secret is now permanently embedded in a string that will be shipped to your error store. Choose the fields; do not render the frame.
- What would you keep in a report instead of the locals?A small dictionary chosen at the call site: a run or request identifier, the input's name, the index or key being processed, and counts. Everything in it is a scalar, is size-capped, and can be reviewed for sensitivity once rather than trusted per frame. Add the rendered traceback for position and cause, and stop there.
Storing an exception to remember an error is like keeping the whole crime scene instead of the photographs: perfectly informative, and you run out of warehouse.
saying these in an interview costs you the question
- Thinks an exception object is small because its message is short
- Believes a fixed-size buffer of exceptions bounds memory
- Assumes locals contain nothing sensitive worth worrying about
- Says formatting locals eagerly fixes the disclosure problem too
- Blames the garbage collector rather than the retained references
- Keeps whole-frame capture on in production for convenience