A document-conversion worker keeps a memoryview slice of each payload; why does memory never drop?
answer
- Not a leak; everything is reachable
- The small thing holds the big thing
- Ask what the view still references
- Copy at the lifetime boundary, not inside
basics
~20 sA memoryview keeps a reference to the object it views, so an eight-byte slice pins the entire source buffer for as long as it is retained. Copy the field out with tobytes() and let the payload be freed.
solid answer
~40 sThis is retention, not a leak: a `memoryview` holds no memory of its own but does hold a reference to its exporter, visible as `view.obj`. Slicing a view does not narrow that reference, so a small window kept in a long-lived list pins the whole payload behind it, and resident memory climbs with the number of *completed* jobs. The garbage collector is not at fault — everything is genuinely reachable. Confirm it by comparing `len(view)` with `len(view.obj)` for the retained objects. The fix is to copy at the lifetime boundary with `tobytes()`, and to scope the view itself with `with memoryview(payload) as view:` or an explicit `release()`, which also unlocks the source `bytearray` for resizing.
code
python · 5 linespayload = bytearray(1_000_000)
pinned = memoryview(payload)[8:16] # 8 bytes that hold 1_000_000 alive
copied = memoryview(payload)[8:16].tobytes()
print(len(pinned), len(pinned.obj), len(copied))
pinned.release()go deeper
Remember the one fact underneath this: a memoryview points at another object's memory and keeps that object alive. If you want to keep a few bytes for later, copy them with tobytes() instead of keeping the view.
Explain the mechanics: the view holds its exporter through .obj, slicing does not narrow it, and len() on the view tells you nothing about what is retained. Know that tobytes() is the copy and release() ends the export.
Show the diagnosis path — retention tracking completed work, reachable objects so collection does not help, inspecting len(view) against len(view.obj) — and state the rule you would enforce: no view outlives the function that created it.
Frame it as an interface contract. A module that hands out views is handing out lifetime obligations; decide where copies are mandatory, whether internal buffers are ever exposed at all, and how that shows up in review and in the service's memory budget.
### The mechanism: a view pins its exporter A `memoryview` owns no memory. It holds a pointer into somebody else's buffer plus a **reference to that exporting object**, reachable as `view.obj`. As long as the view is alive, the exporter is alive. Slicing does not narrow this: `memoryview(payload)[8:16]` is an eight-byte window whose exporter is still the whole `bytearray`, so retaining that window retains every byte of the payload. In the scenario, each conversion job reads a document payload into a `bytearray`, takes a small timestamp field as a view — the field carrying the clock-skew artefact worth reporting on — and appends that view to a results list that lives for the whole run. The list looks like it holds a few hundred bytes; it actually holds one full payload per job, released only when the run ends. Over a six-hour nightly pass, that is every document the queue processed, and resident memory grows monotonically all night while the code looks correct. ### Confirming it rather than guessing The tell is a growth curve that tracks the number of *completed* jobs, not concurrent ones, and that is flat across a restart of the queue but climbs again from zero. Confirm it from the objects themselves: for a retained view, `len(view)` versus `len(view.obj)` states the problem in one line — eight bytes retained, a whole payload pinned. Walking the long-lived container and printing `type(x)` and, for views, `len(x.obj)`, turns a vague "memory grows" into a named culprit. The distinguishing signature against an ordinary leak is that nothing is unreachable: the garbage collector is working perfectly, the objects are genuinely referenced, and collecting more often changes nothing. ### The fix: copy out, or scope the view The repair is one call. `view.tobytes()` materializes an independent `bytes` object of exactly the bytes you meant to keep, after which the payload can be freed on schedule. The rule to state in review is blunt: **a view is for the span of a parse; anything that outlives the parse must be a copy.** Zero-copy is a property of a *scope*, not a property you can store. Scope the view explicitly while you are at it. `memoryview` is a context manager, so ```python with memoryview(payload) as view: field = view[8:16].tobytes() ``` releases the export at the end of the block, deterministically, without waiting for the view object to be collected. `release()` does the same imperatively. Releasing has a second benefit: while any export is outstanding, the source `bytearray` refuses to resize and raises `BufferError`, so a worker that reuses one growable buffer across jobs will start failing the moment a view leaks out of its intended scope. ### Related shapes of the same bug The pattern generalizes past `memoryview`. Any object that keeps a reference to a large container while presenting as small has this property — a closure over a big local, an exception traceback holding frames, a slice-like wrapper. Views are the sharpest case because their whole selling point encourages you to pass them around, and because `len()` on the view actively misleads about the retention. There is a converse failure too: copying when you did not need to. Calling `tobytes()` on every subfield inside the parse loop reintroduces exactly the allocations the view was meant to avoid. The discipline is not "always copy" or "never copy"; it is copy **at the lifetime boundary** — where a value stops being a transient piece of a buffer being parsed and starts being a retained result. ### Preventing it rather than finding it again Two habits keep this bug out. First, make the copy explicit and local: a helper that takes the payload and returns `bytes` gives reviewers one place to check, instead of a view escaping through three call frames into a results list. Second, prefer `with memoryview(payload) as view:` over a bare constructor anywhere the payload is large, because the block makes the intended lifetime visible in the source and enforces it even when the body raises. It is also worth deciding, at module level, whether views cross the boundary at all. A parser that returns views is faster and obliges every caller to understand exporter lifetime; a parser that returns `bytes` is slightly slower and obliges nobody. For a queue worker that retains results, the second contract is almost always the right default, with views kept strictly inside the parse. ### What to say in the interview Name the mechanism (`view.obj` keeps the exporter alive), name the observable (retention tracks completed jobs, and the collector is not at fault), name the fix (`tobytes()` at the boundary, `with` or `release()` for scope), and name the rule you would enforce afterwards: no `memoryview` crosses out of the function that made it. That last sentence is the part that shows you have operated one of these services rather than read about them.
- How would you confirm the diagnosis before changing any code?Show that retention tracks completed jobs rather than in-flight ones, then inspect the long-lived container itself: for each retained item report its type and, for a `memoryview`, both `len(item)` and `len(item.obj)`. Eight bytes retained against a multi-megabyte exporter names the culprit outright. Forcing a collection changes nothing, which distinguishes retention from an unreachable-cycle problem.
- The worker reuses one growable bytearray per job and started raising BufferError. How is that related?Same root cause. While any view of a `bytearray` is outstanding, the object refuses length-changing operations and raises `BufferError: Existing exports of data: object cannot be re-sized`. A view that escaped its intended scope keeps the export alive, so the next job's `extend` fails. Releasing views — via `with` or `release()` — fixes both the retention and the resize failure.
- When is keeping a view instead of copying still the right call?Inside the parse, where the source buffer is alive anyway and the copies would be the dominant cost: walking fields of a large frame, dispatching subviews to handlers, or passing a region to code that consumes foreign memory. The rule is lifetime-based — a view may live as long as its exporter was going to live regardless, and no longer.
Keeping the view is like holding one page of a document open with a paperweight and expecting the archive to shred the other thousand pages. Photocopy the page and the archive can go.
saying these in an interview costs you the question
- Assumes a small slice lets the big buffer be freed
- Blames the garbage collector or reference cycles
- Thinks release() copies the data out of the view
- Caches views long-term and calls it zero-copy
- Treats the BufferError on resize as a bug
- Reaches for a manual gc call as the fix