skip to content

When does multiprocessing.shared_memory beat a Manager proxy for a 200 MB invoice-bitmap payload?

level: seniorimportance: should knowfreq 35%

answer

  1. A proxy call is IPC, not an index
  2. One process owns the objects, others hold proxies
  3. Shared memory maps bytes with no copy
  4. close() per process, unlink() exactly once
  5. size is a minimum, so carry the length out of band

basics

~20 s

When the payload is large and read-mostly. A Manager proxy pickles a round trip to a server process on every access, so a 200 MB bitmap is copied repeatedly; multiprocessing.shared_memory.SharedMemory maps one block that workers read in place, at the price of owning its lifecycle and layout yourself.

solid answer

~50 s

A `multiprocessing.Manager()` proxy is convenient because it looks like an ordinary `list` or `dict`, but every method call is a **pickled round trip to a separate server process**. For an invoice-PDF renderer handing 200 MB page bitmaps to workers, that is the whole payload serialized, piped and rebuilt per access — the copy dominates the render. `multiprocessing.shared_memory.SharedMemory` instead maps one named block into every process; a worker attaches by name and reads through `shm.buf` as a `memoryview` with **no copy at all**, which is exactly right when the data is large and read-mostly. The costs you take on are real: you own creation and destruction (`close()` per process, `unlink()` exactly once), you get no synchronization, and the block carries no structure — length, shape and dtype must travel out of band, because the allocated `size` may exceed what you asked for and a reader that trusts it will silently read past your data. Use a Manager proxy for small shared coordination state; use shared memory for big buffers.

code

python · 12 lines
python
from multiprocessing import shared_memory

block = shared_memory.SharedMemory(create=True, size=16)
block.buf[:5] = b"hello"

view = shared_memory.SharedMemory(name=block.name)   # attach by name
data = bytes(view.buf[:5])                            # length known out of band
print(data, "requested 16, actual size", view.size)   # size may exceed 16

view.close()
block.close()
block.unlink()      # exactly once, by the owner

go deeper

for a junior

Know that a Manager gives you shared list and dict objects that look normal but are served by another process, and that shared memory is the raw-bytes option used for large buffers.

for a middle

Explain the per-call pickled round trip behind a proxy, the nested-mutation trap, and the basic shared-memory lifecycle of attach, close and unlink.

for a senior

Make the call with numbers: estimate copy cost against payload size and access frequency, design a single-writer-then-read-only phase so no lock is needed, and carry length and shape out of band so a short write cannot truncate silently.

for a principal

Own the systemic tradeoff: zero-copy buys throughput and buys you an OS resource with a lifecycle, a leak class and a corruption class. Decide whether the team should absorb that, or whether the pipeline should move the data differently.

### Two very different answers to "share this" **A Manager proxy** is a client/server design. `multiprocessing.Manager()` starts a helper process that owns real Python objects — `list`, `dict`, `Namespace`, `Lock`, `Event` and friends — and hands every other process a *proxy*. The proxy implements the same methods, but each call serializes the method name and arguments, sends them over a connection to the manager process, waits, and unpickles the reply. The ergonomics are excellent: shared state that behaves like normal Python, works across machines with `multiprocessing.managers.BaseManager`, and is safe by construction because only one process actually mutates the object. The price is a round trip per operation. `proxy[i]` is not an index; it is IPC. And crucially, **a proxy is not deep** — mutating a plain list nested inside a proxied list changes a temporary copy that is thrown away. You have to nest proxies or reassign the whole element. **Shared memory** is the opposite trade. `multiprocessing.shared_memory.SharedMemory(create=True, size=n)` asks the OS for a named region and maps it into the caller. Any process that knows the name can attach with `SharedMemory(name=...)` and reach the same physical bytes through `shm.buf`, a `memoryview` you can slice, assign into, or wrap with `struct` or `memoryview.cast`. Nothing is pickled and nothing is copied. ### The renderer, concretely An invoice-PDF renderer holds 200 MB of decoded page bitmaps, and its font-and-template cache runs at an 83% hit rate, so most tasks want to *read* a buffer that is already there rather than produce a new one. Through a Manager proxy, every worker that touches a page pays serialization of that page plus a pipe transfer plus reconstruction — and that cost is per access, not per worker. Through shared memory, the parent writes the bitmaps once into a block, sends each worker the block's **name** plus an offset and a length (a few dozen bytes through the pool's normal argument path), and the worker attaches and reads in place. The read-mostly, high-hit-rate access pattern is precisely the profile that makes zero copy pay. ### What you take on when you choose shared memory **1. Lifecycle, explicitly.** `close()` detaches *this* process's mapping. `unlink()` destroys the block so no one can attach again, and must be called exactly once, by whichever process is designated the owner. Getting this wrong produces the two classic failures: an orphaned block that survives your program and leaks OS memory until reboot, or an `unlink()` while workers are still attached, after which they read a region that has been destroyed. On POSIX, a resource-tracking helper process cleans up blocks at interpreter exit and prints a warning about leaked shared memory objects, which is why a deliberately long-lived block created in a worker needs the `track=False` argument added in Python 3.13 to avoid being cleaned up out from under you. `multiprocessing.managers.SharedMemoryManager` is the ergonomic middle ground: used as a context manager, it hands out blocks and destroys all of them on exit, which removes the leak class entirely for the common scoped case. **2. No structure.** A block is bytes. The requested `size` is a *minimum*; the OS may hand back more — asking for 16 bytes can yield a `size` of 16384. A worker that reads `shm.buf[:shm.size]` and trusts it therefore reads your data plus trailing garbage, and a worker that assumes a fixed record width when the writer wrote fewer bytes silently truncates the payload instead of raising. Neither failure announces itself; you get a corrupt render, not a traceback. Always carry the real length, and any shape or element type, out of band — in the task arguments or in a small header written into the first bytes of the block itself. **3. No synchronization.** Shared memory gives visibility, not mutual exclusion. If more than one process writes, you must add a `multiprocessing.Lock`. The clean discipline for a pipeline is a single-writer phase followed by a read-only phase, so no lock is needed on the hot path at all. **4. Release views before closing.** `shm.buf` exports a buffer; if a `memoryview` derived from it is still alive, `close()` raises `BufferError`. Drop or release your views first — this bites when a view is stashed on an object that outlives the block. ### Choosing, in one line each * **`Queue`/`Pipe`** — the default. Moderate-size results and work items, copied. Simple and safe. * **`Value`/`Array`** — a handful of machine-typed numbers with a lock, when a counter or a small fixed vector must be shared. * **Manager proxy** — small, structured, frequently-coordinated state where convenience beats throughput; also the only one of these that reaches across machines. * **`SharedMemory`** — one large read-mostly buffer, where a copy per access would dominate; you accept lifecycle and layout responsibility to get zero copy. The interview answer that lands is the one that names the crossover: below a few hundred kilobytes, the copy is cheaper than the complexity, and a proxy or a queue is the right call. Once the payload is tens or hundreds of megabytes and is read far more often than it is written, shared memory stops being an optimization and becomes the only design that works.

  • What is the difference between close() and unlink() on a SharedMemory object?
    `close()` detaches the calling process's mapping — it is per process and every attaching process should call it. `unlink()` removes the block's name from the system so nothing can attach again and the memory is released once all mappings are gone; it must be called exactly once, by whichever process owns the block. Skipping `unlink()` leaks an OS-level object beyond your program; calling it early leaves attached workers reading a destroyed region.
  • Why is mutating a plain list nested inside a Manager list proxy a silent no-op?
    Because the proxy only intercepts calls on the outer object. Fetching an element pickles a copy back to your process, so mutating that copy changes something the manager process never sees, and the change is discarded. You either nest proxies — store a proxied list inside the proxied list — or read the element, mutate it locally, and assign the whole element back through the proxy.
  • A worker reads shm.buf[:shm.size] and gets trailing garbage. What went wrong?
    `size` is the size of the allocated block, not the length of your payload — the operating system may round the request up, so asking for a small block can yield a much larger one. The reader has to know how many bytes are meaningful. Carry the length, and any shape or element type, in the task arguments or in a small fixed header at the start of the block; never infer it from `size`.
  • Where would you still prefer a Manager proxy over shared memory?
    For small, structured, frequently-coordinated state: a status dict, a shared set of processed keys, a configuration namespace, a cross-process `Lock` or `Event`. The per-call round trip is irrelevant at that size, the code stays ordinary Python, and there is no lifecycle to manage. A Manager is also the only one of these mechanisms that works across machines, via a manager server bound to an address.

A Manager proxy is phoning a records office and having the file read to you every time; shared memory is being given a key to the room the file sits in, along with the duty to lock up afterwards.

saying these in an interview costs you the question

  • Treating a Manager proxy access as a cheap local lookup
  • Assuming a nested mutation through a proxy propagates
  • Using shm.size as the payload length
  • Calling unlink() in every worker instead of once
  • Expecting shared memory to provide any mutual exclusion
  • Forgetting that a live memoryview blocks close()

context