skip to content

When do you implement `__deepcopy__` on a class, and what must it do with memo?

level: seniorimportance: nice to knowfreq 17%

answer

  1. Only when the default copy is wrong
  2. Resources and caches should not be cloned
  3. Cooperate with the traversal's bookkeeping
  4. Register the clone before recursing
  5. Forward memo into every nested copy

basics

~20 s

Implement deepcopy(self, memo) when the default copy is wrong: the object holds a lock, socket or handle that must not be duplicated, or a cache to drop. Register the new object in memo before recursing, and pass memo down to nested copies.

solid answer

~50 s

`copy.copy` looks for `__copy__(self)` and `copy.deepcopy` looks for `__deepcopy__(self, memo)`; without them, the copy module falls back to generic machinery that shallow-copies `__dict__` or reduces the object via `__reduce_ex__`. You define a hook when that default is wrong — the object owns a resource that cannot or must not be duplicated (a lock, connection, open file), holds a cache that should be dropped or shared rather than cloned, or carries an identity field that must be regenerated for the copy. The contract on `__deepcopy__` has two parts: create the new object and **put it in `memo` keyed by `id(self)` before deep-copying any children**, so a cycle back to this object resolves; and **pass `memo` into every nested `copy.deepcopy` call**, so shared references stay shared across the whole walk. Build the instance with `__new__` rather than `__init__` when the constructor has side effects.

code

python · 24 lines
python
import copy
import threading


class ScoreCache:
    def __init__(self, entries=None):
        self.entries = entries if entries is not None else {}
        self.lock = threading.Lock()

    def __deepcopy__(self, memo):
        clone = ScoreCache.__new__(ScoreCache)
        memo[id(self)] = clone
        clone.entries = copy.deepcopy(self.entries, memo)
        clone.lock = threading.Lock()
        return clone


original = ScoreCache({"acct-1": 0.83})
clone = copy.deepcopy(original)

print(clone.entries)                       # {'acct-1': 0.83}
print(clone.lock is original.lock)         # False
pair = copy.deepcopy([original, original])
print(pair[0] is pair[1])                  # True - memo kept them shared

go deeper

for a junior

Know that a class can customise how it is copied by defining hooks the copy module looks for, and that without them an instance is copied attribute by attribute.

for a middle

Name copy and deepcopy and their signatures, and explain that the default path copies dict without calling init, which is why a shared lock or cache survives into the copy.

for a senior

Justify when a hook earns its place — non-duplicable resources, caches to drop, fields to regenerate — and state the memo contract: register before recursing, forward memo to every nested copy.

for a principal

Decide whether types in your system should be copyable at all: immutable value objects and explicit factory methods often beat copy hooks, which spread copying policy across every class that owns a resource.

### What the copy module does without a hook `copy.copy(obj)` and `copy.deepcopy(obj)` both start by dispatching on the object's type. For built-in containers they use dedicated routines; for an ordinary instance with no hooks they build a new instance of the same class **without calling `__init__`** and populate its `__dict__` — shallowly for `copy.copy`, recursively for `copy.deepcopy`. If the type resists that treatment, they fall back to the pickle-style reduction protocol through `__reduce_ex__`, which is why some objects raise `TypeError` when deep-copied at all. For the large majority of classes that default is exactly right and writing a hook would be noise. You define one when the default is *wrong*. ### The three cases that justify a hook 1. **The object owns something that must not be duplicated.** A `threading.Lock`, a socket, a database connection, an open file handle. Deep-copying a lock does not even work — the reduction protocol refuses it — and even if it did, a clone holding a copy of a lock is a bug: two objects that believe they are mutually excluded are not. The hook gives the copy a *fresh* lock and reconnects or drops the connection. 2. **The object holds derived state that should be dropped or shared.** A memoisation cache, an index, a compiled pattern, a large read-only lookup table. Copying it is pure cost; the honest copy starts with an empty cache or keeps pointing at the shared table. 3. **Some field must differ in the copy.** An identity, a version number, a creation timestamp. If every copy must get a new one, the hook is the only place that happens automatically. ### The memo contract, precisely `__deepcopy__(self, memo)` receives the same memo dict the whole traversal is threading through, and it must cooperate with it in two ways: ```python import copy import threading class ScoreCache: def __init__(self, entries=None): self.entries = entries if entries is not None else {} self.lock = threading.Lock() def __deepcopy__(self, memo): clone = ScoreCache.__new__(ScoreCache) memo[id(self)] = clone # 1. register first clone.entries = copy.deepcopy(self.entries, memo) # 2. thread memo down clone.lock = threading.Lock() # fresh, never copied return clone original = ScoreCache({"acct-1": 0.83}) clone = copy.deepcopy(original) print(clone.entries, clone.lock is original.lock) ``` **Register before recursing.** `memo[id(self)] = clone` must happen before any child is copied. If a child (directly or transitively) refers back to `self`, the walk finds the entry and links to the in-progress clone; without it, the recursion never bottoms out. **Pass `memo` down.** Every nested `copy.deepcopy(child, memo)` must forward it. Omitting the argument starts a fresh traversal, which duplicates objects the outer walk already copied and destroys the aliasing the outer copy was carefully preserving — two fields that shared an object in the original end up holding two different copies. **Prefer `__new__` over `__init__`.** Calling the constructor inside the hook re-runs whatever side effects it has — opening files, registering in a registry, validating — and forces you to reconstruct arguments you may not have kept. Allocating with `__new__` and assigning attributes mirrors what the default machinery does. ### `__copy__`, and the option of not copying at all `__copy__(self)` is the shallow counterpart: no memo, no recursion, just "make me a one-level copy of this object". A class that needs a fresh lock will usually want both hooks, since `copy.copy` will otherwise hand out an instance sharing the original's lock object. There is a third, legitimate answer for genuinely immutable value objects: return `self` from both hooks. Copying something that cannot change is pure waste, and the built-in atomics already behave this way. Use it only when the object really is immutable — returning `self` from a mutable object's copy hook turns every "copy" into an alias and produces exactly the bug the caller was trying to avoid. ### Testing a hook Three assertions cover almost all of it: the copy is not the original (`clone is not original`); the resource is fresh rather than shared (`clone.lock is not original.lock`); and a mutation through one is not visible through the other. Add a cycle to the fixture if the type can participate in one, and assert that deep-copying a structure holding the object twice yields one clone referenced twice — that is the check that catches a forgotten `memo` argument, which is otherwise invisible until a subtle aliasing bug shows up in production.

  • What breaks if __deepcopy__ forgets to forward the memo argument to nested deepcopy calls?
    Each nested call starts its own traversal, so objects the outer walk already copied get copied again. Aliasing that the original graph relied on is lost — two fields that shared one object end up with two different copies — and a cycle passing through the hook can recurse until it raises `RecursionError`. Nothing fails at the call site, which is what makes the bug expensive.
  • Is it ever correct for __deepcopy__ to return self?
    Yes, for a genuinely immutable value object. Copying something that cannot change is wasted work, and the copy module already treats built-in atomics that way. It is only correct when the type has no mutating operations at all; returning `self` from a mutable object's hook silently turns every copy into an alias, producing the exact bug the caller was trying to avoid.
  • Why allocate with __new__ inside a copy hook instead of calling the constructor?
    The constructor may have side effects — opening a file, registering the instance, validating input — that should not re-run for a copy, and it may require arguments the instance no longer keeps. Allocating with `__new__` and assigning attributes mirrors what the copy module's default machinery does, which also skips `__init__` when it builds the new instance.

saying these in an interview costs you the question

  • Writes __deepcopy__ without touching memo at all
  • Calls copy.deepcopy on children without passing memo
  • Registers the clone in memo after recursing
  • Calls __init__ inside the copy hook, re-running side effects
  • Returns self from a mutable object's copy hook
  • Defines __deepcopy__ but leaves copy.copy sharing the lock

context