skip to content

__copy__ and __deepcopy__

copy.copy and copy.deepcopy ask the object first: __copy__ and __deepcopy__ override the default, and the memo dict is what stops a cyclic graph recursing forever. Asked when a deep copy duplicates a lock.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How does a class control what `copy.copy()` and `copy.deepcopy()` return?

level: juniorimportance: must knowfreq 40%

answer

  1. The module asks the object first
  2. Two separate hooks, two functions
  3. One takes an extra dict argument
  4. No hook means the pickle protocol
  5. Return the finished object, not state

basics

~10 s

Define copy(self) for copy.copy() and deepcopy(self, memo) for copy.deepcopy(). Each returns the finished new object, so the class itself decides which attributes are duplicated, which are shared, and which are rebuilt.

solid answer

~40 s

Both functions ask the object before doing anything generic. `copy.copy()` looks for `__copy__` on the type and calls it with no arguments; `copy.deepcopy()` looks for `__deepcopy__` and calls it with the memo dict for the current copy pass. Whatever the hook returns *is* the copy — the module does no further work on it, so the hook owns the whole object and usually builds it with `cls.__new__(cls)` so `__init__` never runs. If neither hook exists, both functions fall back to the pickle protocol: `copyreg.dispatch_table`, then `__reduce_ex__(4)`, rebuilding from that description without ever producing bytes. That fallback is why a class which already customizes pickling normally copies correctly for free, and why you reach for these hooks only when copying must differ from pickling.

code

python · 27 lines
python
import copy
import threading


class Stage:
    def __init__(self, name, counts):
        self.name = name
        self.counts = counts
        self.lock = threading.Lock()

    def __copy__(self):
        new = Stage.__new__(Stage)
        new.__dict__.update(self.__dict__)
        return new

    def __deepcopy__(self, memo):
        new = Stage.__new__(Stage)
        memo[id(self)] = new
        new.name = self.name
        new.counts = copy.deepcopy(self.counts, memo)
        new.lock = self.lock
        return new


s = Stage("parse", {"hits": 83})
d = copy.deepcopy(s)
print(d.counts is s.counts, d.lock is s.lock)

go deeper

for a junior

Be ready to name both hooks and their signatures without hesitation: __copy__(self) and __deepcopy__(self, memo), and that each returns the finished new object rather than a description of it.

for a middle

Explain the lookup order — hook first, then copyreg.dispatch_table, then __reduce_ex__(4) — and why the hooks allocate with cls.__new__(cls) and use type(self) so subclasses copy back as themselves.

for a senior

Show the judgment about when not to write a hook: if pickling is already customized the fallback usually gives correct copies, and every hand-written hook is a second constructor that drifts out of step with the real one.

for a principal

Own the policy question of whether cloning objects is the right pattern at all in a codebase, versus explicit factory functions that take the shared collaborators as parameters and leave nothing to a copy hook's judgement.

## The copy module asks the object first `copy.copy()` and `copy.deepcopy()` are not magic reflection over `__dict__`. They are dispatchers, and the first substantive thing each one does after handling a handful of built-in types is ask the object how it wants to be copied. The two hooks are the answer to that question. **`__copy__(self)`** is called by `copy.copy()`. It takes no arguments beyond `self` and returns the shallow copy. **`__deepcopy__(self, memo)`** is called by `copy.deepcopy()`. It receives the *memo* dict belonging to the current deep-copy pass and returns the deep copy. In both cases the return value is final. The module does not inspect it, fix it up, or copy anything further; the hook is responsible for producing a complete, usable object. ## The lookup order in CPython 3.14 For `copy.copy(x)`, with `cls = type(x)`: 1. atomic types (`int`, `str`, `tuple`, `frozenset`, functions, `type`…) are returned unchanged; 2. the built-in containers `list`, `dict`, `set` and `bytearray` are copied with their own `copy` method; 3. classes themselves are returned unchanged; 4. **`cls.__copy__`**, if present, is called; 5. `copyreg.dispatch_table` is consulted for a registered reductor; 6. finally `x.__reduce_ex__(4)` is called and the object is rebuilt from what comes back. `copy.deepcopy(x, memo)` follows the same shape: atomic types short-circuit, an internal per-type dispatch table handles lists, dicts, tuples and sets, then **`__deepcopy__`**, then `copyreg.dispatch_table`, then `__reduce_ex__(4)`. Two details are worth carrying into an interview. First, `copy.copy()` looks `__copy__` up on the *type*, so the usual special-method rule applies and an instance attribute of that name is ignored; CPython's `deepcopy`, by contrast, looks `__deepcopy__` up on the *instance*, so an instance attribute there does take effect. Second, defining `__copy__` does nothing for `copy.deepcopy()` and vice versa — they are separate hooks, and a class that needs both must define both. ## Writing the hooks The idiomatic body allocates without running `__init__`, because `__init__` typically demands constructor arguments you no longer have and may open resources you are trying to share rather than recreate: ```python def __copy__(self): cls = type(self) new = cls.__new__(cls) new.__dict__.update(self.__dict__) return new ``` Note `type(self)`, not the class name spelled out. A hook that hardcodes `Stage.__new__(Stage)` silently downgrades every subclass copy to a `Stage`, and that bug survives tests that only ever copy the base class. A `__deepcopy__` differs in three ways: it must register the new object in `memo` before recursing, it must pass `memo` down to every nested `copy.deepcopy(child, memo)` call, and it is where you decide, attribute by attribute, what gets duplicated and what gets shared. ## The pickle-protocol fallback When a class defines neither hook, copying still works for the vast majority of ordinary objects. The module falls through to the same description of the object that pickling uses — a reductor from `copyreg.dispatch_table` if one is registered, otherwise `__reduce_ex__(4)`, which by default reports the class and the instance state — and rebuilds a new object from it. Nothing is serialized: no bytes are produced, no file or socket is involved, and the pieces are reused in memory. Two consequences follow. A class that already customizes its pickled form generally gets matching copy behaviour for free, which is the usual reason not to write these hooks at all. And a class holding something that cannot be reduced — a `threading.Lock`, an open socket — raises `TypeError` from a deep copy even though nobody asked for pickling, because the fallback is the pickle machinery. ## When to define them Define `__copy__` or `__deepcopy__` when copying genuinely differs from pickling: an attribute must be shared rather than duplicated, a cache must be dropped and rebuilt lazily, or the default recursion is prohibitively expensive for a large graph. Otherwise leave them out. Every hand-written hook is a second constructor that must be kept in step with the real one, and the most common bug in this area is a hook that quietly stops copying an attribute added to `__init__` a year later. ## What copying never touches, and what it refuses Some types are deliberately outside the whole mechanism. Modules, classes, functions, methods, tracebacks and frames are treated as atomic: `copy.copy()` and `copy.deepcopy()` return the very same object, because duplicating them would be meaningless rather than helpful. That is why a deep copy of a container full of function references stays cheap, and why the functions in the copy are still the functions you registered. At the other end, when no hook, no registered reductor and no usable `__reduce_ex__` can be found, the module raises `copy.Error` — the module's own exception class. In practice you are far more likely to see a `TypeError` escaping from the reduce fallback, complaining that some attribute cannot be pickled, which is the same fallback path reporting the failure one level down. Knowing which of the two you are looking at tells you immediately whether the problem is the object itself or something it holds.

  • A class defines `__copy__` but not `__deepcopy__`. What does `copy.deepcopy()` use?
    Not `__copy__`. The two hooks are independent, so `copy.deepcopy()` skips straight past it to its own fallback: a reductor registered in `copyreg.dispatch_table`, otherwise `__reduce_ex__(4)`, rebuilding the object from that description and deep-copying the state it reports. The practical effect is that a class needing custom behaviour in both directions must define both hooks.
  • Why do these hooks usually allocate with `cls.__new__(cls)` instead of calling the class?
    Because `__init__` is a constructor for a *new* object, not a copier. It generally requires arguments the hook no longer has, and it often does real work — opening a file, starting a thread, registering the instance somewhere — that must not happen again for a copy. `cls.__new__(cls)` allocates a blank instance of the right class, and the hook then fills in exactly the attributes it wants.
  • What breaks when `__deepcopy__` writes its own class name instead of `type(self)`?
    Every subclass copy comes back as the base class, losing subclass attributes and overridden methods, and no exception is raised to say so. Tests that copy only the base class stay green. Always allocate with `cls = type(self); cls.__new__(cls)` so the hook is inherited correctly.

The copy module behaves like a removal firm that first asks the owner how each item should be handled; only when nobody answers does it fall back to its own standard packing procedure.

saying these in an interview costs you the question

  • Thinks `copy.copy()` falls back to `__deepcopy__`
  • Says the hook returns a state dict, not the object
  • Assumes copying always re-runs `__init__`
  • Believes the fallback serializes the object to bytes
  • Hardcodes the class name, breaking subclass copies
  • Defines only `__copy__` and expects deep copies to change

context

open as a page

Why does `__deepcopy__` receive a `memo` dict, and what must it store there?

level: middleimportance: should knowfreq 33%

basics

~20 s

The memo maps id(original) to the copy already made for it in the current deepcopy pass. Store memo[id(self)] = new before recursing into attributes, so cycles terminate and objects reachable by two paths stay one object in the copy.

open as a page

Why does `copy.deepcopy()` raise TypeError on an object holding a `threading.Lock`, and how do you share it instead?

level: seniorimportance: should knowfreq 30%

basics

~10 s

deepcopy recurses into every attribute, and a threading.Lock cannot be reduced, so the pickle fallback raises TypeError: cannot pickle '_thread.lock' object. Define deepcopy to deep-copy the data fields and rebind the lock by reference.

open as a page

What does `copy.replace()` do, and what must a class define to support it?

level: middleimportance: nice to knowfreq 12%

basics

~10 s

copy.replace(obj, **changes), added in Python 3.13, returns a new object of the same type with the named fields replaced. It calls type(obj).replace and raises TypeError when the class does not define that method.

open as a page