When cloning an object with the Prototype pattern, which parts of its state should NOT be copied verbatim, and what should happen to them instead?
answer
- Identity: reset ID, version, createdAt
- Wiring: clear listeners and parent back-refs
- Resources: sockets/locks/threads not copyable
- Derived: drop caches and dirty flags
- Services and immutables: share, don't duplicate
basics
~20 sAnything that identifies the object or ties it to the outside world: database IDs, creation timestamps, version counters, open connections and file handles, locks, caches and registered listeners. These should be reset, regenerated, or re-established rather than duplicated.
solid answer
~60 sA copy is a *different* object, so state that expresses "which object this is" or "what this object is connected to" must not be duplicated. Concretely: **identity** — primary keys, UUIDs, natural keys, `createdAt`/`createdBy`, optimistic-locking version or etag — must be reset or regenerated, otherwise two objects claim to be the same record and persisting the clone overwrites or violates uniqueness. **Wiring** — observer/listener lists, subscriptions, registrations in a parent container, back-references — should be reset to empty and re-established by whoever installs the clone, otherwise one event fires handlers twice. **External resources** — sockets, file handles, threads, mutexes, transactions — cannot be meaningfully copied at all; share them if that is safe, re-acquire them, or forbid cloning. **Derived state** — caches, memo tables, dirty flags, lazily-computed values — should be dropped so the clone recomputes, since the cached value may not be valid for the copy's new context. What remains, and what the clone should genuinely duplicate, is the owned mutable domain state. Documenting this per class is part of the copy contract.
code
pseudocode · 15 linesAccount clone() {
Account c = new Account();
c.balance = balance; // owned value -> copy
c.holder = holder.clone(); // owned mutable -> deep copy
c.currency = currency; // immutable -> share
c.audit = auditService; // service -> share
c.id = null; // identity -> reset
c.version = 0; // opt-lock -> reset
c.createdAt = null; // provenance -> reset
c.listeners = []; // wiring -> reset
c.balanceCache.clear(); // derived -> drop
// c.dbConnection: deliberately NOT carried over
return c;
}go deeper
Name the obvious ones — IDs and open connections — and say the copy is a different object so its identity must differ.
Give the full taxonomy (identity, wiring, resources, derived, services, owned state) with the treatment for each.
Explain concrete failure modes: overwritten rows, defeated optimistic locking, double-fired listeners, double-close, stale caches, and why automatic/reflective copiers get these wrong by default.
Treat copy semantics as part of the type's published contract, enforced by tests and review; push entities toward hand-written copy operations and reserve generic deep copy for pure value objects.
### Why the question exists "Deep copy" is often taught as "duplicate everything reachable", which is wrong in practice. Real objects carry several *kinds* of state, and only one of them wants duplication. | Kind of state | Example | Correct treatment on clone | |---|---|---| | Owned mutable domain state | line items, nested settings, buffers | **Duplicate** | | Immutable values | money amounts, frozen config, strings | **Share** (copying wastes memory) | | Identity | primary key, UUID, natural key, createdAt, version/etag | **Reset / regenerate** | | Wiring | listener lists, subscriptions, parent back-reference, registry membership | **Reset**, re-established by the installer | | External resources | socket, file handle, thread, lock, transaction, connection pool | **Share if safe, re-acquire, or refuse to clone** | | Derived/cached | memo tables, computed totals, dirty flags, lazily-loaded fields | **Drop**, recompute on demand | | Shared services | logger, clock, metrics sink, injected collaborators | **Share** (never duplicate) | ### Identity: the highest-severity mistake A persistence identifier is a claim about which row in a database an object *is*. Copy it and you now have two in-memory objects with the same claim. Symptoms range from a unique-constraint violation on save (the good case — it fails loudly) to an update that silently overwrites the original record with the clone's data (the bad case — data loss with no error). The same applies to optimistic-locking version numbers and etags: a copied version stamp makes the clone appear to be a legitimate update of the original, defeating concurrency control. Audit fields (`createdAt`, `createdBy`) copied verbatim lie about provenance and break compliance trails. Rule: clone identity to *unset* and let the normal creation path assign a new one, or generate a fresh identifier inside the copy operation. ### Wiring: the double-fire class of bugs If the object keeps a list of registered observers and you copy the list, then a single domain event now notifies every handler twice — once via the original, once via the clone — even though the user only sees one object. Worse, the handlers may hold references back to the original, so the clone drives behaviour on the wrong object. Copying a back-reference to a parent is equally wrong: the clone claims to belong to a container that has never heard of it, and the container's invariants (counts, indexes) are now inconsistent. Rule: clones start unwired; whoever inserts the clone into the world wires it. ### External resources: not copyable at all A socket is a kernel-level endpoint; a file handle is an OS descriptor; a lock has an owner and a wait queue; a thread is a running execution. There is no meaningful "copy" of any of these. Duplicating the handle value gives two objects that will both close the same descriptor — a classic double-close / use-after-close bug. Copying a mutex object may produce an unowned or, worse, apparently-held lock. The available choices are: share the same resource (fine when the resource is genuinely shared and thread-safe), re-acquire an equivalent one (open a new connection), or declare the type non-cloneable. ### Derived state: correctness *and* staleness Caches are copied "for free" but may be invalid for the clone: a memoized total computed for the original's context, a lazily-loaded association bound to a session that the clone is not part of, a `dirty` flag that says "I have unsaved changes" when the clone has never been saved at all. Dropping derived state costs a recomputation and removes a whole family of stale-value bugs. (The narrow exception: an expensive, context-independent precomputation is exactly the value a prototype exists to carry — that one you keep.) ### Security-sensitive fields A frequently-missed category: secrets and permissions. Cloning an object that carries a decrypted credential, a session token, or an access-control decision may leak it into a context that should not have it. If the clone will live in a different scope, drop or re-derive those fields. ### Making it a contract, not a habit - Write the policy per field in a short comment or in the type's documentation; a future reader cannot infer "reset" from code that simply omits a field. - Test it: assert the clone has a fresh/unset identity, no listeners, an empty cache, and that saving the clone creates a new record rather than updating the original. - Beware reflective/automatic deep copiers and serialization-based cloning: they copy *every* field by default, so identity, wiring and caches all come along unless explicitly excluded. Whatever exclusion mechanism your language offers (transient markers, ignore annotations, a hand-written copy method) must be applied deliberately. - Prefer a hand-written copy operation or copy constructor for entities; generic deep copy is safest for pure value objects, which have none of these categories.
- What concretely goes wrong if a cloned database entity keeps the original's primary key and optimistic-locking version?Saving the clone is interpreted by the persistence layer as an update to the original row, so the original's data is overwritten — or, if a unique constraint catches it, the save fails at a confusing place far from the clone. The copied version stamp also makes the write pass the optimistic-lock check it should have failed, defeating concurrency control.
- Serialization-based cloning copies every field automatically. How do you keep identity, caches and connections out of the copy?You must mark them excluded explicitly — transient/ignore markers, custom read/write hooks, or a post-copy fix-up step that resets identity, clears caches and drops resource handles. Relying on the default is how these bugs get shipped, which is a strong argument for a hand-written copy method on entities.
Photocopying your passport does not give you a second passport. The photo and name copy across fine, but the passport number, the visas stamped in it, and the chip that authenticates you must not — a duplicate number would make two documents claim to be the same identity.
saying these in an interview costs you the question
- "A correct deep copy duplicates everything reachable" — identity, wiring, resources and caches must not be duplicated.
- Cloning an entity and saving it, expecting a new row while the ID was copied.
- Copying observer/listener lists, then wondering why handlers run twice.
- Duplicating a file descriptor or socket handle, causing double-close or use-after-close.
- Carrying over cached/lazily-loaded values that were valid only in the original's context or session.
- Copying secrets, tokens or authorization decisions into a clone that lives in a different security scope.