Serializing an object and immediately deserializing it is a popular way to get a deep copy. What are the trade-offs of that technique, and what would you prefer at scale?
answer
- Round-trip = generic deep copy, one line
- Slow and allocation-heavy
- Copies identity, wiring, caches; unbounded reachability
- JSON loses reference identity; cycles break
- Bypasses constructors → invariants lost, CVE history
basics
~20 sSerialize-then-deserialize gives a deep copy in one line and handles nested structures generically. But it is slow, copies everything including things you should not copy, silently drops fields the format cannot represent, and can lose shared references or fail on cycles.
solid answer
~1 minRound-tripping through a serializer (binary, JSON, or a language's built-in serialization) is attractive because it is generic and short: any reachable graph is duplicated without writing per-class copy code. The costs are substantial. **Performance**: it allocates an intermediate representation and pays encode+decode, typically an order of magnitude slower than a hand-written copy. **Over-copying**: it follows every reference, so one back-pointer to a context object can drag half the heap along, and identity, caches, listeners and resource handles are copied unless explicitly excluded. **Fidelity loss**: text formats like JSON have no notion of reference identity, so a diamond becomes two copies and a cycle either throws or needs special encoding; types, precision, and unrepresentable members (functions, handles) are silently degraded. **Type safety and security**: deserializing untrusted or reflectively-reconstructed data has a long history of remote-code-execution vulnerabilities, and deserialization normally bypasses constructors, so class invariants are never re-checked. **Coupling**: the copy now depends on the serialization schema, so a format or version change quietly changes copy behaviour. At scale I prefer explicit copy constructors or copy methods on the types that need them, immutability with structural sharing where possible, and copy-on-write for hot paths — reserving serialization round-trips for tests, prototypes, and true process/network boundaries where you are serializing anyway.
code
pseudocode · 14 lines// generic, tempting, and quietly wrong for entities
copy = deserialize(serialize(original))
// - follows every reference (context back-ref -> half the heap)
// - keeps id, version, listeners, cache unless excluded
// - JSON: a.left === a.right becomes two objects; cycles throw
// - constructor/validation not run on the rebuilt object
// explicit and reviewable
Order(Order other) { // copy constructor
this.lines = other.lines.map(Line::new); // owned -> deep
this.total = other.total; // value -> share
this.id = null; this.version = 0; // identity -> reset
this.listeners = []; // wiring -> reset
}go deeper
Say it produces a deep copy easily but is slow and can lose or mangle parts of the object.
List the concrete costs — performance, over-copying, cycles/shared references in text formats, dropped fields — and note it copies IDs and caches too.
Add constructor bypass and invariant loss, deserialization security history, format-versus-copy coupling, and give the preferred alternatives with their conditions.
Frame copying as an ownership-boundary question: unbounded deep copy signals unclear aggregate boundaries; prefer immutability and explicit, tested copy contracts, and confine round-tripping to real serialization boundaries.
### The technique ``` copy = deserialize(serialize(original)) ``` **Serialization** turns an object graph into a linear representation (bytes or text); **deserialization** rebuilds objects from it. Because rebuilding necessarily allocates fresh objects, the result is a deep copy of whatever the format managed to represent. Variants: a language's built-in object serialization, a JSON/YAML mapper, a binary schema format, or a structured-clone facility. ### Why people reach for it - **Zero per-class code.** No `clone()` to write or maintain, no risk of forgetting a newly added field — a genuine advantage, since hand-written copy methods rot when fields are added. - **Handles arbitrary nesting** without writing a traversal. - **Already available** in most stacks. ### Cost 1 — performance You allocate an intermediate buffer, encode every field (often with reflection and string keys), then parse and allocate again. Compared with a direct field-by-field copy this is typically 10–100× slower and allocates far more garbage. Fine once at startup; ruinous inside a per-request or per-frame loop. ### Cost 2 — it copies everything, including what must not be copied Serialization is *reachability-driven*. It has no idea that a primary key is identity, that a listener list is wiring, that a cache is derived, or that a connection handle names an OS resource. Two consequences: - **Unbounded scope.** One back-reference from a child to a shared context/session/container makes the entire application graph reachable. Teams discover this as a mysterious multi-second "copy". - **Semantic corruption.** Copied identifiers, copied version stamps, copied observers, duplicated handles. Excluding them requires deliberate markers (transient/ignore annotations, custom hooks) — so you end up writing per-class knowledge after all, just in a less visible place. ### Cost 3 — fidelity loss - **Reference identity is usually not preserved.** In JSON there is no way to say "this is the same object as that one", so a diamond (`a.left === a.right`) round-trips into two distinct objects, and a cycle either throws or must be encoded with reference markers the format may not support. Binary/built-in serializers often *do* preserve identity via reference tables — a real difference worth knowing per technology. - **Type erosion.** Text formats collapse types: dates become strings, integers may become floats, precision may be lost, and polymorphic subtypes need explicit type discriminators or come back as the base type. - **Unrepresentable members** — closures, handles, thread references — are silently dropped or throw. - **Silent field drops.** A field the mapper cannot see (no accessor, wrong visibility, unknown-property policy) simply vanishes from the copy, and nothing reports it. This turns "never forgets a new field" from an advantage into a hazard, because the failure is silent rather than a compile error. ### Cost 4 — invariants and security Deserialization usually **bypasses constructors**, reconstructing objects field by field. Whatever validation your constructor enforces (non-null, ranges, cross-field consistency) is not applied to the copy. If a serialized form is ever produced elsewhere, deserializing untrusted input is one of the most exploited vulnerability classes there is — arbitrary type instantiation leading to remote code execution. Even in a purely in-process copy, the machinery that makes it possible (reflective, constructor-bypassing reconstruction) is the same machinery that makes those attacks work, and enabling it broadly enlarges the attack surface. ### Cost 5 — coupling and drift The copy's behaviour is now a side effect of the serialization schema. Add a schema version, change a naming policy, mark a field ignored for API reasons — and object copying changes silently, in a completely unrelated part of the system. Copy semantics should not be a downstream consequence of wire-format decisions. ### What to prefer, and when 1. **Design away the need.** Immutable value objects don't need copying; persistent/structurally-shared collections give O(1) "copies". This is the highest-leverage move: it removes an entire bug category rather than implementing it well. 2. **Explicit copy constructor / copy method** on the few types that truly need it. Verbose, but every field decision is visible, reviewable and testable, and it can reset identity, drop caches and clear wiring. Guard the "forgot a new field" weakness with a reflection-based *test* that fails when an unhandled field appears — test-time reflection, not runtime reflection. 3. **Copy-on-write** on hot paths where copies are frequent and mutations rare. 4. **A curated deep-copy utility** with an identity map and an explicit cut set, used for a known family of types — the middle ground when many similar structures need copying. 5. **Serialization round-trip** legitimately, when: you are crossing a process/network/storage boundary anyway; in tests as an independent oracle for a hand-written copy; for quick prototypes; or for plain data-transfer objects that have no identity, wiring, resources or invariants — precisely the case where all the objections above vanish. ### The architectural framing The deeper point is about **ownership boundaries**. Copying is only well-defined for a bounded aggregate with a clear owner. If a system needs unbounded deep copies, that is usually evidence that object ownership is unclear — everything can reach everything. Fixing the boundaries is worth more than optimizing the copy.
- Copy constructors are safe but people forget to update them when a field is added. How do you defend against that?Add a test that reflects over the type's fields and fails when a field is not accounted for by the copy logic (either duplicated or explicitly listed as reset/shared). That keeps reflection at test time, where a failure is loud and cheap, instead of at runtime where it would be generic and silent. Language features that generate copy methods from a declared field list (records/data classes) give the same guarantee at compile time.
- When is a serialization round-trip actually the right choice for copying?When the objects are plain data — no identity, no wiring, no external resources, no constructor invariants — and especially when you are already serializing at a process, network or storage boundary. It is also useful in tests as an independent check against a hand-written copy, and acceptable in throwaway prototypes where the performance and fidelity costs do not matter.
Duplicating a machine by describing it over the phone and having someone rebuild it from the description. Anything the words can express comes through; anything they can't — the serial number's meaning, the fact that two parts were the same part, the plug currently in the wall — is lost or wrongly reproduced. And the rebuild never runs the factory's safety inspection.
saying these in an interview costs you the question
- "Serialize/deserialize is the standard way to deep copy" — it is a shortcut with performance, fidelity, security and scope costs.
- Assuming a JSON round-trip preserves shared references or survives cycles.
- Not realising deserialization commonly bypasses constructors, so class invariants are never validated on the copy.
- Ignoring that reachability makes the copy's scope unbounded — one context back-reference clones the world.
- Treating the technique as safe for objects carrying identity, listeners, caches, secrets or resource handles.
- Letting wire-format/schema decisions silently redefine what copying means.