skip to content

questions

5

When concurrent tasks exchange data as messages, what properties must the message type have for any number of receivers to read it safely without coordinating, and what does immutable have to mean for a value that contains other objects inside it?

level: juniorimportance: must knowfreq 55%

answer

  1. immutable = whole reachable graph, not the top object
  2. array/collection field = the classic leak
  3. copy at construction OR transfer ownership
  4. no retained alias in the sender's scratch buffer
  5. change = derive a new value, share the unchanged parts

basics

~20 s

The message must be fully built before it is sent and never change afterwards, all the way down: nested objects, arrays and collections must be immutable too, and the sender must keep no reference it can mutate. Then every reader sees identical content regardless of timing.

solid answer

~60 s

A message is safe for unsynchronized reading when its content cannot change after it is sent. Three requirements, and the second and third are where designs fail: 1. **Fully initialized before sending.** No field is filled in later, no lazy setup that mutates state on first read. 2. **Transitively immutable.** Immutable is a property of the whole reachable object graph, not the top object. A record with a read-only reference to a mutable list is mutable: whoever holds the list can change what the message says. 3. **No retained aliases.** If the sender built the message from a buffer it still owns and keeps writing to, the message changes under the receiver. Copy at construction, or hand over ownership of the source. Updates are expressed as new values — derive a modified copy and send that — which is why immutable message design pairs with copy-on-write. The costs are allocation and copying for large payloads, and the fact that identity comparisons stop being meaningful; the benefit is that the message needs no locking and can be shared, replayed, cached and logged freely.

code

text · 9 lines
text
# broken: message shares the sender's mutable list
items = mutable_list()
items.add(a); items.add(b)
send(ch, Order(id, items))     # receiver holds the SAME list
items.add(c)                    # the in-flight message just changed

# fixed: copy at construction, expose read-only
class Order(id, items):
    this.items = immutable_copy_of(items)

go deeper

for a junior

Say the message must be fully built before sending and must never change afterwards, and that this has to hold for nested collections too.

for a middle

Add transitive immutability and defensive copying at construction, plus that updates are expressed as new derived values.

for a senior

Raise the alias leak from reused sender buffers, the cost model for large payloads, and when to switch from copying to ownership transfer.

for a principal

Frame it as a boundary contract: immutable value types at every hand-off point so sharing is free and replay/log/cache are sound, with explicitly documented ownership transfer as the escape hatch for large payloads.

## What safe to read without coordinating requires Two concurrent readers can disagree about a value only if something writes to it while they read. Remove the writes after publication and the disagreement is impossible — no lock, no ordering rule, no copy per reader. That is the entire argument for immutable messages, and it is why message-oriented designs lean on immutability so heavily: it converts a timing problem into a type-design problem. The requirement is not immutable-ish. It is: from the moment the message is handed off, no reachable part of it changes, forever. ## Transitive immutability The most common design error is stopping at the outer object. Consider a message holding an order id and a list of line items. Marking the fields read-only prevents replacing the list, not modifying its contents. Anyone holding a reference to that list — including the sender that built it — can add an item after the message is in flight, and a receiver iterating it may see a different order than another receiver, or a partially updated one. So the rule is transitive: every object reachable from the message must itself be immutable, or must be exclusively owned by the message with no outside reference. In practice that means: - Copy incoming mutable collections and arrays at construction, and expose only read-only views. Arrays are the classic leak, because an array field is never deeply read-only. - Do not hand out internal references from accessors; return copies or immutable views. - Prefer value-like types for anything embedded: dates, money, ids. - Beware of a field typed as an interface. Declaring it as a read-only collection type does not guarantee the object behind it is immutable; only construction under your control does. ## No retained aliases Even a perfectly immutable class can be defeated by how the message is produced. If the sender reuses a scratch buffer for each message and wraps it without copying, every message points at the same mutable memory. This is a common real bug in high-throughput message pipelines, because it is invisible in the message type and appears only under load as corrupted or duplicated payloads. The two clean resolutions are: copy on construction (the sender may keep using its buffer), or transfer ownership (the sender drops its reference and never touches the buffer again). Choose one explicitly and document it, because callers cannot infer it. ## Designing for change If a message can never be modified, updates are new messages. A consumer that needs a variant derives it: take the original, produce a copy with one field replaced, and pass that on. Languages with copy-with-modification syntax make this cheap to write; structural sharing makes it cheap to run, because unchanged sub-objects are shared rather than copied — safe precisely because they are immutable. This also gives free properties that mutable messages cannot offer: a message can be broadcast to many consumers without any of them defensively copying; it can be retried or replayed because it still says what it said; it can be logged and the log will not lie; it can be cached and its hash is stable, which is why immutable types make sound map keys. ## Costs, honestly Allocation and copying: for large payloads, per-message copies can dominate, which is where ownership transfer of a mutable buffer becomes the better tool. Memory churn: many short-lived values put pressure on the allocator, though generational collectors handle this well. Ergonomics: deeply nested immutable structures are tedious to update without library support. And identity: two equal messages are not the same object, so any logic keyed on reference identity must move to value equality. ## The rule of thumb Make messages small, transitively immutable value types built once at the boundary. Where the payload is large or performance-critical, do not fake immutability with a shared buffer — use explicit ownership transfer instead, so exactly one task touches the data at a time.

  • A message class has only read-only fields but one of them is an array. Is it immutable?
    No. Read-only means the field cannot be re-pointed, not that the array's elements cannot be written. Anyone with the reference, including the code that built it, can change the contents after the message is sent. Copy the array at construction and expose it only through a read-only view or an accessor that returns a copy.
  • Your payload is a multi-megabyte buffer and copying it per message is too expensive. What now?
    Stop pretending it is immutable and use explicit ownership transfer instead: the producer finishes all writes, sends the buffer, and drops its reference so exactly one task can touch it at a time. Alternatively slice the buffer into disjoint non-overlapping regions handed to different tasks, which is safe because no two tasks address the same bytes.

A message should be a printed and sealed letter, not a shared whiteboard someone can still walk up to and edit while others are reading it.

saying these in an interview costs you the question

  • Believing that marking fields read-only makes the object immutable regardless of what they point to.
  • Exposing an internal collection or array directly from an accessor.
  • Reusing one mutable buffer for every message and calling the message immutable.
  • Thinking immutability is only about thread safety and ignoring that it also enables free sharing, replay and caching.
  • Adding a lazily computed mutable cache field to an otherwise immutable message without making it safe.

context

open as a page

What does it mean to confine mutable state to a single task or thread, what forms does confinement take in practice, and why does confined state need no synchronization?

level: middleimportance: must knowfreq 48%

basics

~20 s

Confinement means only one task can ever reach a piece of mutable state, so concurrent access is impossible and no locking is needed. It comes as local variables, per-task storage, a single owner task, or one owner per data partition — and it holds only while no reference escapes.

open as a page

Explain the copy-on-write technique for sharing a data structure between many concurrent readers and an occasional writer: what invariant lets readers run without locking, and what is its cost model?

level: middleimportance: should knowfreq 40%

basics

~20 s

Readers read a snapshot that is never modified in place. A writer copies the structure, changes the copy, and atomically swaps the shared reference. Readers need no lock and see a consistent, possibly slightly stale, version. Cost: a full copy per write, so only for read-mostly data.

open as a page

Copying is too expensive, so one task must hand a large mutable buffer to another instead. What rules make that handoff safe, and how do you make I no longer own this something stronger than a comment?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Keep the rule that exactly one task may touch the buffer at any time: the sender finishes all writes, hands it over, and drops its reference for good. Use a handoff primitive that also provides ordering, and enforce ownership with move semantics, a detached source, or a one-shot handle — not a comment.

open as a page

You are designing a service where every request mutates some account's state. Compare a share-nothing design, where state is partitioned so exactly one task owns each account, with a shared-state-and-locks design. What do you gain, and where does share-nothing break down?

level: principalimportance: should knowfreq 33%

basics

~20 s

Share-nothing routes each account to its single owner, so updates are serialized per account with no locks, cache-friendly and easy to reason about. It breaks on hot partitions, on operations spanning two accounts, and on ownership changes, which need fencing to prevent two owners.

open as a page