In an Observer-based system, what problems arise from notification ordering, reentrant updates, and update storms, and how do you control them?
answer
- registration order is accidental, not a contract
- snapshot / copy-on-write before dispatch
- no writes back into the subject from update()
- begin/endUpdate batching; skip no-op notifications
- diamond ⇒ glitch ⇒ topological dirty-flag propagation
basics
~20 sObservers are usually called in registration order, which is accidental, so nothing should depend on it. An observer that changes the subject during update() can trigger recursive notification or infinite loops, and one change can fan out into thousands of updates. Fixes: snapshot iteration, reentrancy guards, batching, and change coalescing.
solid answer
~50 sThree related hazards. **Ordering**: the delivery sequence is an artifact of when listeners registered; observers that assume they run first (or last) break when wiring changes. If order genuinely matters, model it explicitly with priorities or a dependency graph, or split into phases. **Reentrancy**: an observer that writes back to the subject re-enters notify() mid-loop, giving recursion, repeated or lost notifications, and possibly a concurrent-modification error if a listener subscribes/unsubscribes during dispatch. Guard with snapshot (copy-on-write) iteration, a re-entrancy flag that queues nested changes for after the loop, and a rule that update() must not mutate its subject. **Storms**: a chain A→B→C means one edit fans out multiplicatively, and in diamond graphs an observer can be notified twice, once from an inconsistent intermediate state — the glitch problem. Mitigate with equality checks before notifying, transactional begin/endUpdate batching that emits one coalesced event, dirty-flag + single-pass recomputation in topological order, and debouncing/rate limiting for high-frequency sources.
code
pseudocode · 30 linesclass Subject {
observers = CopyOnWriteList()
private notifying = false
private pending = Queue()
private batchDepth = 0
set(value) {
if (value == current) return // 1. no-op suppression
current = value
if (batchDepth > 0) { dirty = true; return } // 2. batching
publish(Changed(value))
}
beginUpdate() { batchDepth++ }
endUpdate() { if (--batchDepth == 0 && dirty) { dirty = false; publish(Changed(current)) } }
private publish(e) {
if (notifying) { pending.add(e); return } // 3. reentrancy guard
notifying = true
try {
var event = e
while (event != null) {
for (o of observers.snapshot()) // 4. safe iteration, no lock held
try { o.update(event) }
catch (ex) { log(o, ex) } // 5. per-observer isolation
event = pending.pollOrNull()
}
} finally { notifying = false }
}
}go deeper
Know that observers are typically called in registration order, that you should not depend on it, and that changing the subject inside update() can cause loops. Mention copying the list before iterating.
Add concrete mechanisms: snapshot/copy-on-write iteration, a reentrancy flag, suppressing notifications when the value did not change, and begin/endUpdate batching for bulk edits.
Cover the glitch/diamond problem and two-phase dirty-flag propagation in topological order, exception isolation per observer, never dispatching under a lock, and debouncing high-frequency sources with an explicit backpressure policy.
Treat notification semantics as a published contract: define delivery order guarantees (or explicitly none), reentrancy behavior, at-most-once vs coalesced delivery, error containment, and batching boundaries — then enforce them in one shared dispatch mechanism instead of letting every subject reinvent them, and instrument notification rates to catch accidental N² wiring before it reaches production.
## 1. Ordering Most implementations iterate the listener collection in insertion order. That order depends on module initialization, dependency-injection graph shape, lazy loading, and which screen the user opened first — none of which are part of any contract. **Symptoms**: a validator that must run before a persister works in dev and fails in prod; a cache invalidator that runs *after* a view refresh so the view renders stale data; a test that passes alone and fails in a suite. **Remedies** - *Design ordering out.* Make observers independent and commutative: each reacts only to the event payload, never to another observer's side effects. - *If order is required, make it explicit.* Listener priorities/ordinals, an explicit phase model (validate → apply → notify-views → audit), or an ordered pipeline (which is Chain of Responsibility, not Observer). - *Split into distinct events.* Instead of relying on ordering within one event, publish `Validated`, then `Applied` — sequencing becomes a property of the subject, not of registration. - *Never document "registration order" as a guarantee* unless you are prepared to keep it forever, including under concurrent registration. ## 2. Reentrancy **Structural modification during dispatch.** An observer may unsubscribe itself ("I only wanted the first event") or subscribe a new one. Mutating the collection while iterating it corrupts the iteration or throws. Standard fixes: - Iterate an immutable **snapshot** taken before the loop (cheap array copy), or use a **copy-on-write** collection whose writers replace the array. Cost: an observer removed during dispatch may still receive the in-flight event — usually acceptable, but it must be documented, since a *disposed* observer receiving one last event is a classic crash. - Queue structural changes and apply them after the loop. **Recursive notification.** An observer's `update()` sets state on the same subject, which calls `notify()` again from inside the first `notify()`. Consequences: unbounded recursion and stack overflow if there is a cycle (A observes B observes A); observers receiving events out of causal order (the nested event is fully delivered before the outer loop finishes); and difficult-to-read stack traces. Remedies: - **Rule first**: `update()` should be side-effect-local — do not write back into the subject you are observing. - **Reentrancy guard**: a `notifying` flag; nested changes are appended to a pending queue and drained after the current loop returns (this also serializes causality). - **Cycle detection / depth cap** in development builds — fail loudly rather than overflow the stack. - **Convergence requirement**: if write-backs are inherent (constraint solving, layout), the update function must be idempotent and monotone so the loop terminates. **Locks and alien calls.** Calling an observer while holding the subject's lock is calling an *alien method* with a lock held: the observer may take another lock (or call back into the subject) and deadlock. Take the snapshot under the lock, release it, then dispatch. ## 3. Update storms and the glitch problem **Fan-out.** With a chain (model → viewmodel → view) or a fan (one model, 40 widgets), a single write becomes tens or thousands of calls. A bulk operation (import 100k rows, one notification each) can hang a UI thread for minutes purely in notification overhead. If each observer does a re-layout or a query, cost is multiplicative. **The glitch / diamond problem.** A changes; B and C both derive from A; D derives from B and C. Naive depth-first propagation notifies D once after B updates (while C still holds the old value — a state that is *inconsistent and never truly existed*), then again after C updates. D may render garbage, trigger an alert on a bogus intermediate value, or perform an expensive computation twice. This is the central problem reactive frameworks solve with topological ordering. **Mitigations** 1. **Don't notify on no-ops.** Compare old and new value; equal means silence. This alone removes a surprising share of storms. 2. **Batching / transactions.** `beginUpdate()` … `endUpdate()` suppresses notifications and emits one coalesced event (or one "bulk changed" event) at the end. Bulk APIs (`addAll`, `replaceRange`) should notify once, not N times. 3. **Dirty-flag two-phase propagation.** Phase 1 marks dependents dirty (cheap, no user code); phase 2 recomputes each dirty node exactly once, in dependency (topological) order. This is how spreadsheets and modern reactive/signal systems avoid glitches and duplicate work. 4. **Coalescing / debouncing / throttling.** For high-frequency sources (mouse move, market ticks, file watchers), collapse to the latest value per frame or per interval. Sample rather than deliver everything. 5. **Deltas instead of "everything changed".** A `RowsInserted(index,count)` lets a grid update 3 rows; an `Invalidated` forces a full rebuild. 6. **Async hand-off with a bounded queue.** Move slow observers off the publisher's thread with an explicit drop/coalesce policy under backpressure — never an unbounded queue. 7. **Instrument.** Count notifications per second and per source; a notification counter growing quadratically with data size is the signature of an accidental N² wiring. ## 4. Exception isolation (the fourth hazard people forget) In a naive loop, an exception from observer #2 aborts the loop: #3..#N never see the event, and the exception surfaces in the caller that merely assigned a field. Wrap each dispatch, log with the failing observer's identity, optionally auto-remove disposed observers, and consider an aggregated error rather than the first one. Decide deliberately whether an observer failure should invalidate the state change — usually it should not, which is another argument for publishing after the state change has committed. ## Checklist to recite Snapshot the list; don't hold locks across dispatch; don't mutate the subject inside update(); guard reentrancy with a flag plus a pending queue; suppress no-op notifications; batch bulk changes; propagate in topological order with dirty flags; debounce high-frequency sources; isolate exceptions per observer; and never rely on registration order.
- What is the 'glitch' problem in observer/reactive graphs and how is it eliminated?In a diamond dependency (D depends on B and C, both derived from A), naive depth-first propagation notifies D after B updates but before C does, so D observes an inconsistent intermediate state that never logically existed — and is then notified a second time. The fix is two-phase propagation: mark dependents dirty, then recompute each node exactly once in topological order, so every node sees a fully consistent snapshot and fires at most once per change.
- Why is it dangerous to notify observers while holding the subject's lock?You are invoking an alien method — arbitrary user code — under your lock. It may acquire other locks in a different order (deadlock), call back into the subject and try to re-acquire (deadlock or unintended reentrancy), or simply run slowly and serialize the whole subject. Copy the listener list under the lock, release it, then dispatch outside.
- An observer needs to unsubscribe itself while handling an event. How should the subject support that safely?Iterate an immutable snapshot (or a copy-on-write list) so the mutation cannot corrupt the loop, and document whether the departing observer may still receive the in-flight event. Alternatively queue structural changes and apply them after the dispatch loop; for one-shot listeners provide an explicit 'once' registration so the subject removes it before dispatching.
A phone tree. If everyone calls everyone the moment they hear anything, one piece of news becomes thousands of calls, some people get called twice with half the story before it's confirmed, and if two people call each other the line never clears. Real phone trees fix this exactly as good observer systems do: a fixed order, one call per person, and waiting until the message is final before dialing.
saying these in an interview costs you the question
- Relying on registration order as if it were a guaranteed contract.
- Mutating the subject inside update() and being surprised by recursion, duplicated events, or stack overflow.
- Removing a listener from the live collection during iteration instead of snapshotting.
- Letting one observer's exception abort the notification loop for everyone else.
- Emitting one notification per element from a bulk operation instead of a single coalesced event.
- Holding the subject's lock across observer callbacks.
- Solving storms with an unbounded async queue, which converts a CPU problem into an out-of-memory problem.
- Assuming a diamond dependency graph delivers each observer exactly one consistent update without explicit topological propagation.