What does the rule 'never invoke code you do not control while holding a lock' mean, and why is shrinking a critical section a weaker argument for deadlock safety than it first appears?
answer
- open call = call out with no lock held
- alien code brings locks you cannot order
- never hold a lock across I/O, RPC, sleep, blocking put
- copy state under lock, release, then act
- shorter section = probability; removing a co-hold = correctness
basics
~20 sAn open call means releasing your lock before calling callbacks, plugins or remote services, because their internal locks add wait edges you cannot order. Shrinking a critical section only narrows the timing window; it lowers probability, not possibility.
solid answer
~60 sCalling a callback, plugin, virtual method, listener or remote service while holding a lock lets **code you cannot see** acquire locks you did not order, and lets an unbounded wait (I/O, a queue, a remote call) happen while you hold something everyone else needs. Both are edges missing from your lock graph. An **open call** is the discipline of invoking such code with no locks held: take the lock, copy the state you need into local variables, release, then call out and act. Where atomicity is required, publish an event to a queue and let a consumer do the outward work. Shrinking hold time is worth doing for throughput and for reducing the window in which an unlucky interleaving lands. But it is a **probabilistic** argument: the two acquisition orders still exist, so the cycle is still reachable, and a rare deadlock is worse than a frequent one because it appears in production and not in tests. Only removing an ordering violation, a co-held lock, or an unknown edge is a correctness argument.
code
text · 13 lines# risky: alien code runs under our lock
lock(state)
for listener in listeners:
listener.onChange(state.value) # may take unknown locks, may block
release(state)
# open call: nothing alien runs while we hold anything
lock(state)
snapshot = state.value
targets = copyOf(listeners)
release(state)
for listener in targets:
listener.onChange(snapshot) # stale-by-design, documentedgo deeper
Know the rule itself: do not call code you do not control, and do not do I/O, while holding a lock; copy what you need and release first.
Explain why alien code adds invisible edges to the lock graph, and restructure a listener notification into a snapshot-then-notify open call.
Separate probability arguments from correctness arguments, handle the staleness the snapshot introduces, and give the outbox pattern when atomicity is demanded.
Make it enforceable - held-lock assertions at outward boundaries, hold-duration metrics, and module boundaries where locks are never allowed to cross - and set the policy that rare deadlocks are treated as high severity.
## What an open call is A critical section is code executed while a lock is held. An **open call** is a call made with no lock held. The rule says: whenever you must invoke code whose locking behaviour you do not control - a user-supplied callback, a listener, an overridable method, a plugin, a library entry point, a remote service - do it as an open call. The reason is that your deadlock reasoning depends on knowing every lock a thread can hold at once. Calling out under a lock imports an unknown set of acquisitions into your critical section. If that code takes lock X, your thread now co-holds your lock and X, in an order nobody designed and no reviewer of either module can see. If another path takes X first and then yours, the cycle already exists and is simply waiting for the interleaving. ## The second hazard: unbounded waits under a lock Even if the callee takes no locks, it may block: a network round trip, a disk read, a blocking queue put, a sleep, a wait for another thread's result. Holding an exclusive resource across an unbounded wait converts a remote problem into a local outage. Every thread that needs your lock is now queued behind a dependency you do not control, and if the remote side ends up waiting on this service, you have built a genuine cross-process circular wait. The general form of the rule is therefore stronger than callbacks: **never hold a lock across I/O, a remote call, a sleep, or any wait whose duration you cannot bound.** ## The safe pattern The standard restructuring is copy-then-call: 1. Acquire the lock. 2. Read or mutate only your own in-memory state; build a local snapshot of whatever the outward call needs. 3. Release the lock. 4. Perform the outward call using the snapshot. This usually changes semantics slightly - the snapshot may be stale by the time the call runs - and that must be handled deliberately: recheck a version or sequence number afterwards, make the outward action idempotent, or accept the staleness where it is harmless. When true atomicity between the state change and the outward effect is required, the answer is not to hold the lock longer but to make the outward effect asynchronous and durable: write an intent record or enqueue an event under the lock, and let a separate consumer perform the call with retries. Iterating a shared collection while calling listeners is the same pattern in miniature - copy the listener list under the lock, release, then notify. ## Why shrinking hold time is a weaker argument Reducing time under a lock is genuinely valuable. It raises throughput, cuts queueing delay, and shrinks the window in which the dangerous interleaving can occur. Some people stop there and call the deadlock fixed. It is not fixed. If two code paths acquire the same pair of locks in opposite orders, the cycle is reachable no matter how brief each hold is. All you changed is the probability of hitting it. And probability is a treacherous currency here: a deadlock that occurs once per million operations passes every test suite, survives code review, and then hangs production at peak traffic - exactly when the outage is most expensive and the evidence is hardest to collect. "Rare" is worse than "frequent" for a bug of this class. The distinction to articulate is between a **correctness** argument and a **probability** argument. A global acquisition order, a single lock, immutability, confinement or a removed co-hold are correctness arguments: they make the cycle non-existent. Shorter critical sections, retry with backoff timing tweaks, and sleeps are probability arguments. In an interview, saying "we shortened the critical section so it went away" signals that the candidate cannot tell the two apart. Saying "shortening helped throughput, but the fix was removing the second lock from that path" signals that they can. ## The narrow case where hold time is a correctness lever There is one exception worth naming: if reducing the critical section removes a **co-hold** entirely - the section no longer holds lock A at the moment it asks for lock B - then hold-and-wait is gone for that path and the cycle really is eliminated. That is not "shorter"; that is structurally different, and it is worth saying explicitly, because it is the version of the advice that actually prevents deadlock rather than hiding it.
- The outward call must appear atomic with the state change - the state must not be visible as updated unless the call happened. How do you get that without holding the lock across the call?Split it into two durable steps. Under the lock, apply the state change and record an intent - an outbox row or a queued event - as part of the same atomic unit. Outside the lock, a consumer reads the intent and performs the call, retrying until it succeeds, with the call made idempotent so retries are safe. You trade strict atomicity for eventual consistency, which is the standard price for not holding an exclusive resource across an unbounded wait.
- How would you detect violations of this rule automatically?Instrument lock acquisition to record what each thread holds, and assert in debug or test builds that the held set is empty when entering known outward boundaries - the HTTP client, the callback dispatcher, blocking queue operations, sleeps. Static analysis can catch the common shapes where a call to an interface method appears lexically inside a critical section. In production, timing lock hold duration and alerting on outliers surfaces the same violations empirically, since a hold that lasts as long as a network round trip is almost always a lock held across I/O.
saying these in an interview costs you the question
- Claiming a shorter critical section fixes a deadlock, when the conflicting acquisition orders still exist.
- Holding a lock across a network call or blocking queue operation because 'it is usually fast'.
- Invoking user-supplied callbacks under a lock and assuming they take no locks of their own.
- Treating a rare deadlock as low severity, when rarity mainly means it will surface in production rather than in tests.
- Refusing to release the lock because the outward call must be atomic, without considering an outbox or event-queue design.