How does narrowing a critical section reduce contention, and what work should you move out of a lock?
answer
- Lock only what touches shared state
- Move I/O, logging, computation, allocation out
- Compute outside, publish inside
- Keep check-then-act atomic — don't over-shrink
- Never call foreign code under a lock
basics
~20 sHold the lock for as short a time as possible. Only the steps that actually touch shared data need the lock. Move slow or unrelated work — I/O, network calls, logging, big computations, object creation — outside the locked block so other threads wait less.
solid answer
~50 sA critical section is the code that runs while a lock is held. Because the lock serializes that code, the longer it's held, the longer other threads wait — so contention is roughly proportional to lock hold time times acquisition rate. Narrowing the critical section means protecting only the operations that truly need mutual exclusion (the reads and writes of shared state) and pushing everything else out: I/O and network/database calls, logging, expensive pure computation, and object allocation. A common pattern is to compute a value outside the lock, then take the lock only to publish it. You must keep enough inside the lock to preserve correctness — the whole check-then-act must stay atomic where invariants depend on it — so you can't blindly shrink past the real invariant. The payoff is shorter hold times, lower contention, and better throughput, often with no change to the locking scheme itself.
go deeper
Knows to keep synchronized blocks small and not do slow work inside them.
Can identify specifically what to move out (I/O, logging, computation, allocation) and apply the compute-then-publish pattern.
Reasons about the correctness boundary — preserving multi-variable invariants and atomic check-then-act — while shrinking, and ties hold time to throughput/Amdahl.
Treats hold-time minimization as the first-line, lowest-risk optimization before lock-splitting/lock-free, and reasons about foreign-code-under-lock hazards and tail-latency effects at system scale.
## The lever: lock hold time When a lock is **contended**, the total time threads spend waiting is driven by two things: **how often** the lock is acquired and **how long** it is held each time. You usually can't reduce how often you need shared state, but you can almost always reduce **how long you hold the lock** — and that directly cuts the waiting other threads experience. Think of a lock as a single-lane toll booth. Every car must pass through one at a time. If each car stops to fill out paperwork *at the booth*, the queue explodes. If drivers fill out the paperwork *before* they reach the booth and only pause to drop it off, throughput soars. **Narrowing the critical section is moving the paperwork out of the booth.** ## What MUST stay inside the lock Only the operations that need **mutual exclusion** to stay correct: - reads and writes of the **shared mutable state** the lock protects, - any **check-then-act** or **read-modify-write** that must be **atomic** (e.g. 'if absent, put'; 'increment'; 'if balance >= amount, subtract'). Splitting these across the lock boundary reintroduces a race. ## What should be moved OUT Anything that doesn't touch the shared state, especially if it's slow or unbounded in time: - **I/O**: file, network, and database calls (these can take milliseconds to seconds — catastrophic to do under a lock), - **logging** (often does I/O and string building), - **expensive pure computation** that only depends on local inputs, - **object allocation / serialization / formatting**, - **calls into unknown code** (listeners, callbacks) — holding a lock while calling foreign code risks deadlock and unbounded hold times. ## The canonical pattern: compute-then-publish ``` // BAD: heavy work under the lock synchronized (this) { Result r = expensiveCompute(input); // slow, no shared state cache.put(key, r); // the only line that needs the lock } // GOOD: compute outside, publish inside Result r = expensiveCompute(input); // contention-free synchronized (this) { cache.put(key, r); // tiny critical section } ``` ## The correctness boundary — don't over-shrink Narrowing is safe only as long as every **invariant** that spans multiple shared variables stays inside one critical section. If updating `x` and `y` must appear atomic to other threads, you cannot release the lock between them. Likewise, if you compute outside the lock from shared state you read inside, that value may be **stale** by the time you publish — for cache-style code this is usually fine (you recompute or last-writer-wins), but for invariant-critical updates it is a bug. The skill is shrinking to *exactly* the invariant, no smaller. ## Why it works without changing the lock design Narrowing reduces the **serial fraction** of your program (Amdahl's law) without splitting locks or going lock-free. It's the **first, cheapest** optimization to reach for, and often enough on its own. Finer-grained locking and lock-free structures come *after* you've already minimized hold time.
- Give an example where shrinking the critical section would introduce a bug.A bank transfer that checks 'balance >= amount' then subtracts: if you release the lock between the check and the subtraction, two threads can both pass the check and overdraw. The check-then-act must stay atomic inside one critical section.
- Why is doing I/O under a lock especially bad?I/O has high and unpredictable latency (ms to seconds). Holding the lock for that whole time serializes every other thread behind a slow, externally controlled operation, turning a brief lock into a throughput killer and a source of huge tail latency.
saying these in an interview costs you the question
- Doing I/O or network/database calls inside a synchronized block
- Splitting an atomic read-modify-write across the lock boundary to 'shrink' it
- Holding a lock while invoking a callback/listener (deadlock + unbounded hold)
- Assuming shrinking is always safe regardless of multi-variable invariants