A service stops getting faster — or gets slower — as you add worker threads, while the machine's processors are far from saturated. How do you confirm that lock contention is the cause and identify the specific hot lock?
answer
- threads up, throughput flat, processors idle → serialised
- normal profiler misses it: contention is off-CPU
- off-CPU / blocked-time / lock-event profiling
- sampled thread dumps = poor man's contention profiler
- spin locks invert the signature: high processor use
basics
~20 sProfile where threads wait, not just where they run: use off-CPU or blocked-time profiling, or sample thread dumps and count how often threads sit blocked on the same lock. Rising context-switch rates and flat throughput with growing thread counts confirm contention; the top blocked-on stack names the lock.
solid answer
~50 sAdding threads without gaining throughput means the bottleneck is serialised, and unsaturated processors rule out a compute limit. Two candidates remain: waiting on a lock, or waiting on I/O. A normal sampling profiler will not show this, because it samples running threads and the contended time is spent *not* running. I use off-CPU or blocked-time profiling — lock-instrumentation events giving blocked duration, the monitor being contended, and the owner — or, when nothing better is available, repeated thread dumps as a poor man's profiler: sample every second and count the fraction of samples where threads are blocked on each lock. The lock appearing in most blocked samples is the hot one. Corroborating signals: high voluntary context-switch rates, latency that grows with concurrency while service time stays flat, and throughput that plateaus at the level predicted by the serial fraction. If stacks show socket reads instead of lock waits, it is a dependency problem, not contention.
code
text · 10 linestake 60 dumps, one per second, count states per lock:
lock samples blocked share
inventoryCache monitor 43 72%
metricsRegistry monitor 6 10%
other 3 5%
not blocked 8 13%
=> inventoryCache is the serialisation point;
inspect its owner frames to see what is done while heldgo deeper
Recognise the symptom — more threads, no more throughput, idle processors — and know that threads are waiting rather than computing, most likely on a shared lock.
Explain why an on-CPU profiler misses contention and describe sampling thread dumps or blocked-time profiling to find the lock, plus obvious fixes like shrinking the critical section.
Bring in the scaling experiment and Amdahl reasoning, differential diagnosis against I/O waits, spin loops and false sharing, and an ordered remedy list with re-measurement after each change.
Frame it as capacity strategy: quantify the serial fraction, decide how much scalability is worth buying, weigh sharding and lock-free complexity against risk, and design the observability that makes contention visible before the next scaling wall.
## Why the symptom points at serialisation Amdahl's law says the speedup available from parallelism is bounded by the serial fraction: with a fraction *s* of work serialised, speedup can never exceed 1/*s* no matter how many workers you add. A throughput curve that flattens as threads are added is that bound made visible. Worse, throughput that *declines* indicates the added threads are paying coordination costs — context switches, cache-line ping-pong, scheduler churn — while gaining nothing. Unsaturated processors eliminate the compute-bound explanation. What remains is time spent waiting: on a lock, or on I/O. ## Why an ordinary profiler hides it Most sampling profilers answer "which code is on a processor?" Contended threads are precisely the ones *not* on a processor, so they are invisible or under-weighted. The tools that reveal contention answer a different question: - **Off-CPU / blocked-time profiling** — samples or traces threads while they are descheduled, attributing wait time to the stack that caused the wait. This is the direct instrument. - **Lock instrumentation** — runtimes and OS tracers can emit an event per contended acquisition carrying the lock identity, the blocked duration and often the owner. Aggregating by lock gives a ranked list of contention hot spots. - **Wall-clock profiling** — sampling all threads by elapsed time rather than processor time, so waiting shows up in the profile. - **Repeated thread dumps as a sampling profiler** — take a dump every second for a minute and count states. If 60% of samples show worker threads blocked on the same monitor, that monitor is the bottleneck. Crude, needs no special tooling, and works in environments where you cannot install anything. ## Corroborating system signals - **Voluntary context switches** rising with thread count: threads repeatedly block and are descheduled — the fingerprint of contention on a blocking lock. - **Latency decomposition**: if queue-wait time grows while service time stays flat, requests are waiting for capacity that a serialised section is withholding. - **Scaling experiment**: measure throughput at 1, 2, 4, 8, 16 workers. A plateau tells you the serial fraction; a decline tells you coordination overhead now dominates. - **Utilisation gap**: workers busy from the application's point of view but processors idle means the busy time is wait time. ## Differential diagnosis - **Contention on a blocking lock** — threads descheduled, low processor use, high voluntary context switches, blocked-on stacks concentrated on one monitor. - **Spin-lock or busy-wait contention** — the opposite processor signature: usage is *high* while throughput is flat, because waiting threads burn cycles spinning. A lock-free retry loop under heavy contention looks the same. - **I/O or dependency waiting** — stacks sit in socket or file read frames; downstream latency metrics move in step; adding threads may actually help until the dependency saturates. - **False sharing** — no lock at all: independent variables on the same cache line cause coherence traffic. Processor usage stays high, throughput scales badly, and the tell is a high cache-miss or coherence-event rate from hardware counters rather than any blocked time. - **Memory management or allocator pressure** — heavy garbage-collection or allocator lock contention can masquerade as application contention; check runtime memory metrics before blaming your own locks. ## Reading the hot lock correctly Once you have the contended lock, quantify it: how long is the critical section, how often is it entered, and what is the arrival rate? A short critical section entered a million times a second serialises just as thoroughly as a long one entered rarely; the product of hold time and entry rate is the serial fraction. Also check *why* it is held so long — a remote call, a disk write, an allocation or a callback inside a critical section is the usual culprit, and moving that work outside the lock is the highest-leverage fix. ## Remedies, in rough order of preference 1. **Do less inside the lock.** Compute outside, mutate inside; never perform I/O, remote calls or user callbacks while holding a lock. 2. **Remove the sharing.** Per-thread or per-partition state removes contention entirely; aggregate on read. 3. **Shard the lock.** Split one lock into N by key/hash so unrelated operations do not serialise. 4. **Weaken the exclusion.** Read-write separation, immutable snapshots with copy-on-write for read-mostly data, or optimistic versioning. 5. **Batch.** Amortise acquisition over many items rather than acquiring per item. 6. **Only then consider lock-free structures**, which trade contention for retry cost and much higher implementation risk. Always re-measure after each change: contention has a habit of moving to the next-narrowest point rather than disappearing.
- Your throughput is flat as threads increase but processor usage is high, not low. Does that rule out lock contention?No — it changes which kind. Spin locks, busy-wait loops and lock-free compare-and-swap retry loops consume processor while waiting, so contention on them looks compute-bound. Distinguish by profiling: if the hot on-CPU frames are the retry or spin path rather than useful work, that is contention burning cycles. False sharing shows a similar high-usage, poor-scaling signature and is identified with hardware cache-coherence counters instead.
- You have identified the hottest lock. What do you check before changing the locking strategy?Quantify the serial fraction: how long the lock is held and how often it is acquired, because their product determines the ceiling. Then look at what happens inside the critical section — a remote call, a disk write, an allocation or a user callback under the lock is usually the real defect, and moving it out is far cheaper and safer than redesigning the locking. Only if the section is already minimal do sharding, read-write separation or immutable snapshots become the right next step.
- After sharding the hot lock, throughput improves only slightly. What is the likely explanation?The bottleneck moved rather than disappeared: with the first serialisation point widened, the next-narrowest one now dominates — another lock, the allocator, a connection pool, or the downstream dependency. This is why each change must be followed by a fresh measurement. It can also mean the sharding key is skewed, so most traffic still lands on one shard, which shows up as uneven per-shard contention.
A supermarket adding more shoppers but keeping one checkout: the shop is not short of space, everyone is simply queueing at the same till.
saying these in an interview costs you the question
- Using a normal on-CPU sampling profiler and concluding there is no contention because nothing hot appears.
- Adding more threads or enlarging the pool to fix a serialised bottleneck.
- Assuming idle processors mean the system is healthy or under-loaded.
- Jumping to lock-free data structures before shortening the critical section or removing I/O from inside the lock.
- Ignoring that a very short critical section entered extremely often serialises just as much as a long rare one.