skip to content

Explain the contention problem with a shared AtomicLong and how striping in LongAdder addresses it at the hardware level.

level: seniorimportance: should knowfreq 40%

answer

  1. Cache line = 64B unit of coherence; writers need exclusive ownership
  2. Shared AtomicLong = true sharing → cache-line bouncing + CAS retries
  3. LongAdder stripes value into per-thread cells
  4. @Contended pads each Cell to its own line (avoids false sharing)
  5. Cost: memory + non-atomic sum(); gain: write scalability

basics

~20 s

Many threads updating one AtomicLong keep failing their compare-and-swap and retrying because they all touch the same memory. That memory bounces between CPU caches. LongAdder gives each thread a different cell, so they touch different memory and stop fighting.

solid answer

~50 s

AtomicLong updates a single memory location with a CAS retry loop. Under high contention, that one location lives in one cache line, and every core that writes it must take exclusive ownership of that line, invalidating the copy in other cores' caches. The line ping-pongs between cores (cache-line bouncing), and CAS attempts keep failing because the value changed underneath them, forcing retries. Throughput stalls because the threads effectively serialize on one line. LongAdder breaks this by striping: it keeps an array of Cell objects, routes each thread to a cell by a per-thread hash, and pads each Cell (@Contended) so no two cells share a cache line. Now different threads write different lines — their CASes succeed first try and no invalidation traffic crosses them. It grows the array toward the number of CPUs when it still sees contention. The cost is more memory and a non-atomic sum() that walks all cells.

go deeper

for a junior

Understands that many threads on one counter slow each other down and that LongAdder splits the counter so they don't collide; precise cache mechanics not expected.

for a middle

Describes the CAS retry loop and that LongAdder uses multiple cells to reduce collisions; may know sum() walks all cells.

for a senior

Explains cache-line bouncing / coherence, true vs false sharing, the base+cells+probe design, and @Contended padding; states the memory and read trade-offs precisely.

for a principal

Connects this to broader contention engineering (NUMA effects, per-core/sharded counters, when striping's memory cost is prohibitive at huge counter counts, alternatives like thread-local aggregation or sampling), and to where the JDK reuses the pattern (ConcurrentHashMap counters).

## The unit of cache coherence: the cache line CPUs don't cache individual bytes; they cache **cache lines**, typically 64 bytes. When a core writes to memory, the **cache-coherence protocol** (e.g. MESI) requires that core to hold the line in an **exclusive/modified** state — meaning every other core's cached copy of that line is **invalidated** first. Reading it back elsewhere then re-fetches it. This is invisible to your code but governs concurrent performance. ## Why a shared AtomicLong stalls `AtomicLong.incrementAndGet()` is a CAS loop: read value `v`, compute `v+1`, atomically compare-and-swap. CAS itself requires exclusive ownership of the line holding the value. So with N threads on N cores all incrementing the **same** AtomicLong: 1. They all want exclusive ownership of the **one** cache line that value sits on. 2. Only one core can own it at a time; ownership **bounces** core-to-core ("cache-line bouncing" / "ping-pong"). Each transfer is tens to hundreds of cycles. 3. While a thread was computing `v+1`, another already changed `v`, so its **CAS fails** and it must reread and retry. The net effect: the lock-free counter degrades into something close to serialized — adding threads barely raises throughput and can lower it. Note this is **true sharing** (everyone genuinely contends one value), distinct from *false* sharing where unrelated variables happen to share a line. ## How striping fixes it `LongAdder`/`LongAccumulator` (and `ConcurrentHashMap`'s internal counters use the same idea) **stripe** the single hot value across many: - A `base` field handles the uncontended case (CAS it directly). - A lazily-allocated array of **`Cell`s**, each a tiny holder of one `long`. - Each thread carries a **probe** (a thread-local hash); `add()` uses it to pick a cell. Threads that collide on a cell trigger the probe to be rehashed and the array to grow (toward ~the number of CPUs), spreading them out. - Each `Cell` is annotated **`@jdk.internal.vm.annotation.Contended`**, which makes the JVM **pad** it so each cell occupies its own cache line. That's the crucial detail: without padding, adjacent cells would share a line and you'd recreate the bouncing as *false sharing*. Now thread A CASes cell 0's line and thread B CASes cell 3's line **simultaneously**, on different lines, with no coherence traffic between them. CASes succeed first try; throughput scales with cores instead of collapsing. ## The trade-off restated You converted one hot location into many cool ones. Costs: (1) **memory** — an array of padded cells (64+ bytes each) instead of 8 bytes; (2) **reads** — `sum()` must walk base + every cell and is therefore **not an atomic snapshot**; (3) tiny extra indirection. You gain near-linear write scalability under contention. That's why the rule is *write-heavy/read-rare → LongAdder; need a single frequently-read or CAS-able value → AtomicLong*. ## Mental model Think of one cashier (AtomicLong) with a long line versus many cashiers (LongAdder cells) each serving a few customers; to count total sales you walk to each register and add them up (sum()), and that total is a moment-by-moment approximation while sales keep happening.

  • Why are LongAdder's cells padded with @Contended?
    So each cell sits on its own cache line. If two cells shared a line, writing one would invalidate the other in other cores' caches — false sharing — re-introducing the exact bouncing that striping is meant to remove. Padding keeps each thread's writes coherence-independent.
  • Is the AtomicLong problem true sharing or false sharing?
    True sharing — all threads genuinely contend the same logical value. False sharing is when logically-independent variables accidentally share a cache line. LongAdder eliminates the true sharing by splitting the value, then uses padding to avoid re-introducing false sharing among the cells.

saying these in an interview costs you the question

  • Saying AtomicLong uses locks (it uses lock-free CAS; the problem is contention, not locking)
  • Confusing true sharing (the AtomicLong case) with false sharing
  • Omitting cache-line / coherence reasoning and only saying 'it's faster'
  • Not mentioning @Contended padding when explaining why striping actually helps
  • Claiming striping has no downsides (it costs memory and a non-atomic read)

context