You're profiling a service and a shared metric counter shows up as a hotspot under load. Walk through how you'd decide between AtomicLong, LongAdder, and other options.
answer
- Confirm it's write contention before changing anything
- Write-heavy/read-rare → LongAdder; read-often or needs CAS → AtomicLong
- LongAdder has NO atomic read / CAS — only add/sum/reset
- Check the metrics library first (Micrometer/Dropwizard often LongAdder-backed)
- At millions of counters, padded-cell memory matters → shard / thread-local / sample
basics
~20 sFirst confirm it's really write contention, not something else. If many threads just increment a counter and you read it rarely, switch to LongAdder. If you need to read or compare-and-set a single exact value, keep AtomicLong. Or use a metrics library that already handles it.
solid answer
~50 sI'd start by confirming the diagnosis: is the counter actually hot because of write contention (many threads incrementing), or is it sized/read incorrectly? If it's genuinely a write-contended, read-rarely counter — a classic metric tally — LongAdder is the textbook fix: it stripes writes across padded cells so threads stop bouncing one cache line, and sum() being a non-atomic snapshot is fine for metrics. I'd keep AtomicLong when the value must be read frequently and exactly, or when I rely on compareAndSet/getAndIncrement semantics (sequence/ID generators, flags) — LongAdder can't do those. Before hand-rolling anything I'd check whether the metrics library already provides a striped or sampled counter (Micrometer, Dropwizard), since reinventing it adds risk. At extreme scale I'd weigh LongAdder's per-counter memory (padded cells near CPU count) against alternatives: sharded counters keyed by something natural, thread-local accumulation flushed periodically, or sampling. I'd validate the choice with a contended benchmark, not intuition.
go deeper
Knows the rule of thumb: many threads incrementing a metric → LongAdder; a single exact value you read → AtomicLong.
Can justify the switch via contention and the sum() trade-off, and recognizes LongAdder lacks compareAndSet so it's wrong for sequence generators.
Diagnoses whether contention is the real cause, chooses correctly across AtomicLong/LongAdder/LongAccumulator, and knows to lean on the metrics library; validates with a benchmark.
Reasons about system-scale trade-offs — per-counter memory at millions of counters, NUMA/coherence locality, sharding vs thread-local vs sampling, and library abstractions — and drives a measured, not assumed, decision.
## Step 1 — Verify it's really contention A counter showing in a profiler isn't automatically a *contention* problem. Distinguish: - **True write contention**: many threads doing `incrementAndGet()` on one shared `AtomicLong`, visible as time burned in the CAS loop / cache-coherence stalls (perf counters: high L2/L3 traffic, HITM events). This is what striping fixes. - **Volume without contention**: a single thread incrementing a lot — striping won't help; the work is inherent. - **Read pattern**: code that reads the counter on a hot path may benefit from AtomicLong's cheap exact read, which LongAdder *worsens*. Measure with a contended JMH benchmark or a flight-recorder profile under realistic thread counts before changing anything. ## Step 2 — Match the tool to the access pattern | Need | Choice | Why | |---|---|---| | Write-heavy, read-rare tally (requests, hits, errors) | **LongAdder** | Striped cells kill contention; non-atomic sum() is fine for metrics | | Running reduction other than sum (max latency, min free) | **LongAccumulator** | Same striping, custom associative operator | | Single value read often / atomically | **AtomicLong** | One location is cheap to read exactly; no cell walk | | Need compareAndSet / getAndIncrement semantics (sequence, ID, flag) | **AtomicLong** | LongAdder has no CAS-able value, only add/sum | | Low contention | **AtomicLong** | Less memory, simpler, no benefit from striping | Key limitation to remember: **LongAdder offers no atomic read or CAS** — it's add/sum/reset only. If your logic needs "increment and act on the exact new value," that's AtomicLong territory. ## Step 3 — Prefer existing abstractions Most services already have a metrics stack. **Micrometer** counters, **Dropwizard Metrics**, etc. already implement striped or sampled counting internally (often LongAdder-backed). Reaching for the library counter is usually better than hand-rolling: it's tested, integrates with reporting, and won't surprise the next engineer. Hand-rolling LongAdder is right when you own a tight internal hot path with no metrics layer. ## Step 4 — Consider scale-out trade-offs LongAdder isn't free: each instance can hold an array of **padded cells sized toward the number of CPUs** (each cell ~64+ bytes). With *one* hot counter that's nothing; with *millions* of counters (e.g. per-key statistics) the padding overhead can dominate memory. Alternatives at that scale: - **Sharded counters** keyed by a natural dimension (per-partition, per-core) summed on read — you control the shard count. - **Thread-local accumulation**: each thread keeps a private counter, flushed/aggregated periodically — zero contention, but reads lag and you must handle thread death. - **Sampling / approximate counting** when exactness isn't required. - On NUMA hardware, cross-socket coherence is even costlier, strengthening the case for per-core/per-socket locality. ## Step 5 — Validate Whatever you pick, prove it with a benchmark that reproduces the production thread count and access mix, and re-profile. "LongAdder is faster" is only true under contention; at low contention it can be slightly slower and uses more memory. Decisions here should be measured, not assumed. ## Summary judgement For the common case — a hot, write-heavy metric counter read occasionally — the answer is LongAdder (or the metrics library's counter). Keep AtomicLong for single-value, read-or-CAS semantics. Escalate to sharding/thread-local/sampling only when per-counter memory or NUMA effects make striping itself the bottleneck.
- Your counter feeds an ID/sequence generator, not a metric. Does LongAdder still apply?No. A sequence generator needs the exact new value atomically (getAndIncrement) and often compareAndSet semantics. LongAdder exposes neither — only add/sum. Use AtomicLong (or a dedicated striped sequence) so each caller gets a unique, exact value.
- Why might you NOT replace every AtomicLong with LongAdder by default?LongAdder uses more memory (padded cells), gives no atomic read or CAS, and at low contention is no faster (sometimes slightly slower). Blanket replacement bloats memory and removes capabilities you may rely on. Replace only contended, write-heavy, read-rare counters, ideally measured.
saying these in an interview costs you the question
- Replacing AtomicLong with LongAdder without confirming write contention exists
- Using LongAdder where compareAndSet / exact getAndIncrement is required
- Ignoring per-counter memory cost when there are huge numbers of counters
- Assuming LongAdder is universally faster regardless of contention level
- Reinventing a striped counter when the metrics library already provides one