skip to content

Race Conditions

A race is unsynchronized access to shared mutable state, usually a check-then-act or read-modify-write sequence that is not atomic. Interviewers hand you a lazy-init or counter snippet and ask what can go wrong and why it passes tests anyway.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

4

What is a race condition in Java, and why does it lead to nondeterministic, data-dependent bugs?

level: juniorimportance: must knowfreq 85%

answer

  1. Timing-dependent correctness = race
  2. count++ is read-modify-write (3 steps)
  3. check-then-act / lazy init
  4. Needs: shared + mutable + >=1 writer + no sync
  5. Nondeterministic + data-dependent = hard to reproduce

basics

~20 s

A race condition is when two or more threads touch the same shared data at the same time without coordination, and the result depends on which thread happens to run first. Because timing varies, you get different, sometimes wrong, results each run.

solid answer

~40 s

A race condition occurs when the correctness of a program depends on the relative timing or interleaving of multiple threads accessing shared mutable state, and at least one access is a write. The threads aren't synchronized, so the OS/JVM scheduler is free to interleave their operations in any order. Most runs are fine; occasionally the operations interleave badly and you read stale data or lose an update. The bug is nondeterministic (depends on scheduling) and data-dependent (only shows up under certain values and load), which makes it hard to reproduce and debug. You fix it by making the conflicting access atomic and properly published: synchronization (locks/synchronized), atomic classes, or by removing shared mutable state (immutability, confinement). Classic forms are check-then-act and read-modify-write.

code

java · 9 lines
java
// count++ is NOT atomic. Two threads can lose an update.
class Counter {
    private int count = 0;        // shared mutable state
    void increment() { count++; } // read, add, write -> 3 steps
    int get() { return count; }
}

// 1000 increments from many threads may yield < 1000.
// Fix: AtomicInteger.incrementAndGet() or a synchronized block.

go deeper

for a junior

Can define a race condition as two threads touching shared data with bad timing, and recognize count++ isn't atomic.

for a middle

Explains the read-modify-write and check-then-act shapes, names the required ingredients (shared mutable state, a writer, no sync), and knows the fixes (synchronized/atomic/immutable).

for a senior

Reasons about specific interleavings, explains why the bug is nondeterministic and data-dependent, and distinguishes a race from related issues (visibility, deadlock).

for a principal

Frames races as a design property — unsynchronized shared mutable state — and drives them out by architecture (immutability, confinement, message passing) rather than sprinkling locks, and reasons about it under the Java Memory Model.

## What problem are we even talking about? A **thread** is an independent path of execution inside your program; modern programs run several threads at once so they can do work in parallel. **Shared mutable state** is any data (a field, an array, a counter, a `HashMap`) that more than one thread can both read and *change*. A **race condition** is a bug where the *correctness* of the program depends on the exact order in which threads happen to run — and that order is not under your control. ### Why order is not under your control Threads are scheduled by the operating system and the JVM. The **scheduler** can pause any thread between (almost) any two operations and let another thread run. The set of all possible orderings of the threads' steps is called an **interleaving**. With even two threads doing a few steps each, there are many possible interleavings. A correct concurrent program must be correct under *every* interleaving; a racy one is correct under *most* but wrong under a few. Because the scheduler picks interleavings effectively at random (depending on CPU load, timing, number of cores), the bug appears **nondeterministically** — it works 999 times and fails the 1000th. ### A concrete example: the lost update Consider `count++` where `count` is a shared `int`. This single line is *not* one indivisible operation. It is really three steps: 1. **Read** the current value of `count` from memory into the CPU. 2. **Add** 1 to that value. 3. **Write** the new value back to `count`. Now suppose two threads, A and B, both run `count++` when `count` is 0: - A reads 0. - B reads 0 (A hasn't written yet). - A computes 1, writes 1. - B computes 1, writes 1. Two increments happened, but `count` is 1, not 2. One update was **lost**. This is a **read-modify-write** race: the read, the modify, and the write must happen as one indivisible unit (be **atomic**) but weren't. ### The other classic shape: check-then-act ``` if (map.get(key) == null) { // CHECK map.put(key, compute()); // ACT } ``` Thread A checks (key absent), then thread B checks (still absent), then both act — both compute and put, doing the work twice or overwriting each other. The decision (the *check*) became stale before the *act* used it. **Lazy initialization** (`if (instance == null) instance = new Thing();`) is the most famous check-then-act bug: two threads can both see null and both create the object. ### Why "data-dependent" Whether the bad interleaving actually corrupts anything often depends on the *values* and the *load*. A counter only loses updates when two increments truly overlap; a `HashMap` only corrupts its internal structure (possible infinite loop, lost entries) when a resize happens during a concurrent write. Low load → no overlap → looks fine in testing; production load → overlap → corruption. That's why these bugs hide in tests and surface in production. ### Three operations, one rule The root cause is always the same: **a compound operation on shared mutable state that must be atomic, but isn't, with no coordination between threads.** Note that not every shared access is a problem — concurrent *reads* of unchanging data are fine. You need a race only when there's shared state, *at least one writer*, and no synchronization ordering the accesses. ### How you fix it (preview) 1. **Don't share** — confine the state to one thread, or make it **immutable** (can't change → can't race). 2. **Make the compound op atomic** — `synchronized`, an explicit `Lock`, or atomic classes (`AtomicInteger.incrementAndGet()`, `ConcurrentHashMap.computeIfAbsent`). 3. **Publish safely** — ensure that when one thread writes, others actually *see* the new value (this is the related *visibility* problem, handled by `volatile`/locks under the Java Memory Model). A race condition is fundamentally a *design* defect (unsynchronized shared mutable state), not a typo, which is why "just add a lock here" sometimes only narrows the window instead of closing it.

  • If a race only sometimes corrupts data, why is it still a serious bug?
    Because correctness now depends on luck (scheduling and load). It will eventually hit the bad interleaving in production, under exactly the high-load conditions where it does the most damage, and it's nearly impossible to reproduce on demand for debugging.
  • Does having multiple CPU cores cause races, or just expose them?
    Races exist conceptually even on one core because the scheduler can preempt between steps; multiple cores make them more likely and add memory-visibility effects, so they expose races more often, but the root cause is unsynchronized shared mutable state, not the core count.

saying these in an interview costs you the question

  • Saying count++ is a single atomic operation
  • Claiming concurrent reads of immutable data cause races
  • Thinking a race always crashes loudly (it usually corrupts silently)
  • Believing it reproduces reliably so it can't be a race
  • Confusing a race condition with a deadlock (different bug)

context

open as a page

Explain the check-then-act race condition with a concrete Java example, and show why it is unsafe even with thread-safe collections.

level: middleimportance: must knowfreq 72%

basics

~20 s

Check-then-act is when you test a condition and then act on it as separate steps. Another thread can change things between the check and the act, so your action is based on stale information. Example: "if absent, put" can let two threads both put.

open as a page

What techniques make a read-modify-write operation (like an increment) atomic in Java, and what are the trade-offs between them?

level: seniorimportance: should knowfreq 64%

basics

~20 s

Make the read, change, and write happen as one indivisible step. You can use a synchronized block or a Lock, or an atomic class like AtomicInteger with incrementAndGet, or a concurrent collection's atomic methods like compute. Each guarantees no other thread interleaves in the middle.

open as a page

How would you detect, reproduce, and design out race conditions in a Java codebase before they reach production?

level: principalimportance: should knowfreq 48%

basics

~20 s

Reduce shared mutable state first (immutability and confinement), so most code can't race. For what's left, use thread-safe constructs, review for unsynchronized shared access, and use stress/concurrency tests and analysis tools (jcstress, static analyzers) to surface the rare interleavings.

open as a page