What is a race condition in Java, and why does it lead to nondeterministic, data-dependent bugs?
answer
- Timing-dependent correctness = race
- count++ is read-modify-write (3 steps)
- check-then-act / lazy init
- Needs: shared + mutable + >=1 writer + no sync
- Nondeterministic + data-dependent = hard to reproduce
basics
~20 sA race condition is when two or more threads touch the same shared data at the same time without coordination, and the result depends on which thread happens to run first. Because timing varies, you get different, sometimes wrong, results each run.
solid answer
~40 sA race condition occurs when the correctness of a program depends on the relative timing or interleaving of multiple threads accessing shared mutable state, and at least one access is a write. The threads aren't synchronized, so the OS/JVM scheduler is free to interleave their operations in any order. Most runs are fine; occasionally the operations interleave badly and you read stale data or lose an update. The bug is nondeterministic (depends on scheduling) and data-dependent (only shows up under certain values and load), which makes it hard to reproduce and debug. You fix it by making the conflicting access atomic and properly published: synchronization (locks/synchronized), atomic classes, or by removing shared mutable state (immutability, confinement). Classic forms are check-then-act and read-modify-write.
code
java · 9 lines// count++ is NOT atomic. Two threads can lose an update.
class Counter {
private int count = 0; // shared mutable state
void increment() { count++; } // read, add, write -> 3 steps
int get() { return count; }
}
// 1000 increments from many threads may yield < 1000.
// Fix: AtomicInteger.incrementAndGet() or a synchronized block.go deeper
Can define a race condition as two threads touching shared data with bad timing, and recognize count++ isn't atomic.
Explains the read-modify-write and check-then-act shapes, names the required ingredients (shared mutable state, a writer, no sync), and knows the fixes (synchronized/atomic/immutable).
Reasons about specific interleavings, explains why the bug is nondeterministic and data-dependent, and distinguishes a race from related issues (visibility, deadlock).
Frames races as a design property — unsynchronized shared mutable state — and drives them out by architecture (immutability, confinement, message passing) rather than sprinkling locks, and reasons about it under the Java Memory Model.
## What problem are we even talking about? A **thread** is an independent path of execution inside your program; modern programs run several threads at once so they can do work in parallel. **Shared mutable state** is any data (a field, an array, a counter, a `HashMap`) that more than one thread can both read and *change*. A **race condition** is a bug where the *correctness* of the program depends on the exact order in which threads happen to run — and that order is not under your control. ### Why order is not under your control Threads are scheduled by the operating system and the JVM. The **scheduler** can pause any thread between (almost) any two operations and let another thread run. The set of all possible orderings of the threads' steps is called an **interleaving**. With even two threads doing a few steps each, there are many possible interleavings. A correct concurrent program must be correct under *every* interleaving; a racy one is correct under *most* but wrong under a few. Because the scheduler picks interleavings effectively at random (depending on CPU load, timing, number of cores), the bug appears **nondeterministically** — it works 999 times and fails the 1000th. ### A concrete example: the lost update Consider `count++` where `count` is a shared `int`. This single line is *not* one indivisible operation. It is really three steps: 1. **Read** the current value of `count` from memory into the CPU. 2. **Add** 1 to that value. 3. **Write** the new value back to `count`. Now suppose two threads, A and B, both run `count++` when `count` is 0: - A reads 0. - B reads 0 (A hasn't written yet). - A computes 1, writes 1. - B computes 1, writes 1. Two increments happened, but `count` is 1, not 2. One update was **lost**. This is a **read-modify-write** race: the read, the modify, and the write must happen as one indivisible unit (be **atomic**) but weren't. ### The other classic shape: check-then-act ``` if (map.get(key) == null) { // CHECK map.put(key, compute()); // ACT } ``` Thread A checks (key absent), then thread B checks (still absent), then both act — both compute and put, doing the work twice or overwriting each other. The decision (the *check*) became stale before the *act* used it. **Lazy initialization** (`if (instance == null) instance = new Thing();`) is the most famous check-then-act bug: two threads can both see null and both create the object. ### Why "data-dependent" Whether the bad interleaving actually corrupts anything often depends on the *values* and the *load*. A counter only loses updates when two increments truly overlap; a `HashMap` only corrupts its internal structure (possible infinite loop, lost entries) when a resize happens during a concurrent write. Low load → no overlap → looks fine in testing; production load → overlap → corruption. That's why these bugs hide in tests and surface in production. ### Three operations, one rule The root cause is always the same: **a compound operation on shared mutable state that must be atomic, but isn't, with no coordination between threads.** Note that not every shared access is a problem — concurrent *reads* of unchanging data are fine. You need a race only when there's shared state, *at least one writer*, and no synchronization ordering the accesses. ### How you fix it (preview) 1. **Don't share** — confine the state to one thread, or make it **immutable** (can't change → can't race). 2. **Make the compound op atomic** — `synchronized`, an explicit `Lock`, or atomic classes (`AtomicInteger.incrementAndGet()`, `ConcurrentHashMap.computeIfAbsent`). 3. **Publish safely** — ensure that when one thread writes, others actually *see* the new value (this is the related *visibility* problem, handled by `volatile`/locks under the Java Memory Model). A race condition is fundamentally a *design* defect (unsynchronized shared mutable state), not a typo, which is why "just add a lock here" sometimes only narrows the window instead of closing it.
- If a race only sometimes corrupts data, why is it still a serious bug?Because correctness now depends on luck (scheduling and load). It will eventually hit the bad interleaving in production, under exactly the high-load conditions where it does the most damage, and it's nearly impossible to reproduce on demand for debugging.
- Does having multiple CPU cores cause races, or just expose them?Races exist conceptually even on one core because the scheduler can preempt between steps; multiple cores make them more likely and add memory-visibility effects, so they expose races more often, but the root cause is unsynchronized shared mutable state, not the core count.
saying these in an interview costs you the question
- Saying count++ is a single atomic operation
- Claiming concurrent reads of immutable data cause races
- Thinking a race always crashes loudly (it usually corrupts silently)
- Believing it reproduces reliably so it can't be a race
- Confusing a race condition with a deadlock (different bug)