A service with an unsynchronized shared flag has run correctly for years, then starts failing after the team moves to a different CPU architecture and raises the compiler optimization level. Explain how code with a data race can appear correct for years and then break, and what you would do about it.
answer
- a race has no guarantee — only luck
- TSO forbids 3 of 4 reorderings; ARM forbids none
- optimizer + JIT expose it later, on hot paths
- nanosecond window; logging closes it (heisenbug)
- detector first, then real synchronization
basics
~20 sA race has no defined behaviour; it only ever appeared to work. Strongly ordered CPUs and low optimization levels hide most reordering, and the failure window is nanoseconds wide. A weaker memory model plus a more aggressive optimizer exposes what was always permitted. Fix the synchronization; do not tune around the symptom.
solid answer
~50 sRacy code never had a guarantee — it had *luck*, supplied by three things. First, the old architecture: a total-store-order CPU forbids three of the four reorderings, so most racy publication patterns happen to work there, while ARM/POWER-class cores permit them. Second, the optimizer: at low levels the compiler leaves loads and stores roughly where you wrote them; at higher levels it hoists loads into registers, sinks stores and unrolls loops, and a JIT does the same only after a method gets hot — so failures can appear minutes into a run. Third, timing: the vulnerable window is often nanoseconds, so tests almost never hit it, and adding logging or a debugger closes it. The response: reproduce under a thread sanitizer / race detector rather than by stress alone, then fix the race with real synchronization — a lock or an ordered atomic — not with sleeps, extra logging, larger buffers or lowered optimization. Add the detector to CI so the class of bug cannot return.
go deeper
Say that a data race is undefined behaviour and that seeming to work is not the same as being correct; the fix is proper synchronization.
Name the concrete maskers — strongly ordered CPU, low optimization level, narrow timing window — and explain how each disappears after the change.
Drive a diagnosis: race detector rather than stress reproduction, read for the unsynchronized-shared-state shape, reproduce on the weak architecture, then fix with a lock or ordered atomic and reject the non-fixes explicitly.
Address blast radius and process: how much other code shares this shape, containment versus cure, detectors in CI, and a structural push toward ownership/message passing so the class of bug cannot recur.
## The premise to correct first 'It worked for years' is not evidence of correctness. A data race — two threads accessing the same location, at least one writing, with no synchronization ordering them — has no defined behaviour in any modern memory model. The program was never guaranteed to work; the platform simply happened to produce the desired outcome. When the platform changes, the guarantee that was never there stops being simulated. ## Why the old platform hid it **Strong hardware.** x86-class cores implement total store order: they forbid store→store, load→load and load→store reordering, leaving only store→load. Most naive publication and flag-check patterns depend on exactly the reorderings TSO forbids, so they work by accident. ARM, POWER and RISC-V permit all four, so the same binary logic fails immediately — often at low load, not just under stress. **Low optimization.** At `-O0`-style settings the compiler keeps loads and stores near where you wrote them and rarely caches a shared variable in a register. Turn optimization up and it hoists the flag load out of the spin loop (infinite loop), sinks or merges stores (publication inversion), or eliminates a 'redundant' re-read entirely. Managed runtimes add a schedule dimension: the interpreter is conservative and the JIT is not, so a method that ran correctly ten thousand times starts failing when it is compiled and inlined. **Narrow windows and observer effects.** The interleaving that fails may require two cores to hit specific instructions within a few nanoseconds. A test suite may execute the path a million times and never lose the race. Then production changes core count, adds a faster disk, removes a log line — and the window opens. This is also why the bug 'disappears when I add a print statement': the printf is a synchronization-heavy, slow operation that changes the timing and often incidentally emits barriers. ## Diagnosis 1. **Do not start by reproducing under load.** Run the workload under a dynamic race detector (thread-sanitizer-class tooling exists for most ecosystems). It reports races by *analysis*, flagging the unsynchronized pair even when the bad interleaving did not occur. 2. **Read the code for the shape**, not the symptom: a shared flag or pointer written by one thread and read by another with no lock, no ordered atomic, no queue in between. 3. **Reproduce on the weak architecture** with optimizations on — the environment where the reordering is legal — rather than trying to force it on the strong one. 4. **Check the diff for what actually changed**: often the racy code is old and the trigger was a compiler bump or an instance-type change, which is useful for scoping how much other code is at risk. ## The fix, and the non-fixes Fix it with real synchronization: guard the state with a lock, or make the flag an ordered atomic with a release write and acquire read, or replace shared mutable state with a message/queue handoff so only one thread owns the data. Non-fixes that teams reach for under pressure, and why each is wrong: - **Lowering the optimization level or pinning the old CPU** — hides the symptom, freezes you on obsolete infrastructure, and the race is still there. - **Adding sleeps or retries** — changes probability, not legality; fails again under different load. - **Adding logging** — accidental barriers plus slower timing; the classic heisenbug mask. - **Marking one side only** — ordering requires both a publishing write and a matching read; half a pair guarantees nothing. - **'We'll just make the loop longer/shorter'** — no relationship to the memory model at all. ## Preventing recurrence Run a race detector in CI on a concurrency-focused test suite, even if only nightly. Establish a rule that shared mutable state must be reached through a lock, an ordered atomic, or a queue, and make that a review checklist item. Where possible, remove the sharing: confine state to one owner and communicate by messages, which makes the whole class of bug unrepresentable. And treat every 'works on our machines' concurrency claim as untested until a detector agrees. ## Interview delivery Lead with 'a race is undefined behaviour, so it never worked — it was masked', name the three maskers (strong architecture, low optimization, narrow window), describe detector-first diagnosis, then fix with real synchronization and explicitly reject the tempting non-fixes. That structure is what makes this a senior answer rather than a war story.
- Why can a race detector report a bug even though the failing interleaving never happened during the run?Dynamic race detectors track the happens-before relationships between accesses rather than waiting for a bad outcome. If two threads touch the same location, at least one writes, and no synchronization orders them, the tool reports it regardless of which order actually occurred. That is why detection beats stress testing: you no longer need to win a nanosecond-wide race to see the defect.
- The team proposes pinning the old CPU family and compiler version until there is time to fix it. What do you say?It is an acceptable short-term containment measure only if it is time-boxed and tracked, because it does not remove the defect — it restores the conditions that were masking it. It also blocks routine infrastructure upgrades and leaves you exposed to any other change that shifts timing, such as core count or load. The real fix is small and local, so the containment should be measured in days.
saying these in an interview costs you the question
- Claiming the code was correct and the new hardware or compiler is buggy
- Fixing it by lowering optimization, adding sleeps, or adding log lines
- Believing a test suite that passes a million iterations proves the absence of a race
- Assuming every architecture provides the same ordering guarantees
- Synchronizing only the writer, or only the reader, and declaring the pair fixed