A concurrency bug reproduces in production but vanishes the moment you add logging around it or attach a debugger. Why does that happen, and how do you gather evidence without destroying the conditions you need?
answer
- probe effect: observation shifts timing
- log call = lock + I/O = accidental barrier
- debugger breakpoint serialises threads
- sample, count, ring-buffer, post-mortem dump
- 'fixed by adding logging' = not fixed
basics
~20 sObservation changes timing. Logging and breakpoints add delay and internal synchronisation that widen or close the racy window, so the interleaving that triggered the bug stops occurring. Collect evidence with low-overhead, always-on instruments instead: sampling, counters, in-memory ring buffers, and post-mortem dumps.
solid answer
~50 sThis is the probe effect, and the resulting bug is called a heisenbug. Two mechanisms cause it. First, **timing**: a log statement costs orders of magnitude more than the racing operation, so it shifts the relative arrival of threads and the narrow window simply stops being hit. A debugger is worse — a breakpoint suspends threads and serialises them entirely. Second, **synchronisation as a side effect**: logging takes an internal lock, writes to shared state and performs I/O, which introduces ordering and memory barriers where the buggy code had none. A stale-read or reordering bug can be masked purely by the barrier the logger implies. So I avoid perturbing instrumentation on the hot path: sampling profilers rather than per-event logs, cheap counters and metrics, a lock-free in-memory ring buffer dumped only after a failure is detected, and automatic thread-dump or core-dump capture triggered by a watchdog. Then I try to reproduce the interleaving deterministically off the hot path.
code
text · 12 linesWITHOUT logging (window ~ nanoseconds):
A: read balance (100)
B: read balance (100) <-- both read stale value
A: write 90
B: write 80 <-- lost update
WITH a log line after the read:
A: read balance (100)
A: log(...) ~microseconds, takes appender lock, flushes
A: write 90
B: read balance (90) <-- B now arrives after A finished
B: write 80 <-- correct; bug 'disappears'go deeper
Explain that observation costs time and the racy window is tiny, so adding a log line or pausing at a breakpoint changes the interleaving and hides the bug.
Add the second mechanism — logging implies locks, shared state and memory barriers — and name low-overhead alternatives such as counters, sampling and post-mortem dumps.
Discuss ring-buffer tracing, watchdog-triggered capture, invariant assertions, and moving reproduction off the hot path with forced delays, randomised interleaving or deterministic replay.
Treat perturbation budget as a design constraint: decide what is always on and at what overhead, ensure evidence is captured automatically before recovery actions, and insist that timing-masked bugs are fixed at the synchronisation level rather than papered over.
## The probe effect A heisenbug is a defect whose behaviour changes when you try to observe it. In concurrency it is the norm rather than the exception, because the bug's occurrence depends on a timing window that observation itself perturbs. **Cost asymmetry.** The racy region might be a few nanoseconds — two instructions between a read and a write. A formatted log line is typically hundreds of nanoseconds to microseconds, and with I/O far more. Inserting one is like inserting a mountain into the gap you were trying to hit; threads no longer arrive at the window together. **Accidental synchronisation.** Logging is not passive. It usually takes a lock on an appender or buffer, touches shared mutable state, and performs I/O. Every one of those implies memory ordering: caches get flushed, barriers get issued, operations that were free to be reordered no longer are. A visibility bug — thread B reading a stale value written by thread A — can be masked entirely by the barrier a nearby log call implies, even though nothing in the buggy code changed. **Compiler and runtime effects.** Adding a statement can change what the optimiser does: a value cached in a register may now be re-read from memory, loop-hoisted code may stop being hoisted, an inlined method may stop being inlined. All of these can eliminate the reordering or caching that produced the bug. **Debuggers are the extreme case.** A breakpoint suspends threads. Stepping serialises execution completely, which is the one interleaving in which no race can occur. Even a conditional breakpoint evaluates a predicate on every hit, adding large delays. Attaching a debugger to a race is often self-defeating. **It can also flip the other way.** Instrumentation sometimes *creates* failures — for example, extra delay makes a previously-tight timeout expire, or the logging lock becomes the new contention point. "It only fails with tracing on" is just as much a probe effect. ## Gathering evidence without perturbing The principle is: make observation cheap, constant-cost and off the critical path. - **Sampling instead of tracing.** A sampling profiler perturbs at a fixed low rate regardless of event frequency, whereas per-event logging scales its perturbation with the very code you are studying. - **Counters and metrics.** Incrementing a per-thread counter and aggregating periodically is far cheaper than emitting an event, and often enough to answer "did this branch ever run?" or "how many retries occurred?". - **In-memory ring buffers.** Record small fixed-size records (timestamp, thread, event id) into a preallocated per-thread ring with no locking or formatting, and only serialise the buffer *after* a failure is detected. This preserves ordering evidence at a fraction of logging's cost. - **Post-mortem capture.** A process dump, core dump or heap snapshot taken at the moment of failure holds the state without having watched it get there. Combine with a watchdog that triggers capture on a health-check failure or a stalled progress counter. - **Assertions on invariants.** A cheap check that a value is within a legal range costs a comparison and fires at the moment corruption occurs, rather than requiring you to watch every step. - **Post-hoc correlation.** Existing traces, metrics and logs collected before the incident are non-perturbing by definition. Mining them for coincidence — a deploy, a load spike, a dependency slow-down — is often more productive than adding new instrumentation. ## Reproducing off the hot path Once you have a hypothesis about the interleaving, move the investigation somewhere perturbation is allowed: - **Force the window open deliberately** in a test build — insert delays or yield points precisely where you believe the race is, which makes an improbable interleaving near-certain. This is intentional probing, applied where it helps rather than where it hides. - **Randomised interleaving** in tests: schedule-perturbation tools that inject random preemptions explore many orderings quickly. - **Deterministic replay or a controllable scheduler**, which lets you script the exact suspected ordering and verify whether the bad state is reachable. - **Run on different hardware.** Memory-ordering effects vary between architectures; a bug that never shows on a strongly-ordered machine may appear immediately on a weakly-ordered one, and vice versa. ## The judgment to state Disappearing under observation is *evidence*, not an annoyance: it strongly suggests a genuine timing or memory-ordering dependency rather than a logic error. Treat the disappearance as a diagnostic clue that narrows the hypothesis space, and never treat "we added logging and it stopped happening" as a fix — the window is still there, merely harder to hit, and it will return under different load or on different hardware.
- Someone reports that the bug disappeared after they added detailed logging, and proposes shipping that logging as the fix. How do you respond?The race is still present; the logging merely made the window harder to hit by adding delay and implicit synchronisation. It will resurface under different load, on different hardware, or after any change that alters timing — and now it will be even rarer and harder to diagnose. Treat the disappearance as evidence of a timing or memory-ordering dependency and fix the underlying synchronisation, then decide on logging independently for its own value.
- What instrumentation would you leave permanently enabled so a heisenbug is diagnosable the next time it occurs?Constant-cost, non-serialising instruments: a low-rate sampling profiler, per-thread counters for the branches and retries you care about, progress metrics per pool, and a preallocated lock-free ring buffer whose contents are serialised only after a failure is detected. Pair those with a watchdog that automatically captures thread dumps and, where feasible, a process dump when health checks or progress metrics stall. The cost is a small fixed percentage rather than a cost proportional to event frequency.
- Is deliberately inserting delays into the code ever a legitimate technique here?Yes, but in test builds rather than production. Inserting a yield or short delay exactly where you suspect the race widens the window and turns an improbable interleaving into a near-certain one, which converts an unreproducible bug into a reliable failing test. This is the probe effect used deliberately in the direction that helps. It is a diagnostic and regression-locking tool, not something to ship.
Measuring the temperature of a drop of water with a large cold thermometer: the instrument changes the thing it measures, so the reading describes the measurement, not the drop.
saying these in an interview costs you the question
- Declaring the bug fixed because it stopped reproducing once logging was added.
- Attaching a debugger to a race and concluding there is no bug when stepping shows correct behaviour.
- Assuming only timing changes matter, and missing that logging introduces locks and memory barriers.
- Adding per-event tracing on the hot path, whose overhead scales with the very events under investigation.
- Restarting the process before capturing dumps, then hoping the next occurrence is easier.