skip to content

Why does counter.value += 1 on a multiprocessing.Value race, and how do you fix it?

level: middleimportance: should knowfreq 45%

answer

  1. Shared visibility is not atomicity
  2. It is a read, an add and a write
  3. The GIL does not cross processes
  4. The wrapper guards each half, not the pair
  5. Widen the critical section with get_lock()

basics

~10 s

counter.value += 1 is a read, an add and a write, and another process can land between them, so increments are lost. Hold the value's own lock around the whole read-modify-write with with counter.get_lock():.

solid answer

~50 s

`multiprocessing.Value("i", 0)` puts one C integer in a memory block mapped into every process, so writes really are visible everywhere — but visibility is not atomicity. `counter.value += 1` compiles to a load of the attribute, an addition, and a store back; two processes can both load 41, both compute 42, and one increment vanishes. By default `Value` is created with `lock=True`, which wraps the raw storage in a synchronized object whose individual `.value` reads and writes are guarded — but a `+=` is *two* guarded operations with an unguarded gap between them. The fix is to widen the critical section over the whole sequence: `with counter.get_lock(): counter.value += 1`. The same reasoning applies element-wise to `multiprocessing.Array`. And the cheaper design is usually to avoid the shared counter entirely: let each worker count locally and contribute one total at the end.

code

python · 15 lines
python
import multiprocessing as mp

def bump(counter, n):
    for _ in range(n):
        with counter.get_lock():
            counter.value += 1

if __name__ == "__main__":
    counter = mp.Value("i", 0)
    procs = [mp.Process(target=bump, args=(counter, 10_000)) for _ in range(4)]
    for p in procs:
        p.start()
    for p in procs:
        p.join()
    print(counter.value)   # 40000, every run

go deeper

for a junior

Know that a shared counter needs a lock and that with counter.get_lock(): around the increment is the standard fix; you are not expected to reason about the bytecode.

for a middle

Explain the read-modify-write decomposition, why a default synchronized Value still races on +=, and why the GIL is irrelevant once separate processes are involved.

for a senior

Show judgment about contention: measure the lock cost in a hot loop, accumulate locally or shard per worker, and be able to say when RawValue is a defensible optimization.

for a principal

Argue about whether shared mutable state across processes should exist at all in the design, versus returning partial results and aggregating once — and what that choice costs in complexity and in failure modes.

### What Value gives you and what it does not `multiprocessing.Value(typecode, initial)` allocates a small block of memory that the operating system maps into the parent and every child, and stores one C-typed value there — `"i"` for a C int, `"d"` for a double, and so on. Because the block is genuinely shared, a store by one process is visible to the others. That is the part people expect. What it does not give you is atomicity of compound operations. `counter.value += 1` is three steps: read the current value out of shared memory, add one in the local process, write the result back. Between the read and the write, any other process may run the same three steps. Both read 41, both write 42, and one unit of work disappears. Run four processes each adding 25,000 with no synchronization and the total is reliably short of 100,000 and different every run. A common confusion is to reach for the GIL here. The GIL serializes bytecode **within one interpreter**. Separate processes have separate interpreters and separate GILs, so it offers no protection at all across processes. Even within one process it would not help, because the interleaving can happen between bytecodes. ### The lock that is already there By default `Value` is created with `lock=True`, which means the object you get back is a *synchronized wrapper* around the raw storage. Reading `.value` and assigning `.value` each take the lock internally, so a single read or a single write is safe and you will never see a half-written value. But `+=` is a read *and* a write, and the wrapper releases the lock in between. Guarding each half does not guard the pair. The wrapper exposes that lock, and the fix is to take it around the whole sequence: ```python with counter.get_lock(): counter.value += 1 ``` Now the read, the add and the write happen inside one critical section, and no other process can observe or modify the intermediate state. The same wrapper also exposes the underlying raw object for the rare case where you want to work on it directly. Three variations are worth knowing. Passing `lock=False` gives you an unsynchronized value with no `get_lock()` at all — you must supply your own `multiprocessing.Lock` or accept the race. `multiprocessing.RawValue` and `multiprocessing.RawArray` are the never-synchronized forms, useful when you have proved that only one process writes. And you may pass an existing lock, which is how several `Value` objects can be updated together under one lock so that a pair of counters stays mutually consistent. ### Arrays are the same story, per element `multiprocessing.Array("i", 10)` is a fixed-length block with one lock for the whole array. `arr[3] += 1` has the identical read-modify-write hazard, and `with arr.get_lock():` guards it. Note that the single lock also serializes unrelated elements, so a workload where every worker hammers a different slot of the same array will contend on a lock it does not logically need. That is a reason to shard: give each worker its own `Value`, or its own slice, and combine at the end. ### The cost, and the design that avoids it A lock acquisition here is a real OS-level synchronization primitive — a semaphore, not a cheap in-process flag. Taking it once per increment in a hot loop can dominate the runtime and can make the multiprocess version slower than a single process, which is a genuinely embarrassing benchmark result to explain in an interview. So the senior answer is not only "use `get_lock()`" but "reduce the number of times you need it". Standard patterns: * **Local accumulation.** Each worker keeps a plain Python integer and touches the shared `Value` once, at the end. One lock acquisition per worker instead of per item. * **Return the number instead of sharing it.** Have workers return partial counts through the pool's result path and sum them in the parent. No shared mutable state exists at all, which is the cheapest kind to reason about. * **Shard the storage.** One slot per worker in an `Array`, written only by its owner, summed by the parent after the join. No lock on the hot path because there is exactly one writer per slot. ### The vocabulary trap Interviewers often probe whether you know *which* lock is in play. `threading.Lock` protects threads inside one process and is not even picklable, so it cannot be handed to a child. `multiprocessing.Lock` is backed by an OS semaphore and works across processes. `asyncio.Lock` protects coroutines on one event loop and does nothing here. The lock behind `Value.get_lock()` is the multiprocessing kind. Naming that distinction cleanly is usually worth more in the interview than the fix itself, because it shows you understand *why* the shared block needs its own primitive rather than borrowing one from the threading world.

  • Value is created with lock=True by default, so why isn't += already safe?
    Because the default wrapper guards each individual `.value` read and each `.value` assignment, not the pair. `+=` reads under the lock, releases it, computes, then writes under the lock again — and another process can slip into that gap. Atomicity has to span the whole read-modify-write, which is what `with counter.get_lock():` provides.
  • Does the GIL protect a shared counter across processes?
    No. The GIL serializes bytecode inside a single interpreter; separate processes have separate interpreters and separate GILs, so it imposes no ordering between them at all. Cross-process mutual exclusion needs an OS-level primitive, which is what `multiprocessing.Lock` and the lock inside a synchronized `Value` are built on.
  • When would you deliberately choose RawValue or lock=False?
    When the access pattern removes the need for mutual exclusion — typically one designated writer per location and readers that only look after a join or an explicit barrier, or a worker writing exclusively into its own slot of a sharded array. You are then paying no synchronization cost on the hot path. It is a claim you have to be able to defend, because nothing will detect a violation for you.

Two clerks both read a wall counter showing 41, both write 42. The number is genuinely shared; what was missing was the rule that only one clerk may be at the wall from read to write.

saying these in an interview costs you the question

  • Believing shared memory makes += atomic
  • Thinking the GIL serializes writes across processes
  • Assuming lock=True already guards a compound update
  • Taking a lock per item in a hot loop without noticing the cost
  • Passing a threading.Lock to a child process

context