skip to content

What is a memory fence (barrier), what do the four barrier kinds — LoadLoad, LoadStore, StoreStore and StoreLoad — each prevent, and why is one of them far more expensive than the others?

level: seniorimportance: should knowfreq 40%

answer

  1. four pairs: LL, SS, LS, SL
  2. StoreLoad = drain the store buffer = expensive
  3. acquire = LoadLoad + LoadStore after a load
  4. release = StoreStore + LoadStore before a store
  5. compiler barrier ≠ hardware barrier — need both

basics

~20 s

A fence forbids specific reorderings across a point in program order. LoadLoad keeps earlier loads before later loads; StoreStore keeps earlier stores before later stores; LoadStore keeps a load before a later store; StoreLoad keeps an earlier store visible before a later load — that one must drain the store buffer, so it is the costly full fence.

solid answer

~60 s

A fence is an instruction (and simultaneously a compiler directive) that forbids reordering of certain memory operations across it. The classic taxonomy names the pair it constrains: - **LoadLoad** — loads before the fence complete before loads after it. - **StoreStore** — stores before the fence become visible before stores after it. - **LoadStore** — loads before the fence complete before later stores become visible. - **StoreLoad** — stores before the fence become globally visible before later loads execute. StoreLoad is the expensive one: satisfying it means stalling until the core's store buffer has drained and the write is visible to other cores, which is the very optimization the buffer exists to provide. The others usually cost little or nothing on strong architectures. Composed: a release barrier is LoadStore+StoreStore placed before a store; an acquire barrier is LoadLoad+LoadStore placed after a load; a full fence is all four. Note also that a *compiler* barrier only stops code motion at compile time and emits no instruction — you generally need both.

code

text · 5 lines
text
Publisher:                     Consumer:
  obj.field = 42                 p = shared_ptr      // acquire load
  StoreStore barrier             LoadLoad barrier
  shared_ptr = obj               if (p) read p.field
  // = release store             // = acquire load

go deeper

for a junior

Know that a fence stops reordering across a point and that ordinary code should use locks rather than fences directly.

for a middle

Name the four kinds and give one concrete use each, especially StoreStore for publishing an object and LoadLoad for reading a pointer then its target.

for a senior

Explain why StoreLoad alone is expensive (store-buffer drain), decompose acquire and release into barrier pairs, and separate the compiler and hardware halves.

for a principal

Position explicit fences as an encapsulated, verified implementation detail of a small library, and discuss the portability cost of code tuned to a strongly ordered architecture.

## What a fence is A memory fence (barrier) is a point in the instruction stream across which certain memory operations may not be moved. It constrains two things at once: the compiler, which must not migrate the affected accesses across it, and the hardware, which must not complete them out of the required order. In practice a language-level ordered operation compiles into whatever combination of compiler restriction plus machine barrier instruction the target needs. A fence orders operations; it does not 'flush caches' and it does not push values anywhere by itself. Caches are already coherent. What a fence really controls is when a core is permitted to let its own operations become globally visible relative to each other. ## The four kinds Name each barrier by the ordered pair `A→B` it enforces: nothing of type A before the fence may be observed after anything of type B that follows it. **LoadLoad.** Earlier loads must complete before later loads. Needed when you read a pointer or an index and then dereference or read the data it selects; without it, a weakly ordered core (or the compiler) may issue the dependent-looking second load speculatively from a stale value. **StoreStore.** Earlier stores become visible before later stores. This is what makes the publication idiom work: fill an object's fields, StoreStore, then publish the reference. Without it, the reference can be visible while the fields are not. **LoadStore.** Earlier loads complete before later stores are visible. This prevents work done inside a critical section from being observed after the lock has been given up. **StoreLoad.** Earlier stores are globally visible before later loads execute. This is the only barrier that forbids the store-buffering pattern (write my flag, then read yours) and it is the only one that x86-class hardware does not give you for free. ## Why StoreLoad is the expensive one Store buffers exist precisely so a core can retire a store and continue without waiting for cache-line ownership. A StoreLoad barrier says: no more running ahead — the buffered stores must be globally visible before the next load is satisfied. That stalls the pipeline for the full write-visibility latency, typically tens of cycles, and it cannot be hidden. The other three barriers are usually satisfied by the architecture's own rules on strong models and cost nothing; on weak models they still map to lighter-weight instructions than a full fence. ## How the barriers compose into the ordering people actually use - **Release** (paired with a store): LoadStore + StoreStore *before* the store. Everything you did earlier is visible before the store is. - **Acquire** (paired with a load): LoadLoad + LoadStore *after* the load. Nothing you do afterwards is observed before the load. - **Full fence / sequential consistency:** all four, StoreLoad included. That decomposition explains a fact that surprises people: a release store followed by an acquire load in the same thread does *not* give you a full fence, because neither includes StoreLoad. Hence the store-buffering litmus outcome remains legal under acquire/release. ## Compiler barriers versus hardware barriers A compiler-only barrier stops the optimizer from moving accesses but emits zero instructions; on a weakly ordered CPU the hardware still reorders. A hardware barrier without the compiler restriction is equally useless, because the compiler may have already hoisted the load above the fence at compile time. Every correct ordering primitive supplies both. When someone proposes a hand-written 'do not optimize this' hack, that is usually the missing half. ## When you should be writing fences at all Almost never directly. Locks, ordered atomic operations and higher-level synchronization emit the right barriers for the target architecture, and they document intent far better than a bare fence. Explicit fences belong in the internals of a lock-free data structure, written once, tested with a race detector on both strong and weak hardware, and encapsulated behind a normal API. ## Interview delivery Give the four names with their `A→B` meanings, single out StoreLoad as the store-buffer-draining full fence, then show that acquire = LoadLoad+LoadStore and release = StoreStore+LoadStore. Mentioning the compiler-versus-hardware split as two halves of the same requirement lands well.

  • Why is a compiler-only barrier not sufficient on ARM but often appears to work on x86?
    x86's TSO model already forbids load→load, store→store and load→store reordering in hardware, so on x86 the only reordering left for those cases is the compiler's, and stopping the compiler can be enough. ARM permits all of those in hardware, so the same code needs real barrier instructions. This is the classic reason code 'ported and broke'.
  • If fences only order operations, what actually makes another core see the new value?
    Cache coherence. Once a store leaves the store buffer and enters the coherent cache hierarchy, other cores are guaranteed to see it — they cannot keep a stale copy indefinitely. Fences do not push data; they control the order in which your core's operations become visible relative to one another.

saying these in an interview costs you the question

  • Describing a fence as 'flushing the cache' or 'writing through to RAM'
  • Believing all four barrier kinds cost the same
  • Thinking a compiler barrier alone gives hardware ordering
  • Assuming a release store followed by an acquire load equals a full fence
  • Reaching for explicit fences in ordinary application code instead of locks or ordered atomics

context