Compare acquire/release ordering, sequentially consistent ordering, and relaxed ordering for atomic operations: what does each guarantee, what does each cost, and when would you choose the weaker ones?
answer
- relaxed = atomicity only, no ordering
- release publishes prior writes; acquire receives them
- pairing on the same location or nothing happens
- acquire/release is one-way → store buffering still legal
- SC = one total order everyone agrees on, costs StoreLoad
basics
~20 sRelaxed guarantees atomicity only — no ordering with anything else. Acquire/release is a one-way pairing: a release store publishes everything written before it to any thread doing a matching acquire load. Sequential consistency adds a single total order all threads agree on, requiring the expensive full fence.
solid answer
~60 sThink of three tiers. **Relaxed:** the operation is indivisible (no torn or lost update) but is not ordered against any other memory operation. Correct only for values that carry no relationship to other data — an approximate counter, a statistics tally, a one-way stop flag whose payload you re-check. **Acquire/release:** a *pairing*. A release store forbids earlier operations from sinking below it; an acquire load forbids later operations from rising above it. When an acquire load observes the value written by a release store, everything the writer did before the release is visible to the reader after the acquire. This is exactly the publication/handoff pattern and it is what a lock's unlock/lock pair provides. It is one-way and cheap — often free on x86, and lightweight dedicated instructions on ARM. **Sequential consistency:** additionally there exists one global total order over all SC operations that every thread agrees on. It is the only tier that forbids the store-buffering outcome, and it costs a StoreLoad fence. Default to SC, drop to acquire/release for handoff-shaped code, use relaxed only where you can argue no ordering is needed.
code
text · 8 linesPublisher Consumer
obj.a = 1 p = load_acquire(ptr)
obj.b = 2 if (p != null)
store_release(ptr, obj) read p.a, p.b // guaranteed 1, 2
Break either half:
store_relaxed(ptr,obj) + load_acquire -> no guarantee
store_release(ptr,obj) + load_relaxed -> no guaranteego deeper
Know the three names and that stronger means more guarantees and more cost; in practice use locks and leave the tiers to library authors.
State the pairing rule for acquire/release with the publication example, and that relaxed gives atomicity only.
Add that acquire/release is one-way and therefore permits store buffering, that only SC forbids it at the price of a StoreLoad fence, and give the per-architecture cost picture.
Frame it as a policy: strongest by default, weaken only inside encapsulated, benchmarked, race-detector-verified components, with the ordering argument written down next to the code.
## Why there are tiers at all Ordering costs performance, and different algorithms need different amounts of it. A memory model that offered only the strongest guarantee would make every atomic operation pay for a full fence; one that offered only the weakest would make correct code impossible to write. So models offer a small ladder, and your job is to pick the lowest rung that still makes your algorithm correct. ## Relaxed: atomicity without ordering A relaxed atomic operation guarantees exactly one thing: the read or write happens indivisibly, so you never observe a torn value and a read-modify-write never loses a concurrent update. It says nothing about ordering relative to any other memory operation, in either thread. The compiler and hardware may move ordinary loads and stores freely across it. Legitimate uses are narrow: counters and metrics whose value is consumed only after the threads have joined; a monotone 'stop requested' flag where the reader re-validates everything it needs anyway; reference-count increments (the *decrement* that reaches zero still needs ordering before destruction). If any other data's visibility depends on the atomic, relaxed is wrong. ## Acquire/release: one-way ordering that comes in pairs Release and acquire are meaningful only as a matched pair on the same location. - A **release store** means: everything this thread did before, in program order, must be visible before this store is. - An **acquire load** means: nothing this thread does afterwards may be observed before this load. When the acquire load reads the value the release store wrote, the two threads are synchronized at that point, and the reader is guaranteed to see all the writer's prior work. That single property underpins publication (build an object, release-store the pointer, acquire-load it, use it), handoff through a queue slot, and initialization flags. The crucial word is **one-way**. A release does not stop *later* operations from moving up above it, and an acquire does not stop *earlier* operations from sinking below. Consequently two threads that each do a release store then an acquire load of the other's variable can still both read the old value — the store-buffering outcome — because neither side includes a StoreLoad barrier. Candidates who claim acquire/release 'is basically sequential consistency without the name' fail on exactly this. Cost: on a total-store-order architecture such as x86, plain loads and stores already have acquire and release semantics, so the tier is nearly free (only the compiler restriction remains). On ARM and similar, dedicated load-acquire/store-release instructions exist and are markedly cheaper than a full barrier. ## Sequential consistency: one order everyone agrees on SC adds a global property: all SC operations across all threads appear in one single total order, consistent with each thread's program order. Every thread agrees on that order — there is no 'thread A saw X first while thread B saw Y first'. This is the model people intuitively assume when they reason by interleaving, and it is the only tier that forbids store buffering and the related symmetric puzzles. The price is the StoreLoad fence, meaning the writing core must let its stores become globally visible before proceeding — typically tens of cycles, unhidable. That is usually irrelevant on a cold path and very relevant in a hot loop. ## How to choose 1. **Default to the strongest.** Use locks, or SC atomics. Correct-by-default is worth far more than the fence. 2. **Drop to acquire/release when the code is handoff-shaped:** exactly one publisher of a value, and readers that need to see the writer's prior work. Producer/consumer slots, initialization flags, lock implementations. 3. **Drop to relaxed only with a written argument** that no other data's visibility depends on this operation — and prefer to confine it inside a small, reviewed data structure. 4. **Never mix tiers by accident.** A release store paired with a *relaxed* load synchronizes nothing; the pairing is what creates the guarantee. ## What interviewers are checking That you can state the pairing rule precisely, that you know acquire/release is one-way and therefore strictly weaker than SC, that you can name the cost of each tier, and that your default is the strong one with weakening justified by measurement rather than instinct.
- Two threads each perform a release store to their own flag and then an acquire load of the other's flag. Can both loads return the old value?Yes. Acquire and release are one-way: neither forbids a store from being delayed past a later load in the same thread. That is the store-buffering pattern, and preventing it requires a full StoreLoad barrier, which only sequentially consistent operations provide. This is the sharpest practical difference between the two tiers.
- What happens if a release store is paired with a relaxed load rather than an acquire load?The synchronization does not happen. The reader may observe the published value while still seeing stale versions of the data written before the release, because nothing stops the reader's later loads from having been executed earlier. Both halves of the pair are required; the guarantee is a property of the pair, not of the writer alone.
Release/acquire is a signed handoff of a package: whatever the sender packed before signing is guaranteed present when the receiver countersigns. Sequential consistency is additionally a single company-wide ledger in which every handoff appears in one agreed order.
saying these in an interview costs you the question
- Treating acquire/release as equivalent to sequential consistency
- Applying release to the writer only and leaving the reader unordered
- Using relaxed ordering for a flag that gates access to other data
- Believing the strongest ordering is always ruinously slow — on strong hardware most tiers cost little except SC stores
- Choosing relaxed ordering for performance without any measurement or containment