skip to content

An engineer proposes replacing locks in a hot code path with hand-written weakly ordered atomic operations and explicit memory fences for performance. As the technical decision-maker, how do you evaluate that proposal and what conditions would you attach?

level: principalimportance: nice to knowfreq 26%

answer

  1. profile first — contention vs ordering cost
  2. share less / hold shorter / batch before weakening
  3. contain in one component with a safe API
  4. verify: detector + litmus + both architecture classes
  5. write the ordering rationale; two maintainers minimum

basics

~20 s

Demand evidence the synchronization is actually the bottleneck, then treat weak ordering as a contained asset: encapsulated behind a small reviewed API, justified in writing per operation, verified with race detectors and tests on both strongly and weakly ordered hardware, and owned by someone who can maintain it.

solid answer

~50 s

I evaluate it on four axes. **Evidence.** A profile showing lock contention or fence cost dominating, plus a benchmark of the proposed alternative on the *target* hardware. Uncontended locks are already cheap; if contention is the problem, reducing sharing (sharding, per-thread state, batching) usually beats weakening ordering and carries none of the risk. **Containment.** Weak ordering may live only inside a small data structure with a normal API, never sprinkled through business logic. The blast radius must be one reviewable file. **Verification.** Race-detector runs, targeted litmus/stress tests, and testing on both a total-store-order and a weakly ordered architecture, because a strong machine cannot demonstrate correctness. Formal or model-checked reasoning if the structure is genuinely novel. **Maintainability.** Every non-sequentially-consistent operation carries a written argument for why that tier suffices. At least two people can maintain it, and it survives a rewrite by someone who was not in the room. If any axis fails, the answer is no — usually with a cheaper structural alternative attached.

go deeper

for a junior

Recognize that hand-written weak ordering is expert territory and that the normal answer is to use locks or library primitives.

for a middle

Ask for measurement first and point out cheaper alternatives — reduce sharing, shorten critical sections, batch — before weakening ordering.

for a senior

Add concrete verification requirements: race detectors in CI, litmus tests, and testing on both strongly and weakly ordered architectures because a strong machine hides the failures.

for a principal

Decide explicitly: state the approval conditions, the containment boundary, the ownership and documentation requirements, the Amdahl ceiling on the gain, and the cost of reverting.

## Framing: this is a risk purchase, not a code style choice Hand-rolled weak ordering buys throughput with a currency the team pays later: bugs that are undefined, intermittent, architecture-dependent, invisible to ordinary testing, and expensive to diagnose in production. The decision is whether the measured gain is worth that, and whether the risk can be contained. ## Axis 1 — Evidence that synchronization is the bottleneck Ask for a profile, not a belief. Uncontended lock acquisition is a cheap atomic operation; the costs people attribute to 'locks' are usually contention (threads serializing on shared state) or context switching. If contention is the cause, note that weakening ordering does not reduce contention at all — the same cache line is still bounced between cores. The cheaper interventions, in order: 1. **Share less.** Shard the structure, give each worker its own accumulator, merge at the end. 2. **Hold shorter.** Move I/O, allocation and formatting outside the critical section. 3. **Batch.** Amortize one synchronization over many items. 4. **Change the structure.** A queue handoff or single-owner design removes the sharing entirely. Only if those are exhausted, and a prototype shows a meaningful end-to-end win (not a microbenchmark delta), does the proposal stay alive. Amdahl's law is worth quoting here: if the synchronized section is 5% of the workload, perfecting it caps the gain at about 5%. ## Axis 2 — Containment Weak ordering must be an implementation detail. The rule I attach: it lives inside one small data structure or primitive with an ordinary, hard-to-misuse API, and no ordering obligation leaks to callers. If using the component correctly requires the caller to remember 'read this field only after that flag', the encapsulation has failed and the design should be rejected. Blast radius is what makes the difference between a bad week and a bad quarter. ## Axis 3 — Verification Testing cannot show the absence of a race, so verification must be structural: - **Race detectors** in CI, running the component's stress tests. - **Litmus-style tests** for the specific orderings the algorithm depends on. - **Both architecture classes.** A total-store-order machine forbids three of the four reorderings, so passing there proves nothing about weakly ordered targets. If the fleet is mixed — or might become mixed — test on both. - **Model checking or a published, peer-reviewed algorithm** for anything genuinely novel. 'We invented a lock-free structure' should raise the bar sharply; 'we implemented a well-known published algorithm faithfully' lowers it. ## Axis 4 — Maintainability and ownership Weakly ordered code is not self-documenting. Conditions I attach: each non-sequentially-consistent operation carries a comment stating the invariant it relies on and why the weaker tier is sufficient; the pairing partner of each release/acquire is named; and at least two engineers can explain the algorithm. Otherwise the first well-intentioned refactor — adding a field, reordering two lines, extracting a helper — silently deletes an ordering guarantee that no test will catch. Also consider organizational fit: a team that reviews this code twice a year will lose the knowledge. That argues for using a well-tested library primitive instead of owning the algorithm. ## Deciding Approve when: a profile shows real, dominating cost; the cheaper structural options are exhausted; the change is confined to one component with a safe API; a race detector and dual-architecture tests run in CI; and the ordering rationale is written down with named owners. Decline when the motivation is 'locks are slow' as a general belief, when the weak ordering would spread across call sites, when the fleet's target hardware cannot be tested, or when one person is the only reader of the algorithm. A good decision-maker also frames the reversal path: if this becomes a maintenance problem, what does going back to a lock cost? If the answer is 'a one-line change inside one file', the risk is bounded and the experiment is reasonable. If the answer is 'a redesign', it is not. ## What interviewers are listening for Not an opinion about locks. They want measurement before optimization, containment of dangerous techniques, an honest statement that testing cannot prove concurrency correctness, an awareness of architecture portability, and attention to who maintains it after the author leaves.

  • The benchmark shows a 30% improvement in a microbenchmark of the data structure. Is that enough to approve?
    No. A microbenchmark measures the component in isolation, usually with unrealistic contention and cache behaviour. What matters is end-to-end effect on the service's latency or throughput, and by Amdahl's law that is bounded by the fraction of total time the component consumes. Ask for the same measurement in a realistic workload before deciding.
  • The whole fleet is x86 today. Does that remove the portability concern?
    It reduces it but does not remove it. Fleets migrate — ARM instances, new laptops for developers, a future accelerator or edge target — and a codebase whose correctness silently depends on total store order will fail all at once when that happens, in code nobody remembers writing. At minimum, run the component's tests on a weakly ordered machine in CI so the dependency is visible rather than latent.

saying these in an interview costs you the question

  • Approving on the belief that locks are inherently slow, without a profile
  • Letting weak ordering spread across business logic instead of one component
  • Treating a passing test suite on strongly ordered hardware as proof of correctness
  • Optimizing a section too small to matter, ignoring Amdahl's law
  • Accepting a bespoke lock-free algorithm with a single author and no written ordering rationale

context