In a distributed system with no shared physical clock, why can't you just compare timestamps from different machines to figure out which of two events happened first?
answer
- clocks drift, NTP has ms-level slop
- happens-before: program order + send→receive
- concurrent = causally unrelated, not simultaneous
- LWW conflict resolution can silently lose data
basics
~20 sDifferent computers' internal clocks are never perfectly synchronized, so their timestamps can disagree about event order even when one event actually caused the other. You need a way to order events based on what actually influenced what, not on unreliable clock readings.
solid answer
~40 sPhysical clocks on separate machines drift relative to each other, and even clocks synced via protocols like NTP only agree to within single-digit milliseconds at best, often worse across WANs. If two events happen within that uncertainty window, wall-clock timestamps can't reliably say which came first, and they say nothing about whether one event could have caused the other. Distributed systems instead need an ordering derived from actual information flow: within one process events are totally ordered by program order, and across processes a message send happens-before its receive. Logical clocks -- Lamport timestamps, vector clocks -- turn this happens-before relation into comparable counters without depending on synchronized physical time at all.
go deeper
Should recognize that clock skew exists and can name at least one consequence (misordered events), even without precise formal language.
Should state the happens-before relation's two defining rules and explain that logical clocks are counters, not timestamps, decoupled from wall-clock time.
Should connect this to a concrete production failure mode (e.g. last-write-wins data loss under clock skew) and explain why the fix is causal tracking, not tighter clock sync.
Should be able to reason about when physical-time approaches (bounded-uncertainty clocks) are the right call instead, and articulate the systemic cost of choosing wrong at design time.
## Why the clocks disagree Every machine in a distributed system runs its own local clock, and those clocks are never perfectly aligned. Even with **NTP** synchronization: - agreement is typically bounded to a few **milliseconds** on a **LAN**; - across data centers and **WANs** it can be tens of milliseconds or worse; - clocks also drift continuously between synchronization rounds, because crystal oscillators run at slightly different rates. If two events on two different machines occur within that uncertainty window, comparing their wall-clock timestamps cannot reliably tell you which one actually happened first -- and this problem gets worse, not better, at scale, because more machines means more pairwise skew to reason about. ## Physical time answers the wrong question But the deeper issue isn't just precision, it's that physical time answers the wrong question. What a distributed system usually needs to know isn't 'which event's clock reading is numerically smaller' but 'could event A have influenced event B'. **Leslie Lamport** formalized this as the **happens-before** relation (denoted `→`) in his 1978 paper 'Time, Clocks, and the Ordering of Events in a Distributed System'. Happens-before is defined by two simple rules plus transitivity: - **(1)** if `a` and `b` are events in the same process and `a` occurs before `b` in that process's local execution order, then `a → b`; - **(2)** if `a` is the sending of a message and `b` is the receipt of that same message (possibly in a different process), then `a → b`. Taking the transitive closure of these rules over all events gives a **partial order** -- partial because many pairs of events are simply incomparable: if neither could have influenced the other, they are called **concurrent**, written `a || b`. Concurrency here is a causal statement, not a claim about simultaneity in real time. ## Why the distinction matters This distinction matters enormously in practice. Two writes to a distributed key-value store issued by different clients who never communicated are causally concurrent even if their physical timestamps differ by seconds, and a system that resolves conflicts by 'whoever has the higher timestamp wins' is making an arbitrary, potentially data-losing choice dressed up as a principled one. **Logical clocks** exist precisely to let a system reason about causality -- did this write build on that one, or were they independent -- without needing physical time to be trustworthy at all. **Lamport timestamps** and **vector clocks** are two ways of turning the happens-before relation into comparable values attached to each event, purely through counters that processes increment locally and piggyback on messages. ## The trade-off The trade-off is that logical time deliberately throws away wall-clock meaning. A Lamport timestamp of 47 doesn't tell you anything about when, in real seconds, an event happened, only that it is consistent with a scenario where it occurred after everything with a smaller Lamport value that it's causally related to. That's a feature for correctness (it never lies about causality) but a limitation for anything that actually needs human-meaningful time -- those still need physical clocks, just used carefully. ## The classic production failure mode The classic production failure mode is trusting client- or server-supplied wall-clock timestamps for conflict resolution or ordering when clock skew is large enough to matter. - **The example.** Cassandra's original last-write-wins column semantics is a well-known example: if a client's local clock is skewed backward, its write can silently lose to an older write from a machine with a fast clock, even though the skewed write happened later in real time and causally after the one it lost to -- a source of real, hard-to-diagnose data loss bugs. - **The heavier mechanism.** Systems that need strict global ordering with physical time solve this with hardware-assisted bounded-uncertainty clocks rather than trusting raw NTP, but that's a heavier mechanism living outside this topic. - **The everyday fix.** The everyday, infrastructure-agnostic fix is to track causality explicitly with logical clocks so ordering decisions are based on what a process actually knew about, not on when a clock happened to tick.
- What does it mean for two events to be 'concurrent' under the happens-before relation, and how is that different from them occurring at the same physical instant?Concurrent means neither event's happens-before chain (same-process order or message send/receive) connects to the other -- there's no causal path between them, regardless of their real-world timing. Two events can be concurrent even if one occurred seconds after the other in wall-clock time, as long as nothing propagated information from one to the other. Conversely two events could fire at the exact same instant and still not be 'concurrent' in this sense if a hidden causal chain connects them via a third process.
- Why doesn't NTP synchronization solve this problem well enough for causal ordering?NTP reduces skew but doesn't eliminate it, typically leaving several milliseconds of uncertainty on a LAN and more across WANs, plus continuous drift between sync rounds. That residual uncertainty is often larger than the time between causally related events in a busy system, so you can still misorder events that are microseconds apart even though they're perfectly synced 'well enough' for human purposes.
- If a system doesn't care about causality and only needs approximate event ordering for logging or metrics, do logical clocks still matter?Not really -- for approximate, human-facing ordering like log timestamps or dashboards, physical clocks with reasonable NTP sync are fine because occasional millisecond-level misordering has no correctness impact. Logical clocks earn their complexity specifically when a program needs to make a correctness decision -- conflict resolution, replication consistency, distributed debugging -- based on whether one event could have caused another.
Like reconstructing the order of two friends' text messages sent while their phones both had the wrong time set -- you can't trust the phone's clock, but if one message quotes or replies to the other, you know which came first regardless of what the clocks say.
saying these in an interview costs you the question
- Says wall-clock timestamps from different machines can be compared directly for correctness-critical ordering
- Conflates 'concurrent' with 'simultaneous in real time'
- Thinks NTP fully solves clock synchronization for causal ordering
- Can't state the two happens-before rules (program order, send-before-receive)
- Assumes logical clocks give you real elapsed time