skip to content

Two servers in different data centers each timestamp an event using their local system clock (wall-clock/time-of-day). Why can't you safely assume that whichever timestamp is numerically smaller happened first in real time?

level: juniorimportance: must knowfreq 70%

answer

  1. crystal oscillator drift in ppm
  2. clock skew (accumulated) vs clock drift (rate)
  3. NTP bounds error, doesn't erase it
  4. step vs slew corrections
  5. last-write-wins skew footgun

basics

~20 s

Computer clocks aren't perfectly synced - they drift apart between corrections and can even jump backward during a correction, so timestamps from two different machines can't be trusted to show which event really happened first.

solid answer

~40 s

Physical clocks on separate machines run on independent, imperfect crystal oscillators that drift due to temperature, aging, and manufacturing tolerance, so even after periodic synchronization they diverge again between sync intervals - accumulated divergence between two clocks is called clock skew. Synchronization protocols like NTP only bound this skew (often low single-digit milliseconds on a LAN, tens of milliseconds or worse over a WAN or under congestion); they don't eliminate it. Clocks can also be stepped backward during a correction, so a later real-world event can end up with an earlier-looking timestamp than one that already happened. Comparing raw timestamps across machines therefore can't reliably establish ordering unless you know the error bound and account for it, or you use a mechanism that doesn't depend on physical clock agreement at all.

go deeper

for a junior

Should recall that computer clocks aren't perfectly in sync and that comparing timestamps from two different machines isn't safe for exact ordering - doesn't need the ppm math or NTP internals.

for a middle

Should articulate clock skew vs drift, know that NTP bounds error but doesn't eliminate it, and name at least one concrete failure mode such as last-write-wins data loss.

for a senior

Should connect this to system design: when raw wall-clock comparison is acceptable (human-facing logs) versus dangerous (conflict resolution, distributed transactions), and name the standard mitigations (logical clocks, bounded-uncertainty clocks).

for a principal

Should reason quantitatively about typical error budgets, explain why hardware-based bounded-uncertainty clocks exist as a solution category, and weigh that trade-off (cost of GPS/atomic hardware vs cost of skew-induced bugs) in a system they're designing.

## Why a hardware clock is only approximate Every machine has a local hardware clock, typically a **quartz crystal oscillator**, that ticks at an approximately fixed frequency and drives a counter the OS turns into wall-clock time (time-of-day, e.g. milliseconds since the Unix epoch). "Approximately" is the key word: the oscillator's real frequency depends on manufacturing tolerance, temperature, supply voltage, and age, so no two machines' clocks tick at exactly the same rate. ## Drift and skew - **Clock drift** is that divergence in tick rate, usually measured in parts-per-million (ppm); a cheap commodity crystal can drift 10-100 ppm, which is up to several seconds per day if left uncorrected. - **Clock skew** is the accumulated difference between any two clocks at a given instant. Because every machine drifts independently and unpredictably, two clocks that agreed perfectly at some moment will disagree later. ## What synchronization does, and what it leaves behind To keep skew bounded, machines periodically synchronize against a reference, most commonly via **NTP**, which nudges the local clock back toward the reference. But synchronization is neither instantaneous nor continuous: - It happens periodically, seconds to minutes apart. - The correction itself carries its own uncertainty from network delay. - Between syncs, drift reaccumulates. Crucially, a correction can move the clock backward if it has drifted ahead of the reference - this is a **step** - or the synchronization daemon can **slew** the clock, temporarily running it faster or slower so it converges smoothly without ever going backward. Slewing is preferred but not universal: large corrections, cold boots, VM live-migration/resume-from-suspend, or manual admin changes can still produce a hard backward step. ## What that means for comparing two timestamps The practical consequence is that a timestamp read on machine A and one read on machine B are each best understood as "true time plus or minus some unknown, time-varying error," not an exact value. If A stamps event X at t=100ms and B stamps event Y at t=99ms, you cannot conclude Y happened first - the real ordering could easily be reversed, because each timestamp's error can be several milliseconds on a LAN and tens of milliseconds across a WAN, worse under congestion, virtualization jitter, or a misconfigured NTP daemon. Even on a single machine, wall-clock time is unsafe for measuring elapsed durations for the same reason a step correction can occur mid-measurement, which is why **monotonic clocks** exist as a separate, purpose-built tool. ## How it shows up in production This shows up in production in specific, painful ways. 1. A distributed store that resolves write conflicts by **"last write wins"** using client-supplied wall-clock timestamps can silently keep the wrong write and lose data if the writer that looks "later" actually had a fast clock - this is a well-documented footgun in systems that offer timestamp-based conflict resolution, and operators are advised to run tight clock discipline and still treat near-simultaneous writes as genuinely ambiguous rather than deterministically resolvable. 2. **Distributed tracing** that stitches spans together purely by wall-clock timestamp can render a child span as starting before its parent, confusing whoever is debugging a latency spike. 3. **Fleet-wide log correlation** by timestamp alone becomes unreliable for reconstructing exact event order without additional correlation IDs or sequence numbers. ## The engineering answer The general engineering answer is: never use raw wall-clock comparison to establish causal or global ordering across machines unless you have an explicit, quantified error bound and actively design around it. Two mechanisms do that: - **Bounded-uncertainty clocks.** This is exactly the idea behind bounded-uncertainty systems like Google's TrueTime, which report an interval guaranteed to contain the true time rather than a single falsely-precise number, and then have downstream logic (like Spanner's commit-wait) actively wait out that uncertainty before trusting an ordering decision. - **Logical or causal ordering.** Where that hardware investment isn't available or needed, the standard alternative is a purely logical or causal ordering mechanism - Lamport clocks, vector clocks, or hybrid logical clocks - that establishes happened-before relationships from message passing rather than physical time agreement. In practice, a well-run fleet with disciplined NTP keeps skew to low single-digit milliseconds, which is perfectly fine for human-facing logs and dashboards, but is not on its own a sound foundation for correctness-critical distributed ordering decisions.

  • If NTP-disciplined skew across a fleet is typically only a few milliseconds, why isn't that 'good enough' for ordering financial transactions?
    A few milliseconds of skew is the same order of magnitude as the gap between rapid successive events in a busy system - two transactions submitted 1ms apart on different nodes can easily have their timestamps swapped by clock error, silently reordering them. For anything where ordering has correctness or financial consequences, you need either a proven, actively-waited error bound like TrueTime's commit-wait, or a causality mechanism that doesn't depend on clock agreement at all.
  • What's the actual difference between clock skew and clock drift?
    Drift is the rate at which one clock's tick speed diverges from true time, for example 50 parts-per-million too fast - it's a property of a single clock over time. Skew is the accumulated difference in reported time between two clocks at a given instant, essentially the integral of their relative drift since they last agreed. Drift causes skew, but skew can also jump from one-time events like a step correction or a VM being paused during migration.
  • Would using nanosecond-resolution timestamps instead of milliseconds fix the cross-machine ordering problem?
    No, resolution and accuracy are different things. A nanosecond-resolution wall-clock read is still only as accurate as the underlying synchronized clock, which might already be off by milliseconds - you've just added more precise digits to a number that was already wrong. Higher resolution only helps once the accuracy/uncertainty-bound problem is separately solved.

Like several wall clocks in different rooms of a house, each running a little fast or slow depending on its own battery and temperature - even if you set every clock to the exact same time this morning, by evening they disagree, and you can't tell from the clock faces alone which room's alarm actually rang first.

saying these in an interview costs you the question

  • Assumes two machines' wall-clock readings are directly comparable with no stated error bound
  • Proposes last-write-wins-by-timestamp as a conflict resolution strategy without acknowledging skew risk
  • Doesn't distinguish drift (a rate) from skew (an accumulated offset)
  • Believes NTP eliminates clock error rather than merely bounding it
  • Suggests higher-resolution timestamps (nanoseconds) fix cross-machine ordering on their own

context