skip to content

questions

6

Two servers in different data centers each timestamp an event using their local system clock (wall-clock/time-of-day). Why can't you safely assume that whichever timestamp is numerically smaller happened first in real time?

level: juniorimportance: must knowfreq 70%

answer

  1. crystal oscillator drift in ppm
  2. clock skew (accumulated) vs clock drift (rate)
  3. NTP bounds error, doesn't erase it
  4. step vs slew corrections
  5. last-write-wins skew footgun

basics

~20 s

Computer clocks aren't perfectly synced - they drift apart between corrections and can even jump backward during a correction, so timestamps from two different machines can't be trusted to show which event really happened first.

solid answer

~40 s

Physical clocks on separate machines run on independent, imperfect crystal oscillators that drift due to temperature, aging, and manufacturing tolerance, so even after periodic synchronization they diverge again between sync intervals - accumulated divergence between two clocks is called clock skew. Synchronization protocols like NTP only bound this skew (often low single-digit milliseconds on a LAN, tens of milliseconds or worse over a WAN or under congestion); they don't eliminate it. Clocks can also be stepped backward during a correction, so a later real-world event can end up with an earlier-looking timestamp than one that already happened. Comparing raw timestamps across machines therefore can't reliably establish ordering unless you know the error bound and account for it, or you use a mechanism that doesn't depend on physical clock agreement at all.

go deeper

for a junior

Should recall that computer clocks aren't perfectly in sync and that comparing timestamps from two different machines isn't safe for exact ordering - doesn't need the ppm math or NTP internals.

for a middle

Should articulate clock skew vs drift, know that NTP bounds error but doesn't eliminate it, and name at least one concrete failure mode such as last-write-wins data loss.

for a senior

Should connect this to system design: when raw wall-clock comparison is acceptable (human-facing logs) versus dangerous (conflict resolution, distributed transactions), and name the standard mitigations (logical clocks, bounded-uncertainty clocks).

for a principal

Should reason quantitatively about typical error budgets, explain why hardware-based bounded-uncertainty clocks exist as a solution category, and weigh that trade-off (cost of GPS/atomic hardware vs cost of skew-induced bugs) in a system they're designing.

## Why a hardware clock is only approximate Every machine has a local hardware clock, typically a **quartz crystal oscillator**, that ticks at an approximately fixed frequency and drives a counter the OS turns into wall-clock time (time-of-day, e.g. milliseconds since the Unix epoch). "Approximately" is the key word: the oscillator's real frequency depends on manufacturing tolerance, temperature, supply voltage, and age, so no two machines' clocks tick at exactly the same rate. ## Drift and skew - **Clock drift** is that divergence in tick rate, usually measured in parts-per-million (ppm); a cheap commodity crystal can drift 10-100 ppm, which is up to several seconds per day if left uncorrected. - **Clock skew** is the accumulated difference between any two clocks at a given instant. Because every machine drifts independently and unpredictably, two clocks that agreed perfectly at some moment will disagree later. ## What synchronization does, and what it leaves behind To keep skew bounded, machines periodically synchronize against a reference, most commonly via **NTP**, which nudges the local clock back toward the reference. But synchronization is neither instantaneous nor continuous: - It happens periodically, seconds to minutes apart. - The correction itself carries its own uncertainty from network delay. - Between syncs, drift reaccumulates. Crucially, a correction can move the clock backward if it has drifted ahead of the reference - this is a **step** - or the synchronization daemon can **slew** the clock, temporarily running it faster or slower so it converges smoothly without ever going backward. Slewing is preferred but not universal: large corrections, cold boots, VM live-migration/resume-from-suspend, or manual admin changes can still produce a hard backward step. ## What that means for comparing two timestamps The practical consequence is that a timestamp read on machine A and one read on machine B are each best understood as "true time plus or minus some unknown, time-varying error," not an exact value. If A stamps event X at t=100ms and B stamps event Y at t=99ms, you cannot conclude Y happened first - the real ordering could easily be reversed, because each timestamp's error can be several milliseconds on a LAN and tens of milliseconds across a WAN, worse under congestion, virtualization jitter, or a misconfigured NTP daemon. Even on a single machine, wall-clock time is unsafe for measuring elapsed durations for the same reason a step correction can occur mid-measurement, which is why **monotonic clocks** exist as a separate, purpose-built tool. ## How it shows up in production This shows up in production in specific, painful ways. 1. A distributed store that resolves write conflicts by **"last write wins"** using client-supplied wall-clock timestamps can silently keep the wrong write and lose data if the writer that looks "later" actually had a fast clock - this is a well-documented footgun in systems that offer timestamp-based conflict resolution, and operators are advised to run tight clock discipline and still treat near-simultaneous writes as genuinely ambiguous rather than deterministically resolvable. 2. **Distributed tracing** that stitches spans together purely by wall-clock timestamp can render a child span as starting before its parent, confusing whoever is debugging a latency spike. 3. **Fleet-wide log correlation** by timestamp alone becomes unreliable for reconstructing exact event order without additional correlation IDs or sequence numbers. ## The engineering answer The general engineering answer is: never use raw wall-clock comparison to establish causal or global ordering across machines unless you have an explicit, quantified error bound and actively design around it. Two mechanisms do that: - **Bounded-uncertainty clocks.** This is exactly the idea behind bounded-uncertainty systems like Google's TrueTime, which report an interval guaranteed to contain the true time rather than a single falsely-precise number, and then have downstream logic (like Spanner's commit-wait) actively wait out that uncertainty before trusting an ordering decision. - **Logical or causal ordering.** Where that hardware investment isn't available or needed, the standard alternative is a purely logical or causal ordering mechanism - Lamport clocks, vector clocks, or hybrid logical clocks - that establishes happened-before relationships from message passing rather than physical time agreement. In practice, a well-run fleet with disciplined NTP keeps skew to low single-digit milliseconds, which is perfectly fine for human-facing logs and dashboards, but is not on its own a sound foundation for correctness-critical distributed ordering decisions.

  • If NTP-disciplined skew across a fleet is typically only a few milliseconds, why isn't that 'good enough' for ordering financial transactions?
    A few milliseconds of skew is the same order of magnitude as the gap between rapid successive events in a busy system - two transactions submitted 1ms apart on different nodes can easily have their timestamps swapped by clock error, silently reordering them. For anything where ordering has correctness or financial consequences, you need either a proven, actively-waited error bound like TrueTime's commit-wait, or a causality mechanism that doesn't depend on clock agreement at all.
  • What's the actual difference between clock skew and clock drift?
    Drift is the rate at which one clock's tick speed diverges from true time, for example 50 parts-per-million too fast - it's a property of a single clock over time. Skew is the accumulated difference in reported time between two clocks at a given instant, essentially the integral of their relative drift since they last agreed. Drift causes skew, but skew can also jump from one-time events like a step correction or a VM being paused during migration.
  • Would using nanosecond-resolution timestamps instead of milliseconds fix the cross-machine ordering problem?
    No, resolution and accuracy are different things. A nanosecond-resolution wall-clock read is still only as accurate as the underlying synchronized clock, which might already be off by milliseconds - you've just added more precise digits to a number that was already wrong. Higher resolution only helps once the accuracy/uncertainty-bound problem is separately solved.

Like several wall clocks in different rooms of a house, each running a little fast or slow depending on its own battery and temperature - even if you set every clock to the exact same time this morning, by evening they disagree, and you can't tell from the clock faces alone which room's alarm actually rang first.

saying these in an interview costs you the question

  • Assumes two machines' wall-clock readings are directly comparable with no stated error bound
  • Proposes last-write-wins-by-timestamp as a conflict resolution strategy without acknowledging skew risk
  • Doesn't distinguish drift (a rate) from skew (an accumulated offset)
  • Believes NTP eliminates clock error rather than merely bounding it
  • Suggests higher-resolution timestamps (nanoseconds) fix cross-machine ordering on their own

context

open as a page

A service measures how long an operation took by calling a wall-clock/time-of-day API (e.g. the equivalent of `System.currentTimeMillis()`) before and after the operation, then subtracting. Under what circumstances can this produce a negative or wildly wrong duration, and what's the correct fix?

level: middleimportance: must knowfreq 55%

basics

~20 s

If the system's clock gets adjusted backward (say by a time-sync correction) while you're timing something, the 'end' reading can look earlier than the 'start' reading, giving a negative or nonsense duration. Use a monotonic clock instead - one guaranteed to only ever move forward.

open as a page

When a server synchronizes its clock via NTP (Network Time Protocol), walk through how NTP estimates the clock offset and round-trip delay from a time server, and explain why the resulting accuracy is fundamentally limited by network conditions rather than by protocol design.

level: middleimportance: must knowfreq 65%

basics

~20 s

NTP asks a time server what time it is, times how long the round trip took, and assumes the trip there and back took equally long to estimate the network delay and correct the local clock. If the trip isn't actually symmetric, the correction is a little off.

open as a page

Google's TrueTime API, used inside the Spanner database, doesn't return a single timestamp for 'now' - it returns an interval [earliest, latest]. Explain the mechanism behind this design, including the role GPS receivers and atomic clocks play, and why returning an interval is more useful than a system trying to report one perfectly accurate timestamp.

level: seniorimportance: must knowfreq 40%

basics

~20 s

Instead of pretending to know the exact time, TrueTime admits 'the real time is somewhere in this small window,' using GPS satellites and atomic clocks spread across data centers to keep that window tiny. This gives the database a guaranteed, honest bound to reason with instead of a false, precise-looking number.

open as a page

In Google Spanner, a read-write transaction's commit protocol includes a 'commit-wait' step where the coordinator delays making the transaction's writes visible until a certain point. Describe what commit-wait actually waits for and why it's necessary to guarantee external consistency - meaning transactions appear to execute in an order consistent with real, wall-clock time.

level: seniorimportance: should knowfreq 25%

basics

~20 s

Spanner picks a commit timestamp for a transaction, then literally pauses before letting anyone see the result, until it's sure real-world clocks everywhere have caught up past that timestamp. This guarantees that anything starting after the commit will see it, at the cost of adding a little delay to every write.

open as a page

A bounded-uncertainty clock system like Spanner's TrueTime relies on a data center having working GPS receivers and atomic clocks feeding its local time-reference servers. Suppose a data center's GPS antennas fail and its atomic clock references start drifting undetected for an extended period. Walk through what happens to (a) the width of the reported clock-uncertainty interval, (b) transaction commit latency, and (c) what the system should be designed to do if the uncertainty can no longer be bounded with confidence.

level: principalimportance: nice to knowfreq 15%

basics

~20 s

If the trusted time sources fail, the system's honesty mechanism widens its uncertainty window the longer it goes without a trustworthy reference, which slows down every write since transactions must wait longer for real time to catch up. If the uncertainty grows too large to trust at all, the system should refuse to make risky guarantees rather than silently return a wrong answer.

open as a page