When a server synchronizes its clock via NTP (Network Time Protocol), walk through how NTP estimates the clock offset and round-trip delay from a time server, and explain why the resulting accuracy is fundamentally limited by network conditions rather than by protocol design.
answer
- T1/T2/T3/T4 four-timestamp exchange
- offset formula assumes symmetric delay
- stratum hierarchy (0=reference, 1=attached, ...)
- slew vs step correction
- leap smear mitigates leap-second jumps
basics
~20 sNTP asks a time server what time it is, times how long the round trip took, and assumes the trip there and back took equally long to estimate the network delay and correct the local clock. If the trip isn't actually symmetric, the correction is a little off.
solid answer
~50 sNTP exchanges four timestamps per poll: client-send (T1), server-receive (T2), server-send (T3), client-receive (T4). It computes round-trip delay = (T4-T1)-(T3-T2) and offset = ((T2-T1)+(T3-T4))/2, which assumes network delay is symmetric in both directions. If it isn't, due to asymmetric routing or one-way congestion, the offset estimate is off by roughly half the asymmetry. Small corrections are slewed, gently adjusting clock rate; large ones are stepped. Clients typically poll several servers across a stratum hierarchy, where stratum 1 servers attach directly to a reference like GPS or an atomic clock, and statistically filter out outlying 'falseticker' servers before combining the rest. On a quiet LAN this yields sub-millisecond to low-single-digit-millisecond accuracy; over the public internet, tens of milliseconds is typical and can spike under congestion, which is why systems needing very tight bounds add dedicated hardware time references instead of relying on NTP alone.
go deeper
Should know NTP is what keeps a machine's clock roughly in sync with real time over the network, without needing the timestamp math.
Should be able to state the four-timestamp exchange and the offset/delay formulas, and explain the symmetric-delay assumption as the fundamental accuracy limit.
Should discuss the stratum hierarchy, outlier filtering across multiple servers, and slew vs step behavior, and connect typical LAN vs WAN accuracy numbers to system design decisions.
Should be able to explain why NTP alone is insufficient for systems needing a tight, provable bound (motivating dedicated hardware references like TrueTime) and reason about operational failure modes like leap seconds and firewall-blocked sync at fleet scale.
## The four-timestamp exchange NTP's core mechanism is a four-timestamp exchange per poll. 1. The client records its own send time (`T1`) and stamps it in the request. 2. The server records its receive time (`T2`). 3. After processing, the server stamps its own send time (`T3`) in the reply. 4. The client records its receive time (`T4`) on arrival. From these four values NTP derives two quantities: - **Round-trip delay** — `delta = (T4-T1) - (T3-T2)`, which is the total time spent on the wire minus the time the server spent processing. - **Clock offset** — `theta = ((T2-T1) + (T3-T4)) / 2`, which is the client's best estimate of how far its clock is from the server's. ## The symmetry assumption is the accuracy ceiling The offset formula implicitly assumes the outbound leg (`T1` to `T2`) and inbound leg (`T3` to `T4`) take equal time. When they don't, because of asymmetric routing, one-way congestion, or different queueing on each path, the true offset differs from the computed one by roughly half the asymmetry - this is the fundamental, protocol-independent accuracy ceiling: no amount of clever software fixes an offset estimate built on an unmet symmetry assumption. ## The stratum hierarchy NTP organizes time sources into a stratum hierarchy to limit how much error compounds across hops. | Level | What sits there | |---|---| | **Stratum 0** | The physical reference devices themselves - GPS receivers, atomic clocks, radio time signals - which are not directly on the network. | | **Stratum 1** | Servers that are machines with hardware directly attached to a stratum-0 reference. | | **Stratum 2** | Servers that sync from stratum 1. | | **Stratum 3** | Servers that sync from stratum 2, and so on. | Each hop adds its own network-delay and processing uncertainty, so accuracy generally degrades as stratum number increases, though a well-provisioned stratum 3 pool can still easily beat a poorly-connected stratum 1 server. ## Choosing which servers to trust A client typically polls multiple servers rather than trusting one, then runs a selection algorithm (historically based on Marzullo's algorithm, refined in modern implementations) that: - **discards** servers whose reported time intervals don't overlap with the majority - these are called **falsetickers**, potentially misconfigured or compromised; - **combines** the surviving **truechimers** into a weighted estimate, favoring lower-delay, lower-jitter sources. ## Applying the correction The local clock discipline then applies the correction one of two ways: | Correction | When | What it does | |---|---|---| | **Slew** | The offset is small | It slews, temporarily speeding up or slowing down the reported clock rate so it converges smoothly without ever jumping backward. | | **Step** | The offset exceeds a threshold (classically around 128ms in reference implementations, configurable) | It steps, jumping the clock directly, which can break code that assumed monotonically increasing wall-clock reads. | ## What accuracy is actually achievable The practical accuracy achievable is bounded by network conditions far more than by any refinement to the algorithm. - On a quiet, low-jitter LAN, NTP commonly achieves sub-millisecond to low-single-digit-millisecond accuracy. - Over the public internet, tens of milliseconds of error is typical, and transient congestion, route changes, or asymmetric peering can push it much higher for a poll or two before the filtering algorithm adapts. - Virtualized environments add another layer of jitter, since a guest OS's sense of elapsed time can be distorted by hypervisor scheduling pauses. This is precisely why systems that need a tight, provable error bound - Google's TrueTime being the canonical example - don't rely on public NTP alone; they deploy dedicated GPS receivers and atomic clock references inside their own data centers, feeding a much shorter, better-controlled network path to each machine, and explicitly publish the resulting uncertainty rather than assuming it away. ## Failure modes in operation Operationally, NTP failure modes show up in recognizable ways: - **Silently stopped daemons.** A host whose NTP daemon has silently stopped (crashed, or blocked by an overzealous firewall dropping UDP port 123) will drift unboundedly, sometimes for weeks, until something notices clock-dependent behavior misbehaving. - **Leap seconds** are a related historical pain point - inserting an extra second into UTC has caused real outages (notably widespread Java/Linux issues around the 2012 leap second) because software assumed time only moves forward at a constant rate; Google's public NTP servers popularized "leap smearing," spreading the extra second out as a tiny rate adjustment over roughly a day instead of a single discontinuous step, precisely to avoid triggering step-related bugs.
- Why does NTP prefer to slew the clock instead of stepping it whenever possible?Stepping jumps the clock instantaneously, which can violate assumptions elsewhere in the system, such as code that measures elapsed time with wall-clock reads or expects monotonically increasing timestamps in logs and can produce negative durations or out-of-order log entries. Slewing achieves the same correction gradually, by briefly running the clock a little faster or slower, so time as observed by applications still only ever moves forward.
- If a client only ever queries a single NTP server, what's the main risk compared to querying several?With one server, there's no way to detect that the server itself is faulty, misconfigured, or compromised (a 'falseticker'), so the client has no cross-check and will faithfully adopt a wrong time. Querying multiple servers lets the selection algorithm discard outliers and combine agreeing sources, which is both more accurate and more robust to a single bad reference.
- How does adding stratum hops (e.g., syncing from a stratum 3 server instead of stratum 1) typically affect achievable accuracy?Each additional hop introduces its own network delay measurement and asymmetry uncertainty on top of the hop before it, so accuracy tends to degrade with increasing stratum number, though this is a tendency rather than a strict guarantee. A well-connected, low-jitter stratum 3 pool can still outperform a poorly-connected or congested stratum 1 server.
Like setting your watch by shouting a time question across a canyon and listening for the echo - you estimate the delay by halving the round trip, assuming your voice and the echo take equally long, but if wind is blowing one direction, that assumption is wrong and your watch ends up a bit off no matter how carefully you time the shout.
saying these in an interview costs you the question
- Thinks NTP measures absolute network delay directly instead of estimating it from a round-trip assumption
- Doesn't know the offset calculation assumes symmetric outbound/inbound delay
- Believes stratum number indicates server importance rather than hop distance from a physical reference
- Assumes NTP always corrects gradually and never jumps the clock
- Can't name a concrete NTP failure mode observed in production (silent drift, leap-second step)