Google's TrueTime API, used inside the Spanner database, doesn't return a single timestamp for 'now' - it returns an interval [earliest, latest]. Explain the mechanism behind this design, including the role GPS receivers and atomic clocks play, and why returning an interval is more useful than a system trying to report one perfectly accurate timestamp.
answer
- TT.now() returns [earliest, latest], not a point
- epsilon grows between timemaster polls
- GPS + atomic clocks = diverse references against correlated failure
- commit-wait waits out epsilon before visibility
- interval gives a provable bound, not just a probable one
basics
~20 sInstead of pretending to know the exact time, TrueTime admits 'the real time is somewhere in this small window,' using GPS satellites and atomic clocks spread across data centers to keep that window tiny. This gives the database a guaranteed, honest bound to reason with instead of a false, precise-looking number.
solid answer
~50 sTrueTime is backed by a fleet of GPS receivers and atomic clocks distributed across Google's data centers, so if one source or even an entire site fails, others still bound the error - this diversity is deliberate against correlated failure. Local 'timemaster' servers combine several references and publish a time plus a growing uncertainty (epsilon) that widens between polls, historically kept to roughly a few milliseconds, due to local oscillator drift. TrueTime.now() returns an interval [earliest, latest] guaranteed to contain true UTC time, rather than a single value with implicit, unstated error. Spanner exploits this honesty directly: a committing transaction picks a timestamp and then performs 'commit-wait,' delaying visibility until TT.now().earliest has passed that timestamp, which guarantees external consistency - if one transaction commits before another starts in real time, the second is guaranteed a later timestamp. The smaller epsilon is, the less waiting is required, which is exactly why the dedicated GPS/atomic hardware exists instead of relying on ordinary NTP's much larger error bound.
go deeper
Should grasp the basic idea that some systems admit clock uncertainty explicitly as a range instead of pretending to know the exact time, without needing commit-wait mechanics.
Should know TrueTime returns an [earliest, latest] interval backed by GPS/atomic hardware and that this is different from ordinary NTP's single-value-with-hidden-error approach.
Should explain why the interval (not just tighter hardware) is what enables a safe wait-and-then-proceed mechanism, and connect epsilon width to real trade-offs like latency.
Should evaluate this as an architectural choice - hardware cost, diverse-reference redundancy design, and the alternative of consensus-based logical ordering - and reason about when this investment is and isn't justified for a system they're designing.
## Reporting the error instead of hiding it The starting insight behind TrueTime is that every physical clock has some error, and pretending otherwise by reporting a single "now" value is a form of dishonesty that pushes the error into whatever system consumes that value, usually silently. TrueTime instead reports the error explicitly: calling `TT.now()` returns a triple, or equivalently an interval `[earliest, latest]`, such that the true, real-world UTC time at the moment of the call is guaranteed to fall somewhere inside that interval. The width of the interval, commonly denoted **epsilon** (so `latest - earliest = 2*epsilon` around a midpoint estimate), is not fixed; it grows the longer it has been since the local machine last synchronized against a trusted reference, because between syncs the local oscillator's drift accumulates just as it does for any other clock. ## What keeps epsilon small What keeps epsilon small in practice is dedicated hardware. Each Google data center hosts a set of "time master" machines equipped with GPS receivers and, in many cases, atomic clocks (rubidium references) as a secondary, independent source. Using two physically different reference technologies is a deliberate defense against correlated failure: | Reference | How it fails | |---|---| | **GPS** | A GPS antenna can be knocked out by a physical fault or signal jamming. | | **Atomic clock** | An atomic clock free-runs on its own internal physics and is unaffected by that particular failure mode, and vice versa an atomic clock can drift out of calibration while GPS remains available. | A local timemaster: - **combines** readings from multiple such references, cross-checking them against each other and discarding any that disagree beyond a sanity threshold; - **serves** its combined estimate to ordinary machines in the data center over a fast, low-latency internal network path, which itself avoids much of the delay-asymmetry problem that limits plain internet NTP. Every machine's local daemon polls its timemasters at short, frequent intervals and grows epsilon linearly with time-since-last-poll, so a poll failure or a slow poll visibly and safely widens the reported uncertainty rather than silently reporting a stale, falsely-confident value. ## The trade-off it makes explicit The trade-off TrueTime makes explicit is: - **Pay:** extra hardware (GPS antennas, atomic clock modules, redundant timemasters per data center), and accept a small, bounded interval width. - **In exchange:** being able to make correctness guarantees that would otherwise require a fundamentally different, non-physical-time mechanism (like a consensus-ordered logical clock funneled through a single sequencer, which doesn't scale the same way globally). Historical published figures put typical epsilon in the low single-digit milliseconds (commonly cited around 1-7ms), which is dramatically tighter than the tens-of-milliseconds-or-worse error typical of public internet NTP, precisely because the reference hardware is local and the network hop to it is short and dedicated rather than routed across the public internet. ## Why an interval beats a best guess The reason an interval beats a single "best guess" timestamp becomes concrete in Spanner's use of it. Spanner needs to order transactions across a globally-distributed database such that the order is **externally consistent**: if transaction T1 completes (returns success to its client) before transaction T2 begins, in real wall-clock time, then T2 must be assigned a later timestamp than T1 and see T1's effects, everywhere, even if T1 and T2 touch entirely different servers with no other communication between them. A naive approach - just assign transactions timestamps from a fast local clock - can't guarantee this ordering, because two clocks that appear to agree might actually be off by more than the real time gap between the transactions (exactly the wall-clock-unreliability problem). TrueTime resolves this by giving Spanner a mechanism to know when it's safe to conclude a chosen timestamp is definitely in the past everywhere: 1. After picking timestamp `s` for a commit, Spanner waits until `TT.now().earliest` exceeds `s` (this wait is called **commit-wait**) before releasing the transaction's results. 2. Because `earliest` is a guaranteed lower bound on true time, once it exceeds `s`, true UTC time is guaranteed to be past `s` everywhere, satisfying the ordering guarantee for any subsequent transaction. Without an interval with an honest, guaranteed lower bound, there would be no safe moment to know "the true time has now definitely passed timestamp s" - a single best-guess value with hidden error offers no such safety margin to wait out. ## Provably bounded, not roughly accurate A concrete way to see why this matters: if TrueTime instead returned a single number that was, say, 90% likely to be accurate to within 2ms, Spanner would either have to 1. accept a small but nonzero probability of an external-consistency violation (unacceptable for a database marketed as strongly consistent), or 2. pad its wait time with an arbitrary, unjustified safety margin with no principled basis for how large it needs to be. The interval turns "roughly accurate" into "provably bounded," which is exactly the property a correctness-critical distributed system needs, at the cost of the specialized infrastructure required to keep that bound small enough to be commercially useful (a wide epsilon, if unaddressed, would make commit-wait latency unacceptably high for a real transactional database).
- Why does TrueTime use both GPS and atomic clocks instead of just one or the other?The two technologies fail in different, largely uncorrelated ways: GPS depends on visible satellites and can be disrupted by antenna damage, jamming, or signal blockage, while an atomic clock free-runs on internal physics and is immune to those specific issues but can itself drift or fail independently. Combining diverse reference types means a single-cause outage is far less likely to take out a data center's entire time reference at once.
- What would happen to Spanner's write latency if epsilon suddenly doubled across a data center?Commit-wait duration is directly tied to epsilon, so doubling it roughly doubles the minimum wait added to every committing read-write transaction's latency in that region, since Spanner must wait until TT.now().earliest passes the chosen commit timestamp. This wouldn't cause incorrect results, since the guarantee still holds, but it would visibly degrade write throughput and latency until the uncertainty shrinks back down.
- Is TrueTime's interval width the same everywhere in the world, or does it vary by location?It can vary, since epsilon depends on the quality and freshness of the local timemaster's synchronization, which in turn depends on how many reference sources are healthy and reachable at that particular site and how recently a machine last polled them. A data center with degraded hardware or worse connectivity to its timemasters will generally see a wider epsilon than a well-provisioned one.
It's like a weather forecaster who, instead of saying 'it will rain at exactly 3:00pm' (falsely precise and often wrong), says 'rain will arrive sometime between 2:55 and 3:05pm, guaranteed' - less flashy, but you can actually plan around a guarantee, whereas a fake-precise single number just hides the uncertainty until it bites you.
saying these in an interview costs you the question
- Describes TrueTime as returning a single precise timestamp rather than an interval
- Doesn't connect the interval to why commit-wait is possible/necessary
- Thinks GPS alone (without atomic clocks or redundancy) is what TrueTime relies on
- Can't explain why a guaranteed bound is different from 'probably accurate'
- Confuses TrueTime's physical-clock uncertainty bounding with logical/vector clock causal ordering