skip to content

A bounded-uncertainty clock system like Spanner's TrueTime relies on a data center having working GPS receivers and atomic clocks feeding its local time-reference servers. Suppose a data center's GPS antennas fail and its atomic clock references start drifting undetected for an extended period. Walk through what happens to (a) the width of the reported clock-uncertainty interval, (b) transaction commit latency, and (c) what the system should be designed to do if the uncertainty can no longer be bounded with confidence.

level: principalimportance: nice to knowfreq 15%

answer

  1. epsilon grows unbounded without a trusted reference
  2. commit-wait latency scales with epsilon
  3. locks held during wait -> throughput hit too
  4. fail-safe: refuse commits past a threshold, don't silently degrade correctness
  5. redundant GPS + atomic + cross-checking prevents correlated failure

basics

~20 s

If the trusted time sources fail, the system's honesty mechanism widens its uncertainty window the longer it goes without a trustworthy reference, which slows down every write since transactions must wait longer for real time to catch up. If the uncertainty grows too large to trust at all, the system should refuse to make risky guarantees rather than silently return a wrong answer.

solid answer

~50 s

As soon as external references are lost, the local uncertainty interval (epsilon) grows continuously with elapsed time since the last trusted sync, driven by the free-running local oscillator's own uncorrected drift, since there's no longer anything actively bounding that drift. Because a system like Spanner's commit-wait duration is directly proportional to epsilon, growing uncertainty translates linearly into growing write latency and, since locks are held throughout the wait, reduced write throughput under contention - a slow, visible degradation rather than a silent one. If epsilon grows past a configured safety threshold, indicating the system can no longer make its correctness guarantee with confidence, the correct design is to fail safe: refuse new commits (or otherwise degrade availability) in the affected zone rather than continue serving requests on an unbounded, no-longer-trustworthy interval, which would silently risk the external-consistency guarantee the whole design exists to provide. This is a deliberate, explicit CAP-style trade-off, choosing consistency and honesty about correctness over availability during the degraded period.

go deeper

for a junior

Should get the basic shape: if the trusted clock source breaks, the system gets less sure of the time and things slow down as a result, without needing the commit-wait mechanics.

for a middle

Should connect epsilon growth to the loss of a trusted reference and know that latency (not incorrect results) is the first visible symptom of degraded clock uncertainty.

for a senior

Should reason through the full chain - reference loss, epsilon growth, commit-wait latency, throughput impact - and know why the design fails loud rather than silently continuing.

for a principal

Should design the operational response: thresholds for refusing commits, monitoring/alerting on epsilon, redundant/diverse reference hardware as prevention, and frame this explicitly as a deliberate consistency-over-availability trade-off comparable to CAP-theorem reasoning.

## Built to fail loud The core design principle at work here is that the system's clock-uncertainty mechanism is built to fail loud, not silent. When GPS antennas at a site fail and atomic clock references start drifting without being cross-checked (because the redundant reference that would normally catch the drift, GPS, is also unavailable), the local timemaster no longer has a trustworthy external anchor. Every machine's local clock-discipline daemon continues to poll the timemaster, but since the timemaster itself has stopped receiving fresh, cross-validated reference data, it can no longer shrink the reported uncertainty back down after each poll the way it normally would. The uncertainty interval instead grows continuously, driven purely by the free-running local oscillator's own drift rate accumulating unchecked, exactly the same physical mechanism responsible for ordinary clock skew between unsynchronized machines, just now playing out inside what's supposed to be a tightly-bounded reference clock rather than an ordinary commodity clock. ## What the widening interval is admitting This growth is not hypothetical or edge-case-only; it is the explicit, designed-for behavior of the uncertainty-tracking mechanism, and it's a feature, not a bug: a widening interval is the system's honest admission that it's less sure of the time than it was a moment ago, which is precisely the alternative to the dishonest behavior of continuing to report a falsely tight interval after its basis for confidence has eroded. Contrast: | Clock | Does it report its own uncertainty? | |---|---| | **An ordinary wall clock or NTP daemon** | Typically has no concept of reporting its own uncertainty at all. | | **A bounded-uncertainty design** | The design's entire value proposition is that it converts silent, unknowable clock error into an explicit, growing, and therefore actionable number. | ## The cost in commit latency The direct consequence for a system like Spanner is on commit-wait latency. Since commit-wait requires waiting until the lower bound of the uncertainty interval passes the transaction's assigned commit timestamp, and that lower bound is now further away from the upper-bound-derived timestamp because the interval itself is wider, every read-write transaction committing in the affected data center pays a longer minimum latency tax, growing roughly in step with epsilon. - Because locks are held for the duration of commit-wait, this isn't merely a latency inconvenience for individual requests; it directly reduces achievable write throughput under contention, since each transaction now occupies its locks for longer before releasing them. - Operationally, this shows up as a gradual, monitorable latency and throughput degradation specifically localized to the affected data center or zone, rather than a sudden outage, giving operators a real signal (rising commit latencies, rising reported epsilon in monitoring) well before the situation becomes critical. ## When epsilon grows past the usable threshold The interesting design question is what happens if the degradation continues long enough that epsilon grows past whatever threshold the system considers still "usable." 1. **A naive design** might simply let commit-wait grow arbitrarily long, technically preserving correctness (the guarantee still holds mathematically no matter how wide epsilon gets, since the wait just gets longer) but at some point the latency becomes operationally unacceptable, effectively an availability failure in practice even if not in name. 2. **A more robust design** treats a sufficiently degraded uncertainty state as a first-class failure condition: past a configured threshold, the affected servers should stop accepting new write commits (or, in a more graduated approach, shed load and signal upstream that the zone is degraded) rather than silently let latency balloon or, in a worse hypothetical design, abandon the guarantee and start committing at unbounded risk. This is a deliberate application of the CAP-style trade-off, explicitly favoring consistency and correctness over availability during a degraded-clock episode, on the reasoning that serving a request with a silently-broken external-consistency guarantee is worse than serving no request at all, especially for a database whose entire value proposition to its users rests on that guarantee holding unconditionally. ## The mitigations that make it rare At a fleet-operations level, the mitigations that make this scenario rare in the first place are worth naming: - **physically redundant GPS antennas and receivers per site**, so a single antenna failure doesn't take out the whole reference; - **a second, independent reference technology (atomic clocks)**, specifically so a GPS-class failure mode doesn't correlate with an atomic-clock-class one; - **cross-checking multiple references against each other**, so a single degraded source is caught and excluded rather than silently trusted; - **monitoring/alerting on reported epsilon itself** as an operational signal, treating a rising uncertainty trend as an incident-worthy event well before it crosses a hard failure threshold. The broader lesson generalizes past this specific system: any correctness guarantee built on a physical resource with a real, nonzero failure rate needs an explicit, monitored degradation path, not just a happy-path design, and the honest, self-reporting nature of a bounded-uncertainty clock is what makes that degradation path possible to build at all.

  • Why is a linearly growing uncertainty interval preferable to the timemaster simply reporting its last-known-good epsilon forever?
    Reporting a stale, falsely tight epsilon would mean downstream commit-wait logic stops waiting long enough, silently breaking the external-consistency guarantee without any signal that anything is wrong. A growing interval keeps the guarantee mathematically sound throughout the outage and produces a visible, monitorable symptom (rising latency) that operators can act on before real correctness is at risk.
  • Could you avoid this whole failure mode by just running more NTP servers across the public internet as a backup, instead of dedicated GPS/atomic hardware?
    Public NTP servers would provide some fallback but at a much wider baseline uncertainty than dedicated local hardware, since they're subject to internet path asymmetry and congestion; falling back to them would still cause a meaningful latency jump compared to normal operation, just a smaller and more bounded one than free-running with no reference at all. It's a reasonable degraded-mode fallback but not a substitute for the tighter guarantee dedicated hardware provides during normal operation.
  • Does this failure scenario affect Spanner's read-only transactions the same way it affects read-write commit-wait?
    Not identically - read-only transactions read at a chosen past timestamp and don't need to wait for real time to catch up the way commit-wait does, so they're less directly latency-impacted by a growing epsilon. However, extremely degraded uncertainty could still affect how far in the past a safe read timestamp needs to be chosen, so read paths aren't entirely immune, just less severely coupled to epsilon than the write path is.

It's like a ship's navigator who, losing satellite positioning, keeps dead-reckoning from the last known fix - each hour without a fix, the navigator's stated 'we could be anywhere in this radius' circle grows wider and wider, which is honest and useful, but if that circle ever grows large enough that no safe course can be plotted with confidence, the right move is to stop and wait for a fresh fix rather than keep sailing on a guess that's quietly become worthless.

saying these in an interview costs you the question

  • Assumes uncertainty stays constant if the reference clock fails, rather than growing over time
  • Doesn't connect growing epsilon to growing commit-wait latency and reduced throughput
  • Proposes silently ignoring the degraded uncertainty and continuing to commit as normal
  • Can't articulate the fail-safe design principle (refuse commits over silently risking correctness)
  • Doesn't mention redundant/diverse reference sources as the preventive mitigation

context