In Google Spanner, a read-write transaction's commit protocol includes a 'commit-wait' step where the coordinator delays making the transaction's writes visible until a certain point. Describe what commit-wait actually waits for and why it's necessary to guarantee external consistency - meaning transactions appear to execute in an order consistent with real, wall-clock time.
answer
- commit-wait: wait until TT.now().earliest > commit timestamp s
- guarantees external consistency, not just serializability
- cost proportional to epsilon (uncertainty width)
- read-only transactions skip commit-wait
- locks held throughout the wait -> throughput cost too
basics
~20 sSpanner picks a commit timestamp for a transaction, then literally pauses before letting anyone see the result, until it's sure real-world clocks everywhere have caught up past that timestamp. This guarantees that anything starting after the commit will see it, at the cost of adding a little delay to every write.
solid answer
~50 sSpanner assigns each committing transaction a timestamp s, then before releasing locks and making results visible, waits until TrueTime's TT.now().earliest exceeds s - meaning the local uncertainty interval has moved entirely past s, so real absolute time is guaranteed to now be past s everywhere. This wait is what converts Spanner's timestamp ordering into external consistency: any transaction that starts after another one commits, in real wall-clock time, is guaranteed to receive a later timestamp and see the prior writes, because by the time the first transaction becomes visible, true time has already passed its assigned timestamp everywhere in the system. The cost is latency directly proportional to twice epsilon, the width of the TrueTime uncertainty interval, added to every read-write commit, typically single-digit milliseconds - a deliberate trade of a small, bounded latency tax for a strong, provable correctness guarantee that would otherwise require a much more expensive global coordination mechanism.
go deeper
Should get the basic idea that Spanner pauses briefly before revealing a write's result to make sure the timing is safe, without needing the timestamp-selection details.
Should know commit-wait involves waiting on TrueTime's interval and that it exists to make timestamp ordering trustworthy, even if the exact earliest/latest mechanics aren't precise.
Should walk through the earliest/latest mechanics correctly, explain why this yields external consistency specifically (not just any consistent order), and connect the cost to epsilon and to write latency/throughput.
Should evaluate this as a deliberate architectural trade-off - comparing it to alternative global-ordering mechanisms (e.g. a single sequencer) and reasoning about when the hardware investment plus latency tax is and isn't worth the guarantee for a system they're designing.
## What external consistency means External consistency, the guarantee Spanner is built to provide, means that if a client observes transaction T1 complete (gets a success acknowledgment) before it starts transaction T2, in real wall-clock time, then the system must behave as if T1 happened before T2 - T2 is guaranteed to be assigned a later commit timestamp and to see T1's effects, even if T1 and T2 touch entirely disjoint sets of servers with no direct communication between them. This is a strictly stronger property than ordinary serializability, which only guarantees some consistent ordering exists, not that the ordering matches real time as perceived by clients outside the system. ## The mechanism The mechanism starts when a read-write transaction is ready to commit. 1. Spanner assigns it a timestamp `s`, chosen to be at least as large as `TT.now().latest` at the moment of choosing (i.e., picked from the "safe" upper end of the current uncertainty interval, ensuring it's not smaller than any timestamp already assigned to an earlier-committing, causally-related transaction). 2. Having chosen `s`, Spanner does not yet release the transaction's locks or make its writes visible to other transactions. 3. Instead it enters **commit-wait**: it repeatedly checks `TT.now()` and blocks until `TT.now().earliest`, the guaranteed lower bound of the current time interval, is strictly greater than `s`. Because `earliest` is a guaranteed lower bound on true UTC time (that's the entire point of TrueTime's interval design), once `earliest` exceeds `s`, real time is now provably past `s` everywhere in the system, not just probably past it. ## Why this particular wait delivers the guarantee Why this specific wait is what delivers external consistency is worth walking through concretely. Suppose T1 commits with timestamp s1, and only after T1's client sees the success response does some other client begin T2. Because T1 didn't return success to its client until after commit-wait completed - that is, until true time had already passed s1 - and T2 only begins after that, T2 necessarily begins its own transaction and gets its own timestamp assignment at a moment when true time is already past s1. Spanner assigns T2 a timestamp s2 also based on `TT.now().latest` at that later moment, which is guaranteed to be greater than s1, since s1 is already in the past by the time T2's timestamp is chosen. So s2 > s1 is guaranteed by construction, not by luck, and any of T2's reads that should see T1's committed state will correctly do so because the ordering timestamps reflect the real happened-before relationship the clients experienced. ## What the guarantee costs The cost of this guarantee is direct and quantifiable: every committing read-write transaction pays a minimum latency penalty roughly equal to the width of the uncertainty interval at commit time, since the transaction must literally wait for real time to catch up to its assigned timestamp before proceeding. - With epsilon historically in the low single-digit milliseconds thanks to TrueTime's dedicated GPS/atomic clock infrastructure, this tax is small enough to be commercially acceptable for a general-purpose transactional database. - But it is not free, and it's added to every single read-write commit, not amortized across a batch. This is precisely why Spanner's designers invested in keeping epsilon as tight as possible rather than accepting whatever a generic NTP deployment would offer, since a wider epsilon translates directly and linearly into worse write latency and, because locks are held throughout the wait, reduced write throughput under contention. ## Read-write against read-only A useful worked comparison: | Transaction | What it pays | |---|---| | **Read-only** | Transactions in Spanner deliberately avoid paying this cost, since Spanner can serve a consistent snapshot read at a chosen past timestamp without needing to coordinate locks or wait for real time to catch up (it just needs to know the read is at a timestamp already guaranteed to be in the past, which existing committed data satisfies). | | **Read-write** | Commit-wait is specifically the price of read-write transactions that need their result visible and correctly ordered relative to real time going forward, which is the harder guarantee. | This asymmetry is a deliberate design choice: Spanner accepts a small, bounded latency cost on writes in exchange for a strong ordering guarantee, rather than either accepting weaker consistency (and the application-level bugs that can follow from stale or reordered reads) or paying for a much more expensive coordination mechanism, like routing every transaction through a single global sequencer, which would not scale the same way across Spanner's globally-distributed deployment footprint.
- Why does Spanner choose the commit timestamp based on TT.now().latest rather than TT.now().earliest?Choosing from the upper end of the current interval (latest) ensures the assigned timestamp is at least as large as any timestamp a causally-prior transaction could plausibly have been assigned, which keeps timestamps monotonically consistent with causal order. Combined with waiting on earliest during commit-wait, this pairing of 'pick high, wait for low-bound to catch up' is what makes the whole scheme correctly ordered.
- Does commit-wait mean two concurrent read-write transactions on completely unrelated data still get serialized against each other?Not necessarily in terms of blocking each other's execution - concurrent transactions on disjoint data can proceed in parallel and each independently pays its own commit-wait delay. What commit-wait guarantees is the timestamp ordering relative to real time, not that unrelated transactions must run one after another; only conflicting transactions require lock-based serialization on top of this.
- If a data center's TrueTime epsilon were reduced to near zero through even better clock hardware, would commit-wait become unnecessary?The wait would shrink toward zero but the mechanism itself would still conceptually be needed, since as long as any nonzero clock uncertainty exists, some wait is required to convert a chosen timestamp into a provable statement about real time. In the theoretical limit of a perfect, zero-uncertainty clock, the wait duration approaches zero, but Spanner's design doesn't assume such a clock is achievable, hence the deliberate architecture around bounded uncertainty rather than assumed perfection.
It's like a courier who, after handing off a signed, time-stamped package, deliberately waits at the door until their own watch's slowest possible reading has caught up to the timestamp on the receipt, before telling anyone the delivery happened - that way, anyone who hears about the delivery afterward can be completely sure it's genuinely in the past, not just probably in the past.
saying these in an interview costs you the question
- Thinks commit-wait waits for a fixed, hardcoded delay rather than for TT.now().earliest to pass the commit timestamp
- Confuses external consistency with plain serializability
- Doesn't connect commit-wait's cost to the TrueTime uncertainty interval width
- Believes read-only transactions also require commit-wait
- Can't explain why waiting for 'earliest' (not just any reading) provides a guarantee rather than a probability