skip to content

Delivery intervals computed as the reading host's clock minus a record's write-side timestamp sometimes come out negative — what is actually being measured?

level: seniorimportance: should knowfreq 52%

answer

  1. two clocks, one subtraction
  2. negative samples are the tell
  3. synchronisation bounds, never removes
  4. time a round trip instead

basics

~20 s

The difference between two hosts' clocks, plus the real interval. A negative result proves the offset between the writing host and the reading host is larger than the journey itself, so every sample from that pair is wrong by that offset — the small ones ruinously.

solid answer

~50 s

Subtracting a time set on one host from a time observed on another does not yield a duration: it yields the duration plus the unknown offset between the two clocks. When the offset happens to exceed the real interval and runs the wrong way, the result goes negative, which is the visible tail of an error that is corrupting every sample quietly. Running a clock synchronisation service bounds that offset but never removes it, and on a fast path the residual is easily the same order as the number you are trying to measure. The honest repair is to stop subtracting across hosts: have one host write a probe record and observe its finished result, so the whole interval is one clock's arithmetic. Where a cross-host number must be kept, treat it as a trend, break it down by writing host, and publish the clock offset beside it.

go deeper

for a junior

Recall that two machines do not agree on the time, so a duration computed from one machine's stamp and another machine's clock is not a duration. A negative result is the giveaway.

for a middle

Explain that the subtraction returns the interval plus the difference of two clock offsets, and that synchronisation narrows that difference without eliminating it.

for a senior

Show the diagnosis — negative samples, host-split humps, step changes at deployment — and the repair: measure a round trip on one host's clock and keep the cross-host number only as a trend.

for a principal

The call is what the platform promises. Publishing an absolute delivery interval across an estate means owning clock discipline or owning probes; promising one without either is a number the organisation will act on and should not.

## A difference of two clocks is not a duration The measurement in question is `observed_on_reader − stamped_on_writer`. Those two numbers come from two different hosts, and each host's clock has its own offset from true time. What the subtraction returns is therefore the real interval **plus** the difference of the two offsets. Nothing in the arithmetic separates them. When that difference is small relative to the interval — a path that takes minutes, hosts that agree within milliseconds — the error is negligible and the number is useful. When the path is fast, the error can dominate completely, and a negative value is simply the case where it dominates and points the wrong way. A negative delivery interval is not a paradox to be explained with reordering or retries; it is proof that the method has failed for that host pair, and a strong hint that the positive samples from the same pair are wrong by a similar amount. ## What clock synchronisation buys, and what it does not - It **bounds** the offset; it does not remove it. Every synchronised host sits somewhere inside an uncertainty band rather than on true time. - The band is not constant. Clocks drift between corrections, and a correction is applied as a step or a slew that itself perturbs short measurements. - A host can lose synchronisation silently and keep serving. Nothing about its records looks different; only its computed ages move. - The band is set by conditions outside the platform, so it is the same for a path that takes ten seconds and one that takes ten milliseconds — which is exactly why the second one is the one it ruins. The practical rule: **a cross-host interval is only trustworthy when the interval is much larger than the offset bound you can actually assert.** ## How skew shows up in the data - Negative samples, usually a small fraction, usually from a recognisable subset of writers. - A distribution with two or more humps instead of one, where the humps split by writing host rather than by workload. - A step change in the interval at a deployment, when hosts are replaced by others with a different offset, even though nothing about the path changed. - A stream whose measured interval is suspiciously constant — the offset is large and stable, and it is what you are plotting. - Cluster-side and reader-side dashboards that disagree by a fixed amount rather than a varying one. ## Measurements ranked by how many clocks they involve | Measurement | Clocks involved | What it can be trusted for | |---|---|---| | One host writes a probe record and observes its finished result | one | an absolute interval, including short ones | | Reader subtracts a record's travelling timestamp from its own clock | two | a trend, and absolutes only when the interval is large | | Two dashboards each computing their own half of the path | three or more | almost nothing absolute; use for direction only | ## Measuring on one clock The repair that actually works is to arrange the measurement so both ends are read from the same clock. A **synthetic probe record** is written on a schedule by a probe writer; the result of processing it is observed by the same host, and the round trip is timed there. No cross-host subtraction occurs, so the offset cancels out of the arithmetic entirely rather than being estimated and subtracted back. The price is coverage: the probe times the probe's journey, not each real record's. That is why the probe and real-traffic ages answer different questions and are usually run together. ## Keeping a cross-host number usable 1. **Publish the offset next to the interval.** If you cannot say how far apart the clocks are, you cannot say what the number means, and the first question in an incident will be unanswerable. 2. **Break the measurement down by writing host.** A bimodal distribution becomes obvious and attributable in one chart. 3. **Bucket coarsely.** If the offset bound is tens of milliseconds, reporting the interval at millisecond resolution invents precision that is not there. 4. **Watch the trend, not the level.** The offset is roughly constant over short spans, so a change in the measured value is real even when the value itself is not. 5. **Never build a decision on an absolute cross-host number of the same order as the offset bound.** That includes any threshold you would be woken by. What separates a senior answer here is refusing the two easy outs: dropping the negative samples and keeping the rest, or declaring the clocks synchronised and moving on. The first hides the evidence while keeping the error; the second mistakes a bound for an equality.

  • How would you tell skew apart from genuinely fast delivery on a low-latency path?
    Look for the signatures rather than the values: negative samples, humps that split by writing host, step changes when hosts are replaced. Then compare against a probe timed on one clock. If the single-clock number and the cross-host number disagree by a constant, you are plotting an offset.
  • Is it acceptable to discard the negative samples and keep the rest?
    No. The negatives are the only visible members of a population that is wrong in both directions; dropping them removes the evidence and biases the remaining distribution upward. Keep them, count them, and treat their share as a health indicator for the measurement itself.
  • The interval is measured in minutes. Does skew still matter?
    Much less. When the real interval is orders of magnitude larger than any plausible offset, the error is noise and the number is usable as an absolute. The discipline is to know which regime you are in rather than to apply one answer everywhere.

Timing a courier by comparing the sender's wristwatch against the receiver's wall clock. If the two are set five minutes apart, a parcel that took one minute appears to arrive four minutes before it was sent. The repair is not a cleverer subtraction — it is one clock: the sender waits for the parcel to come back and times the round trip on the watch already on their wrist.

saying these in an interview costs you the question

  • Says the clocks are synchronised so the subtraction is fine
  • Explains negative intervals as records processed before they were written
  • Drops the negative samples and trusts the remainder
  • Sets a threshold on an absolute number smaller than the offset bound
  • Thinks a monotonic timer fixes a measurement spanning two hosts