skip to content

How does a record carrying its writer's timestamp rather than the node's arrival timestamp change a measured delivery interval?

level: middleimportance: should knowfreq 48%

answer

  1. whose clock stamped it
  2. before acceptance, or at acceptance
  3. batching and retries hide here
  4. a backfill's ages look enormous

basics

~20 s

A write-side timestamp puts the writer's own accumulation, batching and retries inside the measured interval and trusts a clock you do not control. An arrival timestamp is set by the cluster on acceptance, so it is comparable across records but hides everything that happened before the record was accepted.

solid answer

~50 s

Whatever time travels with the record is what the reader subtracts from its own clock, so it decides what the number means. A **write-side timestamp** is set by the writing client when it created or enqueued the record: it makes client-side delay visible — batching, a retry after a rejection, an hour spent buffered during an outage — and it is only as good as that host's clock, so one misconfigured writer poisons a fleet percentile. An **arrival timestamp** is set by the receiving node on acceptance: records in a stream then compare cleanly against each other because they come from the cluster's own hosts, but a record buffered for an hour looks brand new. Platforms differ in which they retain — one, the operator's choice, both, or neither — so where nothing trustworthy is carried, teams put a time field in the record body and own the problem themselves.

go deeper

for a junior

Recall that the age of a record depends on which time travelled with it: one set by the client that wrote it, one set by the node that accepted it. They are not interchangeable.

for a middle

Explain what each timestamp includes and excludes, and give a concrete case where they differ by a lot — a buffered writer, a retry, a replay of historical data.

for a senior

Demonstrate that you would state which timestamp a dashboard uses before interpreting it, monitor per-writer clock offset when trusting write-side values, and check what a cross-cluster copier does to the stamp.

for a principal

The estate-level call is whether the platform guarantees a trustworthy time on every record or leaves it to each team's payload. Guaranteeing it costs bytes and discipline; leaving it out means no two teams' numbers compare.

## The two candidate timestamps Any delivery interval computed from real traffic works the same way: the reader takes a time that travelled with the record and subtracts it from a time it observes itself. Two sources for that travelling time are common. - **Write-side** — stamped by the writing client at the moment it created or enqueued the record. - **Arrival** — stamped by the receiving broker node at the moment it accepted the record. The choice is not cosmetic. It moves a whole segment of the path into or out of the number, and it changes whose clock the number depends on. ## What each one makes measurable | | Write-side timestamp | Arrival timestamp | |---|---|---| | Set by | the writing client's host | the receiving node | | Client accumulation, batching, retries | inside the interval | outside it | | A record buffered an hour before sending | shows its true age | looks freshly written | | Comparability across records in one stream | only if every writer's clock agrees | good, within the cluster's own hosts | | Ordering along the stream | not necessarily increasing | ordinarily increasing | | A replayed backfill of old data | ages of days or months, immediately | the time of the replay | ## Where the two diverge sharply - **A writer that accumulates.** Records are held briefly to be sent together. Write-side counts the hold; arrival does not. - **A rejected write that is retried.** The successful attempt is what gets accepted, so arrival timestamps the retry, while write-side still carries the original attempt's time and exposes the delay. - **A writer that lost its connection.** Records buffered locally for minutes arrive in a burst. On arrival timestamps every one of them looks new, and the staleness the business is feeling is invisible in the measurement. - **A backfill.** Historical records replayed into a stream carry historical write-side times, so the measured age is the age of the data rather than of the path, and a dashboard reading write-side ages lights up with a false incident. ## Whose clock you are trusting A write-side timestamp is set on a host you may not administer — a client fleet, a partner's service, a mobile device. Nothing on the path validates it. A single host whose clock is minutes off contributes records whose computed ages are wrong by exactly that offset, and because percentiles are computed over all of them, the distribution splits into humps rather than simply shifting. Arrival timestamps come from the cluster's own hosts, which are usually administered and synchronised together, so they are wrong in a more uniform and more correctable way. Neither escapes the deeper problem: the reader is still subtracting a time set on one host from a time observed on another. ## What varies between platforms - Some platforms stamp arrival and keep nothing from the writer; some carry a write-side field; some let the operator choose per stream which of the two is retained; some expose both. - Where a record is copied into a second cluster, the copier commonly presents it as a fresh write to the destination, so an arrival timestamp there reflects the copy rather than the original write. Check this before trusting any age measured on the far side. - Where records are removed once acknowledged rather than kept for later readers, the same two candidates appear under different names — the time the sender enqueued it, against the time the server took it — and the trade-off is identical. ## A practical rule 1. Decide, per stream, which timestamp the delivery interval is defined against, and write that down next to the number. 2. If client-side delay must be visible, use write-side and accept the clock exposure — then monitor per-writer offset so you can tell a stale record from a wrong clock. 3. If comparability matters more, use arrival and measure client-side delay separately at the writer, where one host's clock is enough. 4. If nothing trustworthy travels with the record, put an application time field in the body. You then own the format, the clock discipline and the cost of the extra bytes. The answer an interviewer is listening for is not a preference. It is that you know the two numbers measure different spans of the path and depend on different clocks, and that you would say which one a given dashboard is showing before drawing any conclusion from it.

  • What happens to the timestamps when a record is copied into a second cluster?
    It depends on the copier, and that is the point. Many present the record to the destination as an ordinary write, so its arrival timestamp there reflects the copy, not the original write, and ages measured on the far side read near zero. A write-side field in the record survives the copy and keeps meaning.
  • A backfill replays a year of historical records. What does a write-side age show?
    Ages of up to a year, immediately, on a path that is perfectly healthy. Teams either define the backfill stream's interval against arrival timestamps, or exclude replay streams from the measurement, rather than trying to explain the spike every time.
  • Why is one writing host with a wrong clock worse than a general offset?
    A general offset shifts every sample the same way, so the distribution keeps its shape and a trend is still readable. One wrong host injects a second population, so percentiles become bimodal and neither hump describes the path. Breaking the measurement down by writing host is what exposes it.

saying these in an interview costs you the question

  • Treats a record's timestamp as trustworthy without asking whose clock set it
  • Assumes every platform carries a write-side timestamp
  • Thinks arrival stamping captures client batching and retries
  • Reads a backfill's huge ages as a live incident
  • Assumes a timestamp survives a copy into another cluster unchanged