A stream's delivery interval is four minutes while its writes are acknowledged in six milliseconds — what is the interval measuring?
answer
- whole path, not one hop
- acceptance through to finished work
- dwell dominates the total
- per stream and reader group
basics
~20 sThe delivery interval covers the whole path: from a write being accepted to a reader finishing the work on that record. Acknowledgement time covers only the first hop, so the wait in the stream and the reader's own processing are what make up the four minutes.
solid answer
~40 sAcknowledgement time is one request's duration at a broker node — arrival to reply. The delivery interval is the record's whole journey: accepted, then durable and unread while the reader works through everything ahead of it, then delivered on the reader's next read, then processed. Most of that is not spent inside any request, which is why a cluster answering everything in single-digit milliseconds is perfectly consistent with a four-minute interval. The dominant segment is almost always **dwell** — the record sitting in the stream because the reader has not reached it yet. It is measured per stream and per reader group, never as one cluster-wide number, because every reader has its own pace and its own backlog.
go deeper
Recall the two endpoints: an accepted write at one end, a reader having finished the work at the other. Be ready to say that a fast acknowledgement does not mean fast delivery.
Explain the segments — accumulation, acceptance, dwell, delivery, processing — and say which of them the cluster's own panels can see. Name dwell as the segment that usually holds the minutes.
Show that you would measure it per stream and reader group, define where the clock starts, and use the divergence between the interval and the hop latencies to decide which side of the path to investigate first.
The judgment is coverage against cost: measuring delivery honestly everywhere means probe traffic, storage and a reader per stream. Decide which streams carry a published commitment and which get nothing but cluster-side numbers.
## What the delivery interval is The **delivery interval** is one number for the whole path: the elapsed time from a write being accepted by the cluster to a reader finishing the work that record triggers. It is deliberately not a request duration. Every other latency number on a broker dashboard describes one hop — one request arriving at a broker node and being answered — while the delivery interval describes a record's entire journey, most of which is not spent inside any request at all. Teams differ on exactly where the clock starts. Some start it at the writing client's send call, which folds in client-side accumulation and retries; others start it at acceptance, because that is the first moment the platform is accountable for the record. Either is defensible, but the choice has to be written down, because the two differ by everything that happens before the record leaves the writer. ## Why the hop numbers do not add up to it | Number | What it covers | What it cannot see | |---|---|---| | Write request duration at a broker node | one request, arrival to reply | client accumulation before it, and everything after acceptance | | Read request duration | serving one read | how long the record waited unread before that read | | A reader group's record lag | how much is still unread | how long the reader takes to finish each record | | Delivery interval | the whole path, end to end | which segment of the path is at fault | Hop latencies are typically small and stable. The delivery interval is usually dominated by **dwell**, and dwell is not a latency anyone measured — it is the record simply sitting there. That is the whole reason the interval is measured separately instead of being inferred from the panels the cluster already publishes. ## The segments the interval is made of 1. **Accumulation at the writer** — the record waits to be sent in a batch, or is being retried after a rejection. 2. **Acceptance** — the receiving node stores it and waits for however many copies the configured durability demands before answering. 3. **Dwell** — the record is durable and unread. Where readers own a stored read position and records persist after being read, this shows as the reader's unread backlog; where a record is removed once it is acknowledged, the same dwell shows as queue depth and the age of the oldest queued record. Either way it is the same time. 4. **Delivery** — the reader's next read actually returns the record, which on a polling reader includes whatever is left of the poll cycle. 5. **Processing** — the reader does the work, including whatever it calls downstream before it counts the record as done. Only segments 2 and 4 appear on the cluster's own latency panels. Segments 1, 3 and 5 are invisible to the cluster, and they are where the minutes are. ## Two ways to obtain the number - **From real traffic.** The reader compares a time carried on each record against its own clock and reports the result. This covers every record you actually care about, but it subtracts one host's clock from another's, so it is only as trustworthy as the offset between them. - **From a synthetic probe record.** A record is injected on a schedule purely to be timed, and the same host that wrote it observes the finished result, so the whole measurement stays on one clock. It is honest about the path but only measures the probe's path. Estates that care about this usually run both: the probe for a number that is not corrupted by clock offsets, real-traffic ages for coverage of the streams and readers that matter. ## Reading it during an incident - Interval rising while write and read request durations stay flat: the cluster is doing its job and the reader side is not keeping up. - Interval rising together with request durations and node saturation: the cluster itself is the constraint. - Interval small and flat while the business still complains: the path is carrying records promptly, and the question has moved to whether any useful work is being done with them — a different signal entirely. - Interval published for one stream only: you know nothing about the others. There is no cluster-wide delivery interval, because there is no cluster-wide reader. ## Why anyone bothers During an incident the cluster team and the application team routinely disagree, each reading a green panel of its own, because each is measuring its own share of the path. The delivery interval is the number the person actually waiting for the record experiences. It is also the number most likely to end up in a commitment to a downstream team, because it is the only one they can feel.
- Should the delivery interval be measured per stream, or once for the whole cluster?Per stream and per reader group. Each pair has its own writers, its own reader pace and its own unread backlog. A cluster-wide average blends a nightly bulk reader with a payments reader, and the resulting number describes neither: it is routinely four hours for one and forty milliseconds for the other.
- A writing client buffered records locally during a network outage. Does the interval show it?Only if the clock starts before the write is accepted. Measured from acceptance, the buffered hour is invisible and the interval looks healthy while the data is stale. That is exactly why the start point has to be defined explicitly rather than assumed.
- What does a delivery interval of a few milliseconds actually prove?That the path carried one record quickly. It says nothing about whether the reader's work had any effect, nothing about streams you are not measuring, and — if the number came from subtracting two hosts' clocks — possibly nothing at all, since the clock offset can be larger than the interval itself.
saying these in an interview costs you the question
- Quotes the write acknowledgement time as the delivery time
- Adds the hop latencies together and expects the total
- Treats the interval as one cluster-wide number
- Assumes a green cluster dashboard bounds the interval
- Leaves the reader's processing time out as someone else's problem