skip to content

Why is one aggregate request-latency percentile per broker node a weaker signal than separate publish and fetch series?

level: middleimportance: should knowfreq 46%

answer

  1. unrelated work in one number
  2. the mix moves it, not the work
  3. cannot say which side is hurting
  4. split by kind, node and stream
  5. each split costs series

basics

~20 s

Because writes, reads and metadata calls do unrelated work with unrelated costs. One percentile over all of them moves when the traffic mix changes, not only when something slows, and it never says which side of the node is hurting.

solid answer

~50 s

A node's requests fall into kinds that have almost nothing in common: a write may be held until other copies hold it, a read is cheap from memory and expensive off the volume, and metadata calls are numerous and trivial. A percentile taken over that blend has two defects. First, it is **mix-sensitive**: change the proportions — a chatty reader backs off, a batch job starts — and the number moves although every kind is exactly as fast as before. Second, it is **unattributable**: a rise tells you something in the node got slower without telling you whether writers or readers are affected, which is the first question anyone asks. Separate series per kind fix both, at the cost of more series to store — a cost your metrics practice, not the broker, decides how far to carry.

go deeper

for a junior

Recall that a node answers several unrelated kinds of request, and that one latency number covering all of them cannot say which kind is slow.

for a middle

Explain mix-sensitivity concretely: proportions change, the blended percentile moves, and nothing actually got slower — then say which splits recover the meaning.

for a senior

Show the diagnosis a split enables, decide which dimensions are worth the series they cost, and name the client-side fallback when a rented cluster exposes only an aggregate.

for a principal

Set which series the platform team keeps and publishes across the estate, and defend that storage cost against the alternative of an unattributable number on every dashboard.

## One number over a mixed population A broker node answers several unrelated kinds of request on the same connections. Grouped roughly: - **Writes** — accept records onto a stream. On platforms where an accepted write waits until other copies hold it, that waiting is part of the request's measured duration, so a write's cost depends on the state of other nodes. - **Reads** — return records to a reader. Cheap when the records are still resident in memory, substantially more expensive when they must come off the volume, and frequently held deliberately until records or a size threshold arrive. - **Metadata and administrative calls** — where partitions live, who is in a reader group, connection housekeeping. Numerous and fast. A single latency percentile over that blend is a summary of a population whose members are not comparable. Two specific failures follow. ## Failure one: the mix moves the number A percentile describes the population it was taken over. If the proportions change, the number changes even though no individual kind got slower: - a reader that polled constantly backs off, removing a mass of cheap requests, and the blended percentile rises; - a nightly batch reader starts, adding expensive volume-backed reads, and it rises again for a different reason; - a client switches to fewer, larger requests, and the count-weighted picture shifts entirely. So the blended series generates movement that is not a regression and hides movement that is. Reading it next to the request **count per kind** is the minimum defence if you only have the blend. ## Failure two: the number cannot be attributed The first question in any incident is which side is affected: are writers being made to wait, or readers? A blended percentile cannot answer it, and neither can it separate a slow-work problem from a shared-capacity one, since it merges requests whose costs are driven by different things. Splitting by kind turns one ambiguous line into a diagnosis: | Series | Moves when | Points at | |---|---|---| | Write duration | Records are larger, or another copy is slow to hold them | The write path and the copy set behind it | | Read duration | Readers fall behind and records leave memory; larger reads | The read path and how far behind readers are | | Metadata call duration | The coordination side is slow or contended | Cluster metadata, not the data path | | Blended percentile | Any of the above, or the mix | Nothing in particular | ## Which splits are worth keeping Beyond the kind, the dimensions that repeatedly earn their place are: 1. **The node**, so a problem concentrated on one machine is visible instead of averaged across the fleet. 2. **The stream**, so a moved tail can be tied to the traffic responsible. 3. **The two halves of the duration** — waiting for a handler against being handled — which separates a capacity problem from a work problem within each kind. Each split multiplies the number of series stored, and that cost is real; how far to carry it is a decision for your metrics practice rather than something the broker settles. ## What varies across platforms - Some platforms expose durations per request kind and even per stream; some expose one duration for the node; rented services frequently expose an account-level aggregate only. - The set of kinds is not the same everywhere. A design that keeps records until acknowledged and deletes them has request kinds a log-shaped design does not, and the reverse. - Where the node exposes nothing you can split, the fallback is client-side measurement: the writing and reading applications already know which kind of call they made, so the split survives even when the node's own numbers do not. ## The point to make in an interview A blended percentile is not wrong so much as **unanswerable**: when it moves you cannot say whether the work changed or the traffic did, nor who is affected. The split is what converts a latency panel from a mood indicator into the first step of a diagnosis.

  • You only have the blended percentile. What single extra series makes it far more usable?
    The request count per kind over the same window. Most spurious movement in a blended percentile comes from the proportions changing, so seeing that reads doubled or that chatty metadata traffic vanished explains the move without any claim about speed. It does not localise a genuine slowdown, but it separates the false alarms from the real ones.
  • Is splitting by stream as well as by kind always worth it?
    Not always. It is decisive when several tenants share a node, because it ties a moved tail to the traffic responsible, and close to useless on a cluster carrying one workload. It also multiplies stored series, so the usual compromise is to keep the per-kind split everywhere and add the per-stream split on shared clusters where attribution is the recurring question.

saying these in an interview costs you the question

  • Treats one node latency line as the whole health of the node
  • Reads a moved blended percentile as proof something got slower
  • Assumes writes and reads have comparable cost per request
  • Ignores request count per kind when only the blend exists
  • Splits by every available dimension without weighing the storage cost