A node's mean publish latency is 9 ms while its p99 is 700 ms — which number does a writing client feel, and why?
answer
- one average over a mixed population
- cheap requests outnumber costly ones
- the writer waits on its own request
- high percentile, split by request kind
basics
~20 sThe p99. Most requests a node serves are cheap, so the mean tracks them and stays flat, while the tail is where a writer actually waits — and one writer in a hundred waiting 700 ms is what a customer reports.
solid answer
~50 sA broker node serves a mixed population on the same connections: small metadata and administrative calls, reads answered out of memory, and writes that may be held until other copies accept them. The cheap requests are the numerous ones, so the **mean** describes them and can sit flat while the expensive minority gets much worse. A writing client does not experience the average of everyone's requests — it blocks on its own, and the write it is blocked on is typically one of the expensive ones. That is why publish and fetch latency are read as high percentiles, per request kind, rather than as a single average. Note that platforms differ in what they hand you: some export a distribution you can read a percentile from, some export only an average and a maximum, and some rented services expose neither at node level.
go deeper
Recall that the average and the high percentile answer different questions, and that a client waits on its own request rather than on the average of everyone's.
Explain why a broker's request population is lopsided — frequent cheap metadata and near-empty reads against a costly minority — and why that makes the mean insensitive to a real tail.
Show how you would attribute a moved tail: split by request kind, compare median against the high percentile to see whether the distribution stretched, and check whether the traffic mix changed rather than the work.
Decide what the platform team publishes and guarantees per node, and what the estate pays to keep per-request-kind distributions rather than averages on every cluster instead of the few that matter.
## What the two numbers are answering A broker node serves a **mixed population** of requests over the same connections. The **mean** answers "what did the average request cost the node?". A **high percentile** — the value below which 99% or 99.9% of requests fell — answers "how bad was it for the unlucky ones?". On a broker those are almost always answered by *different requests*, which is why the two numbers can move independently. ## Why the mean is dominated by cheap requests The population a node handles is lopsided by construction: - **Small administrative and metadata calls** are frequent and finish almost immediately. Clients ask where partitions live, announce themselves, and keep connections warm. - **Reads answered from memory** — records written seconds ago and still resident — are far cheaper than reads that have to be pulled off the volume. - **Empty or near-empty responses** are common: many readers poll far more often than records arrive, and each of those round trips is a fast request in the population. - **The expensive requests are the minority.** On platforms where an accepted write waits until other copies hold it, that waiting sits inside the write's measured duration; on platforms where it does not, a large write still costs more than a metadata call. So the arithmetic mean is a weighted average in which the cheap majority carries nearly all the weight. A tenfold slowdown confined to 1% of requests barely nudges it. ## What the writing client actually experiences A client that waits for its write to be acknowledged experiences **its own request's duration**, not the node's average. Two consequences follow: 1. The requests a client blocks longest on are exactly the ones the mean under-weights, so the dashboard and the customer disagree. 2. Because a client holds a bounded number of requests outstanding, a slow one delays the ones behind it on the same connection — the tail does not stay confined to the request that was unlucky. This is why a support report of "writes are hanging" routinely arrives while the latency panel looks unchanged. ## How the candidate readings compare | Reading | What it answers well | Where it fails | |---|---|---| | Mean request time | Total work the node is doing per request | Hides a slow minority; moves when the traffic *mix* moves | | Maximum | Proves a bad case existed | One sample, unrepeatable, often a lone outlier | | High percentile per request kind | What the unlucky share of real clients waited | Costs more to store; still only this one hop | | Ratio of high percentile to median | The *shape* — how stretched the distribution is | Not a magnitude; needs both series kept | ## What varies across platforms Do not assume the numbers are on offer in the same shape everywhere: - Some platforms expose a full distribution per request kind; others expose an average and a maximum only, from which no honest tail can be recovered. - Some rented services publish only account- or cluster-level aggregates, so per-node tails are simply not visible to the tenant. - Where the write path waits on other copies, that wait may be counted inside the request's duration or reported separately — establish which before reading the tail as "the node is slow". ## Reading it in practice 1. **Split by request kind first.** A single percentile over writes, reads and metadata calls together is a blend whose movement you cannot attribute. 2. **Keep median and a high percentile side by side.** The median tells you whether the typical request changed; the gap between them tells you whether the distribution stretched. 3. **Treat the maximum as a hint, not a signal.** It is a single sample and reacts to one unlucky request. 4. **Remember this is one hop.** The node's tail is what happened inside the node; what a writer-to-reader path took end to end is a different measurement with its own pitfalls, and a green node tail does not vouch for it. The short version an interviewer wants: the mean describes the requests that are cheap and numerous, the tail describes the requests that people are waiting on, and on a broker those two sets barely overlap.
- The p99 rose but the median did not. What does that combination narrow it down to?The typical request is unchanged, so the work itself did not get uniformly slower. Something is affecting a minority: a subset of streams or clients, requests that hit the volume rather than memory, requests waiting on another copy, or brief contention on the node. Look for a dimension along which that minority differs, rather than for a cluster-wide cause.
- Both the mean and the p99 fell after a client change, and nobody made the node faster. What happened?Most likely the traffic mix changed, not the node. If a client started sending fewer, larger requests, or a chatty reader backed off, the population being summarised is different — cheaper requests are a bigger or smaller share. A latency number over a mixed population moves with the mix, which is why the request count per kind is read alongside it.
- A rented cluster exposes only an average request time. What can you still do?Measure from the client side, where you control instrumentation: record each call's duration in the writing and reading applications and read percentiles there. That number includes the network and the client's own queueing, so it is not the node's request time, but it is the number your users experience and it is honest about the tail.
saying these in an interview costs you the question
- Thinks a flat mean proves no client is waiting
- Reads one latency percentile across every request kind at once
- Treats the maximum as the tail signal to watch
- Assumes a rising tail is always a network problem
- Reads a node's request percentile as the whole path's time