skip to content

A reader group's total unread count is low and flat, yet one downstream report is hours stale — what reading was missing?

level: seniorimportance: should knowfreq 47%

answer

  1. a sum hides its distribution
  2. four healthy shares drown one
  3. maximum, not sum
  4. age, not count, for a stalled share
  5. the breakdown says which, not why

basics

~20 s

The group total is a sum, and a sum hides its distribution. One share of the work can be hours behind while the others are current, leaving the total small. The missing reading is the per-share breakdown, and specifically the maximum unread age across shares.

solid answer

~50 s

Where a stream is divided into parts and each member of a reader group holds a share of them, the group-level gap you see on a chart is an aggregate — usually a sum of counts, sometimes an average. Both hide skew. A share carrying a low record volume can be completely stalled and still contribute only a few hundred records to a total of tens of thousands, so the aggregate stays flat while everything routed to that share goes unprocessed. The reading that exposes it is per-share, and the most useful single number is the **maximum unread age across shares**, not the sum of counts: a maximum cannot be diluted by healthy siblings the way a sum can. Skew of this kind has several possible causes — uneven routing of records across shares, uneven member capacity, or a share whose holder is making no progress — and the breakdown tells you *which share*, not *why*. That second step is separate work.

code

json · 11 lines
json
{
  "readerGroup": "order-enrichment",
  "totalUnreadRecords": 12280,
  "shares": [
    { "share": 0, "unreadRecords": 3100, "oldestUnreadSeconds": 5 },
    { "share": 1, "unreadRecords": 2800, "oldestUnreadSeconds": 6 },
    { "share": 2, "unreadRecords": 3300, "oldestUnreadSeconds": 4 },
    { "share": 3, "unreadRecords": 2900, "oldestUnreadSeconds": 5 },
    { "share": 4, "unreadRecords": 180, "oldestUnreadSeconds": 9930 }
  ]
}

go deeper

for a junior

Understand that a group-level number is a sum over several separate gaps, so one part of the work can be far behind without the headline number moving much.

for a middle

Explain the arithmetic: why four healthy shares dominate a sum, why a low-volume stalled share contributes almost nothing to it, and why an average is no better than a sum here.

for a senior

Show that you reach for the maximum unread age across shares as the detector, and that you use the count-versus-age pattern on the offending share to classify the skew before you act on it.

for a principal

Argue for what a shared stream's owners are required to expose — a per-share distribution rather than a single headline — because an estate where every team publishes only a total cannot see this class of fault at all.

## Why a group total can lie On platforms that split a stream into parts and hand each part to one member of a reader group, there is not one gap — there is one gap per part. What a chart labelled with the group's name shows is an aggregate over those, and the aggregate is nearly always a **sum of unread counts**. Sums are exactly the wrong summary for the failure that matters here. The arithmetic is unforgiving. Suppose a group holds five shares. Four of them are current and carry three thousand unread records each as ordinary in-flight bulk. The fifth has not moved in three hours, but it is a low-volume share and has accumulated only one hundred and eighty records in that time. The total is a little over twelve thousand, sitting exactly where it always sits. Every routing decision that sent a record to the fifth share is three hours stale, and the group-level view shows a flat line. ## Why the maximum, and why age Two choices make the breakdown useful: - **Maximum rather than sum or average.** A maximum across shares cannot be diluted. Four healthy shares dominate a sum and drag an average down; neither can move a maximum. - **Age rather than count.** The stalled share's *count* is small — that is precisely why it hid. Its *age* is three hours, which is enormous by any standard. Age is the reading that scales with how long something has been broken rather than with how busy it is, and that makes it the right reading for detecting a share that is not moving. The combination — the largest unread age over all shares — is the single number that would have surfaced this situation, and it is a strictly stronger signal than the group total for this class of fault. ## What the breakdown does and does not tell you The breakdown identifies *which* share is behind. It stops there, and that boundary is worth respecting in an interview answer, because the causes are genuinely different problems with different fixes: 1. **Routing skew** — records are not spread evenly across shares, so one share legitimately carries far more work than its siblings. The signature is a share with a large count *and* a growing age, while its siblings are fine. 2. **Capacity skew** — the member holding that share is on slower hardware, or holds more shares than its peers, so it drains more slowly. The signature is several shares behind at once, all of them held by the same member. 3. **A share that is not progressing at all** — the signature is the one above: small count, very large age, and an age that climbs at exactly one second per second. The third has its own diagnosis and its own remedies, which are not part of reading the signal. ## Where no per-share breakdown exists This whole reading presumes a platform that splits a stream and attributes a position per part. Designs where consumers simply compete for records from a shared queue have no shares to break down and no per-reader position — there is one depth and one oldest-waiting-record age for the whole queue. Skew of the kind above cannot occur in the same form, because no record is bound to a particular consumer before delivery. What replaces it is a different question entirely: whether a specific record has been handed out repeatedly and never completed, which the depth reading also does not show. The transferable lesson is the same in both worlds: **a single aggregated number is a summary, and every summary throws away the distribution that the interesting failures live in.** ## The practical reading habit | Reading | What it catches | What it misses | |---|---|---| | Group total unread count | Whole-group overload; anything that moves every share at once | Any fault confined to a minority of shares, especially low-volume ones | | Maximum unread age across shares | A share that has stopped or fallen far behind, whatever its volume | How much total work remains to be done | | Per-share counts, listed | Routing skew, and which member is the slow one | Anything about records already handed out but not finished | The habit that follows is simple: treat the group total as the thing you glance at and the per-share distribution as the thing you actually read. A group total is a reasonable summary of load; it is a poor detector of a localised stall, and it was never designed to be one.

  • Why is the maximum unread age a better group-level summary than the average?
    An average is a sum in disguise and is dominated by the healthy majority. With four shares at five seconds and one at three hours, the average is around half an hour — high enough to look odd, low enough to explain away, and it moves toward the healthy value every time a share is added. The maximum reports the worst share exactly, whatever the group size.
  • How would you distinguish routing skew from a share that has simply stopped?
    Compare the count against the age on that share. Routing skew gives a large count with an age that grows slowly, because the member is still draining, just not fast enough. A stopped share gives a small or static count with an age that climbs at one second per second, because nothing is being read at all. The count-versus-age pair separates them cleanly.
  • Does adding more members to the reader group help when the problem is one behind share?
    Usually not on its own. On platforms that bind each part of a stream to one member, a new member can only take work if there are shares to hand it, and it does not make the existing behind share drain faster unless the share is moved to it. The reading here identifies that constraint; acting on it is a separate exercise.

saying these in an interview costs you the question

  • Trusts a group total to reveal a localised stall
  • Uses the average unread age across shares as the summary
  • Assumes a small unread count means a small problem
  • Reads only counts and never the age per share
  • Believes every share must carry equal record volume
  • Concludes which member is at fault from the total alone