skip to content

questions

4

A reader group's lag shows two million unread records, so why does that count alone not say whether anyone is waiting?

level: juniorimportance: must knowfreq 70%

answer

  1. records, or seconds
  2. a count carries no clock
  3. the arrival rate converts them
  4. age of the oldest unread record
  5. count sizes the work, age dates it

basics

~20 s

Unread count states the gap in records; unread age states it in time. Two million records may be seconds old on a fast stream or half a day old on a slow one, so only the age reading shows staleness.

solid answer

~60 s

A reader group's lag is the gap between the newest record on a stream and the point the group has read up to, and it has two honest readings. The **unread count** says how many records remain unread. The **unread age** says how old the oldest unread record is. They are convertible only if you already know the arrival rate, which is exactly the thing nobody looks up mid-incident: two million records is a few seconds on a stream taking hundreds of thousands per second and most of a day on one taking fifty. The age is the reading that answers staleness directly, because it is already in the unit the business cares about. The count still earns its place — it is what you compare against how many readers can work in parallel, and its direction over an observation window is the diagnosis. Where the broker keeps no read position at all, the same pair shows up as queue depth and the age of the oldest waiting record.

go deeper

for a junior

Be able to say what the gap is and that it has two units: how many records are unread, and how old the oldest unread one is. Knowing that a count cannot be read as a delay is most of the answer.

for a middle

Explain why the two readings are not interchangeable: the conversion factor is the arrival rate, and that rate moves. Say which question each reading answers and what each one is blind to.

for a senior

Demonstrate that you reach for the age first during an incident because it is already in business units, and that you keep the count to size the remaining work. Mention the age ceiling imposed by how long records are kept.

for a principal

Frame it as an estate decision: which of the two readings a team is required to publish for a shared stream, so that incidents across different services are described in comparable terms rather than in whichever number that team's platform makes easy.

## What the gap actually is On a platform where a reader records how far it has got, the **read position** is a marker: the point in a stream that a reader group has processed up to. The stream also has a newest end, which the writing side keeps pushing forward. **Lag** is the distance between those two points, and on this subject it always means *reader* lag — not how far a data copy on another machine is behind its leader, and not how long a record takes to travel from write to processing. That distance is a geometric fact, and like any distance it can be measured in more than one unit. Two are used in practice, and they answer different questions. ## The two readings - **Unread count** — how many records sit between the read position and the newest record. A pure integer. Cheap to compute, because both ends are just markers. - **Unread age** — how old the *oldest unread record* is: the wall-clock difference between now and the timestamp of the record sitting immediately at the read position. Expressed in seconds or minutes. | Reading | Unit | Answers directly | Blind to | |---|---|---|---| | Unread count | records | how much work remains; how it compares with how many readers can share it | how long anyone has been waiting | | Unread age | time | how stale the output is right now | how much work it will take to close | The conversion between them is the arrival rate, and that is the catch: the rate is not constant. It has a daily shape, it spikes on a campaign, it collapses when an upstream writer dies. So the count is only translatable into a delay if you already know the rate over the same observation window — and during an incident that is one more thing to go and find. ## Why a count alone is ambiguous Consider three streams, each showing exactly two million unread records: 1. A high-volume click stream taking three hundred thousand records a second. Two million is under seven seconds of material. Nobody is waiting; a reader that paused for a garbage-collection cycle produces this. 2. A moderate order stream taking two thousand a second. Two million is about seventeen minutes. Somebody downstream is noticing. 3. A low-volume settlement stream taking fifty a second. Two million is eleven hours. This is an incident that started yesterday. The same number, three completely different situations. Nothing in the count distinguishes them. The unread age separates them instantly: seven seconds, seventeen minutes, eleven hours. ## Why the age alone is not enough either The age has its own blind spot: it says nothing about volume. An age of thirty minutes on a stream that receives one record an hour means a single record is late. The same thirty minutes on a heavy stream may mean tens of millions of records are queued behind it, which is a different-sized problem even though the staleness reading is identical. The count is what tells you the *size* of what has to be worked through, and therefore whether the readers you have could ever close it. There is also a ceiling effect worth knowing about. On platforms that remove records after a retention period, the oldest unread record can be deleted out from under the reader — at which point the age reading stops climbing and starts reporting the age of whatever is now oldest. An age curve that flattens at a suspiciously round value is often hitting that ceiling rather than recovering. ## Where there is no position to subtract from Not every platform stores a read position. Designs that hand a record to one consumer and delete it on acknowledgement have no marker to measure from, and no notion of "the newest record" as a fixed reference. The same signal still exists, read differently: - **Queue depth** — the number of records waiting to be handed out, which plays the role of the unread count. - **The age of the oldest waiting record**, which plays the role of the unread age. The pair behaves the same way and carries the same ambiguity, so the habit transfers. What does not transfer is per-reader attribution: with no stored position there is nothing that says how far *one* reader is behind, only how much is waiting for all of them together. ## Reading them as a pair The practical habit is to keep both on screen and read them together, because each one's blind spot is the other's strength. A large count with a small age is normal bulk in flight. A small count with a large age is a small amount of very old work, which usually means something is not moving rather than that something is overloaded. Both large is the straightforward backlog. Both small is the steady state — with one caveat worth remembering: a gap that reads as zero can also mean no records are arriving at all, which is a writer-side problem wearing a healthy reader's clothes.

  • Why is the unread age usually the harder of the two readings for a platform to provide?
    The count is arithmetic on two markers the broker already holds, so it costs nothing. The age needs the *timestamp of the record at the read position*, which means actually looking at a record rather than a marker, and it needs a trustworthy timestamp on that record. Where the timestamp is set by the writer's clock rather than the broker's, the age reading inherits that clock's skew.
  • A reader group's unread count is zero. What are the possible explanations?
    Three, and they are not equally good news. The readers are keeping up; or no records are arriving, so there is nothing to be behind on; or the gap is being computed against a position that is not advancing for a reason that makes the arithmetic meaningless. A zero should always be read next to the arrival rate over the same observation window.
  • Does a gap measured in records mean the same thing on every platform?
    The arithmetic does, but the shape it applies to differs. Where a stream is split into parts and each part carries its own position, the group figure is an aggregate over many separate gaps. Where consumers simply compete for records with no stored position, there is one depth for the whole stream and no per-reader gap at all.

A queue at a service counter. The number of people in line is the unread count; how long the person at the front has already been waiting is the unread age. Twenty people is nothing at a fast counter and an afternoon at a slow one, and only the second number tells you whether anyone has been abandoned.

saying these in an interview costs you the question

  • Treats unread count as if it were a duration
  • Assumes a fixed arrival rate when converting records to time
  • Says a gap of zero always means everything is healthy
  • Confuses reader lag with how far a data copy trails its leader
  • Thinks the age reading is available for free everywhere
  • Reports the gap without ever naming which unit is meant
open as a page

A reader group's gap is flat, rising steadily, or sawtoothing over a one-hour observation window — what does each shape mean?

level: middleimportance: must knowfreq 58%

basics

~20 s

Flat means arrival and drain rates match, at whatever height. A steady rise means the drain rate is below the arrival rate, and the slope is the difference. A sawtooth means arrival or reading happens in bursts rather than continuously.

open as a page

A reader group's total unread count is low and flat, yet one downstream report is hours stale — what reading was missing?

level: seniorimportance: should knowfreq 47%

basics

~20 s

The group total is a sum, and a sum hides its distribution. One share of the work can be hours behind while the others are current, leaving the total small. The missing reading is the per-share breakdown, and specifically the maximum unread age across shares.

open as a page

On a broker that deletes each record once it is acknowledged and stores no read position, what plays the role of lag?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

Queue depth — the count of records still waiting to be handed out — plus the age of the oldest waiting record. It is the same signal read without a position to subtract from, so it describes the whole queue rather than any one reader.

open as a page