skip to content

A reader group's record lag sits near zero all day while no business work completes — how can both be true at once?

level: middleimportance: must knowfreq 55%

answer

  1. lag is a reading signal
  2. the report comes from the reader
  3. stalled reader lags; this one does not
  4. compare inputs against completed outcomes

basics

~20 s

Near-zero reader record lag says records are being read and progress is being recorded — not that the work succeeded. A reader can move past every record while the processing throws, writes nowhere, or is rejected downstream.

solid answer

~50 s

Reader record lag compares where records have been written to with how far the reader has reported getting. It is a **reading** signal, not an **outcome** signal. Where a reader owns a stored read position, that position can advance while the work inside the reader fails silently; where a broker deletes a record on acknowledgement, the depth drains to zero for exactly the same reason. The cluster cannot tell a reader that processed a record from one that skipped past it, because both look like the same report. Usefully, this failure is distinguishable from a stalled reader on the very number in question: a stalled reader's lag **grows**, while a silently failing one keeps lag flat and low. To confirm it, compare the count of inputs over a window against a count of completed outcomes emitted by the application; the gap, not the lag, is the evidence.

go deeper

for a junior

Remember that lag measures reading, not doing. If someone says lag is zero so everything is fine, the missing question is whether anything useful actually happened.

for a middle

Explain the mechanism: progress is reported by the reader and taken at face value, so a swallowed error, a no-op branch or a rejecting destination all leave the number flat and low.

for a senior

Walk the diagnosis: compare the write-side rate, the reader's lag and an application outcome count, and rule out the stalled-reader case because its lag would be growing rather than flat.

for a principal

Decide whether every pipeline in the estate must publish an outcome count, or only the ones whose silent failure would cost something, and who pays for the instrumentation.

## What near-zero reader record lag asserts **Reader record lag** is the distance between the newest record written to a stream and the point the reader group has reported reaching — counted in records, or as the age of the oldest unread record. Near zero means: records are being handed to the reader as fast as they arrive, and the reader keeps reporting that it has got past them. Read that carefully, because the second half is where the trouble lives. The report is generated by the reader, about the reader. It is a claim of position, not a receipt for work done. The broker has no way to audit it, because the work happens in a process the broker does not run. ## The shapes that produce the contradiction Only a handful of situations put flat low lag next to zero outcomes, and they are worth being able to list: - **The work failed and the failure was swallowed.** Processing raised an error, something caught and discarded it, and the reader moved on as if it had succeeded. - **The work became a no-op.** A branch introduced by a deploy, or a flag flipped in configuration, means the reader reads everything and does nothing with it. - **The output goes nowhere.** A transform stage writes its results to a stream nobody reads, or to a destination that quietly discards them; the reading side keeps flowing. - **The destination rejects every write.** A downstream store refuses each write — bad credentials, a schema mismatch, a full disk of its own — and the reader records progress regardless. - **A filter inverted.** A predicate meant to drop a small fraction now drops everything, so records are read and almost nothing is emitted. In all five, the transport did its job perfectly. That is the point. ## The same illusion under two platform shapes The number you stare at depends on the platform's shape, but the illusion is identical: | Platform shape | The number that reads healthy | Why it reads healthy | |---|---|---| | The reader owns a stored read position | Reader record lag near zero | The position keeps advancing on the reader's own report | | The broker deletes on acknowledgement | Queue depth drained to zero | Records are removed as the reader reports being finished with them | A candidate who has only operated one of these will describe only one number. Naming both is what shows you understand the mechanism rather than a dashboard. ## Telling it apart from a reader that is merely behind This is the useful diagnostic contrast, and it runs on the same signal: 1. **A stalled or slow reader** — lag climbs steadily, or the oldest unread record gets older and older. The cluster reports this faithfully, and it is somebody else's subject. 2. **A silently failing reader** — lag stays flat and low, forever, while nothing downstream changes. The cluster has nothing to report. 3. **A reader that has died outright** — lag climbs and its progress reports stop. Also visible. So the pathological case is precisely the one where the signal looks best. An operator who only alerts on lag rising has built a detector that is blind to the worst case. ## What to measure instead 1. Emit a **count of completed outcomes** from the application at the point the outcome becomes durable — the row committed, the message sent, the file written. 2. **Reconcile** that count against the count of records read over the same window, allowing for records the pipeline is supposed to drop. 3. Alert on the outcome count **stopping**, not only on it being wrong; absence is the signature of this failure. 4. If you time the whole path with a **synthetic probe record**, route the probe through the processing step, not merely through the broker. A probe that is written and read back proves the transport works, which the lag number already told you. ## The trap in the interview The weak answer is "lag is zero, so the pipeline is fine — check the producer." The producer is usually the wrong end: if nothing were being written, lag would be zero because the stream is empty, and the write-side rate would say so. The strong answer names the class of failure, says why the transport cannot see it, and reaches for an application-published outcome count as the evidence.

  • Would a rising write-side rate with flat lag change your diagnosis?
    It sharpens it. Records are definitely arriving, and the reader is definitely keeping pace with them, so the stream is not empty and the reader is not stalled. That leaves the work itself, or what it writes to, as the failing part — exactly the region the cluster cannot observe.
  • How quickly would this failure be noticed without an outcome count?
    Usually when a human complains — a report that is empty, a customer whose order never shipped — which can be hours or days later. Every automated signal that exists reads normal, so nothing pages. That delay, not the failure itself, is what makes the case worth building a signal for.

A courier network that scans a parcel the moment it is handed to the recipient's doorstep. The tracking page is flawless and every parcel shows delivered on time — which tells you nothing about whether the box was empty, whether anyone opened it, or whether the goods inside were the ones ordered.

saying these in an interview costs you the question

  • Concludes the pipeline is healthy because reader lag is zero
  • Blames the producer whenever no outcomes appear
  • Thinks the broker confirms that processing succeeded
  • Expects a silently failing reader to raise lag like a stalled one
  • Suggests only watching lag more closely or more often