skip to content

questions

4

A reader group member still holds its share and makes no progress: what separates a stuck member from a merely slow one?

level: juniorimportance: must knowfreq 62%

answer

  1. alive is not the same as moving
  2. compare one share against its siblings
  3. does the recorded read position advance?
  4. same record again, on a cycle
  5. capacity raises a rate, not zero

basics

~10 s

A slow member still completes records, just fewer than arrive, so its read position keeps advancing. A stuck member completes nothing: its position is parked. Only the slow case is answered with capacity.

solid answer

~50 s

Both look the same on a chart — a share of the work that is not catching up — but the rate is different in kind. A merely slow member has a positive completion rate that is below the arrival rate, so its recorded read position keeps advancing and the share would drain if writes stopped. A stuck member has a rate of zero: it holds its assignment and finishes nothing, so the position parks. The separating question is simply *does anything complete over a fixed observation window*, not *how big is the gap*. Liveness is not the test, because a stuck member normally keeps answering its liveness signal — a handler waiting on a dependency or failing on one record is still a live process. The distinction matters because adding capacity raises a low rate and does nothing at all to a rate of zero.

go deeper

for a junior

Recall the two states and the one question that separates them: a slow member finishes records too slowly, a stuck member finishes none. Check whether the recorded read position moves at all.

for a middle

Explain why the broker does not raise the alarm: a frozen member normally keeps answering its liveness signal, so presence and progress are different facts. Mention the progress deadline and the redelivery cycle it produces.

for a senior

Show the diagnosis order under pressure — compare the flat share against its siblings, look for one record reappearing on a cycle, test whether removing one record restores flow — and say plainly that capacity is the wrong lever for a rate of zero.

for a principal

Frame it as an observability and policy question: per-share progress must be observable, because an aggregate hides a single frozen share until its oldest unread record approaches the edge of the retained window.

## Two different failures that draw the same line A **reader group** is the set of readers sharing one division of work over a stream, and each member holds an **assignment** — its share of that work. When one member's share stops draining, two broad things can be true of it, and they take opposite remedies. - **Merely slow.** The member completes records, just fewer per second than arrive. Its recorded **read position** keeps advancing. The backlog of unread records on that share grows because arrivals outrun it, not because nothing is happening. If writes stopped, the share would eventually drain. - **Stuck.** The member holds the share and completes nothing. The position is parked at one point. If writes stopped, the share would still not drain, because the completion rate is zero rather than low. Both produce the same first symptom, and the instinctive response — add capacity — helps exactly one of them. ## Why nothing reports a stuck member as failed A broker notices a member that has **gone**: its **liveness signal** stops arriving and its share is handed to another member. A stuck member normally keeps answering that signal. A handler waiting on a downstream call that never returns, a thread parked on a lock, an exhausted worker pool, a record that fails on every attempt — none of these stop the periodic proof that the process exists. Where the design also enforces a **progress deadline** — the timer that decides a member has stopped because it is holding work without showing progress — the work is eventually taken back. That is usually not a recovery: the same record goes to the next member, which meets the same fault. What the operator sees is a cycle, and the cycle's period is a clue, because it is the timer rather than anything about the workload. ## Discriminators you can apply in minutes 1. **Does anything complete at all?** Over one fixed observation window, does the member's recorded read position advance by any amount? Any advance means slow. Zero, with unread records still waiting, means stuck. 2. **How do siblings behave?** Other members of the same reader group, holding comparable shares of the same stream, are the control. Every share flat points at the whole group or at something upstream of it; one flat share among draining ones points at that member or at what it is holding. 3. **Is one record reappearing?** A second copy of the same record delivered again, on a regular cycle, with no completion in between, is the signature of work being handed around rather than done. 4. **Does flow resume when one record leaves?** Divert or discard exactly one record and watch. Instant resumption identifies both the failure and its scope: one record was holding the share. 5. **Is the record even the cause?** A member can be stuck with nothing wrong in the stream — a dependency that never answers, a lock nobody releases, a process suspended by the platform beneath it. Check the handler's own state before blaming the data. ## A flat share can also be innocent - **Nothing is arriving.** A share with no unread records and a position at the newest record is idle, not stuck. Always confirm that unread records are actually waiting. - **The reader was stopped deliberately.** A deployment, a scale-to-zero, a paused process during maintenance: the share is not moving because nobody is holding it. - **The share is empty by routing.** Where work is divided by key, a share can legitimately receive almost nothing while its siblings are busy. That is skew, not a stall. ## What the platform's shape changes | the question | where a reader owns a recorded read position | where work is pulled from a shared queue and removed on acknowledgement | |---|---|---| | Is it moving? | the share's read position advances | the count of waiting records falls | | What marks a blocked record? | the position parks at one point | the same record returns after each holding window | | Can it be attributed to one member? | yes — shares are held one member each | less directly; members compete, so attribute to the record | | What does adding members do? | nothing beyond the parallelism ceiling | adds competitors, which helps only the slow case | The underlying distinction survives both designs: a positive rate below demand, against no rate at all. Only the instrument changes. ## Why the label decides the bill - Capacity — more members, larger batches, a faster handler — raises a completion rate. It is the remedy for slow, and it has nothing to raise when the rate is zero. - On a stuck member, capacity can even look like it worked: the other shares drain faster, the group's aggregate gap bends downwards, and the frozen share is exactly where it was. - A stuck member has a bounded blast radius — its own share — so the aggregate picture understates it. The smaller that share, the longer the freeze hides. - A slow member is a planning problem that can usually wait for the next capacity review. A stuck member is an incident now, because whatever it is holding is ageing towards the edge of the retained window.

  • If a stuck member keeps answering its liveness signal, what does eventually make the broker act?
    Designs that enforce a progress deadline treat holding work without showing progress as failure and take the share back, independently of liveness. That produces a cycle rather than a recovery, because the next member usually meets the same fault. Designs without such a timer will leave a live-but-frozen member holding its share indefinitely, which is why the completion rate, not liveness, has to be the signal an operator watches.
  • Can one member be stuck while the reader group's aggregate gap looks acceptable?
    Routinely. A frozen share contributes only its own unread records, so with many shares the aggregate barely bends while one share's oldest unread record ages without limit. This is why per-share observation exists at all: an aggregate hides a bounded failure, and the bound is what makes it survivable long enough to become a data-loss problem at the edge of the retained window.
  • Does restarting the stuck reader tell you anything useful?
    It is a cheap test with a clear reading. If the member resumes and runs normally, the fault was process-local — a leaked lock, an exhausted pool, a hung connection. If it resumes, reaches the same point and freezes again, the fault travels with the work rather than the process, and restarting is not a remedy. Either way, record which happened before doing anything irreversible.

saying these in an interview costs you the question

  • Says a stuck member is detected automatically because its liveness signal stops
  • Calls any flat share stuck without checking that unread records are waiting
  • Adds members to a reader group and expects a frozen share to move
  • Assumes stuck always means a bad record, never a blocked dependency
  • Reads only the group total, never one member's share against its siblings
  • Treats a bigger backlog as proof of being stuck rather than behind
open as a page

A reader is frozen on one record it cannot handle: what are the exits, and what does each one cost in correctness?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Four exits: fix the handler, set the record aside, discard it, or wait for a transient cause to clear. Fixing loses nothing but is slowest; discarding is fastest and loses the record silently. Adding capacity is not an exit.

open as a page

On a reader's assigned share, does one record that cannot be handled block the records behind it or only itself?

level: middleimportance: should knowfreq 55%

basics

~20 s

It depends on the unit of recorded progress. Where a share's progress is one advancing read position, the record blocks everything behind it on that share. Where each record is acknowledged on its own, it blocks only itself.

open as a page

Across an estate of streams, what should a standing policy say about who may discard a record to unblock a frozen reader, and what must be captured first?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Settle it before the incident: classify each stream by tolerable loss, pre-authorise the cheap exits, require the record to be captured durably before any discard, name an owner for anything set aside, and set a deadline before the retained window decides for you.

open as a page