One reader's own maximum record size is smaller than the node's — what happens to a record that lands between the two?
answer
- accepted and stored is not readable
- the smallest gate can be inside the client
- one application affected, cluster green
- stall or redelivery loop, not a write error
basics
~20 sThe record is accepted, stored and perfectly healthy from the cluster's point of view, but that one reader cannot take it. Depending on the design it stalls on that record or fails to complete it repeatedly, while every other reader is unaffected.
solid answer
~40 sThis is the failure that arrives long after the write. The write succeeded, the cluster is green, other consuming applications are fine, and one reader simply stops making progress — or, on a design that hands records out and deletes them on acknowledgement, keeps receiving the same record and never completing it. The symptom is a stalled or looping reader, not a size error, unless somebody reads the client's own error text. It is hard to attribute because the cause is client-side configuration that was correct until the day a record larger than it was finally produced, and because the cluster's own signals show nothing wrong. Designs differ: some deliberately return an oversize record anyway so a reader cannot wedge itself, others refuse and leave the reader retrying the same request.
go deeper
Recall that the reading side has its own maximum record size, so a record the cluster accepted and stored is not automatically one every reader can take.
Explain the mechanics: the reader's ceiling is client configuration, it binds only when a record larger than it finally exists, and the result is a reader that stops advancing or repeats work rather than an error on the write path.
Demonstrate the diagnosis under pressure: one consuming application affected while the cluster is green, evidence read from the client's own error rather than inferred, and awareness that the fix needs a release and that the stored records remain.
Treat the reading side's ceiling as part of the estate's agreed contract, raised ahead of any headroom the writing side is given, and published where the next consuming application will find it.
## The gate on the read side The last of the size gates lives in the consuming application. A reader asks the node for records and will accept a response, and a record inside it, only up to a ceiling of its own. That ceiling is client configuration: it is shipped with the application, it is owned by whichever team runs it, and the cluster operator usually cannot see it at all. So a record can be **accepted, replicated and durable** — everything the writing side and the cluster care about — and still be unreachable to a particular reader, because the smallest gate on that reader's path is inside the reader. ## Why it stalls rather than skipping The outcome depends on the shape of the platform, and this is where candidates over-generalise from whichever one they have run: - Where readers work through a durable stream **in order**, an oversize record sits directly in the way. The reader asks for the next records, cannot take what comes back, and asks again — the same request, the same answer. Nothing advances, and the reader looks hung rather than failing. - Some platforms specifically defend against this by **returning the oversize record anyway**, above the reader's stated ceiling, so a reader can never be permanently wedged by one large record. On those, the symptom is a memory spike or a surprised client rather than a stall. - Where records are handed to competing consumers and removed on acknowledgement, there is nothing to stall. The record is delivered, the consumer cannot process it, it is not completed, and it comes back — a **redelivery loop** on one record while the rest of the traffic drains normally. The common thread across all three is that the cluster is behaving exactly as configured. Nothing is broken; a gate is simply in the wrong place. ## Why nobody spots it quickly | Who is looking | What they see | |---|---| | The writing application | success; the record was accepted | | The cluster operator | a healthy node, records stored, no refused writes | | Other consuming applications | nothing at all — their ceilings are larger | | The affected team | a reader that has stopped, or keeps repeating work | Three properties make attribution slow. First, the **time gap**: the ceiling may have been in place for a year and only binds on the day the biggest record so far is produced. Second, the **blast radius is one application**, so the incident starts in a team that does not own the cluster and does not suspect a size ceiling. Third, the **symptom is the wrong shape** — a reader that is not advancing looks like a slow consumer, a dependency problem or a wedged process, and the size reason is usually sitting in the client's own error output where nobody has looked yet. ## What the operator can actually do 1. **Confirm it is size**, by reading the failing reader's own error text and comparing the size of the record it is stuck on with the reader's configured ceiling. Do not infer it from the shape of the stall. 2. **Check whether other readers of the same stream are fine.** A single affected consuming application is strong evidence for client-side configuration rather than anything on the cluster. 3. **Raise the reader's ceiling and release it.** This is the real cost: a cluster-side ceiling changes in an afternoon, while this one needs a build, a review and a deployment from a team that may be in another time zone. 4. **Decide what to do about the records already stored.** They are not going away, and they will be met again by any reader that still has the low ceiling. Waiting for retention to remove them is a choice, not a fix. 5. **Ask whether the writing side should have produced that record at all.** Raising a reader is the immediate fix; agreeing the largest record any producer may write is the durable one. ## Preventing it - Treat the reader's ceiling as part of the same agreement as the node's, and raise it **before** anything on the writing side is allowed to use new headroom. - Publish the agreed maximum where a team onboarding a new consuming application will read it, since a fresh application with stock configuration is the most common way a low gate reappears. - Watch the **tail** of the record-size distribution rather than the average, because this failure is a tail event by construction. - When the cluster's ceiling is raised, treat every existing reader as a hop that has not been raised yet, and confirm rather than assume. ## Where platforms differ Beyond the stall-or-return difference above, the reader's ceiling is sometimes expressed per record and sometimes per response, so the same configuration behaves differently depending on how many records come back at once. On designs where a reader can be moved back over stored records, an oversize record remains in the path on every replay until a ceiling changes; on designs that remove a record once it is acknowledged, the same record eventually leaves through the failure-handling path instead.
- Why do some platforms deliberately return a record larger than the reader asked for?To make it impossible for one large record to wedge a reader permanently. If the node refuses, an in-order reader retries the identical request forever and never advances; returning the record regardless converts a permanent stall into a one-off memory cost the client can survive. It trades a bounded surprise for an unbounded outage, which is usually the better trade.
- The affected reader is raised and restarted, but the records it choked on are still stored. What follows?That reader now gets through them, but any other consuming application still holding the old ceiling will meet the same records whenever it reaches them, including a new one starting from the beginning of the retained data. The records do not become smaller, so every remaining reader is a gate that still has to be raised or replaced before it reads that far.
saying these in an interview costs you the question
- Assumes a stored record is readable by every reader.
- Treats a stalled reader as always a slow-processing problem.
- Thinks a reader skips a record it cannot accept.
- Believes raising the node's ceiling reaches the readers too.
- Expects the cluster's own signals to show this failure.