skip to content

Which pipeline failures can no signal a broker cluster publishes ever reveal, and what makes the cluster blind to them?

level: seniorimportance: should knowfreq 48%

answer

  1. draw the line at hand-off
  2. transport observes bytes, not consequences
  3. output can go nowhere and look alive
  4. the broker would have to know the purpose

basics

~20 s

Anything that happens after hand-off is invisible: a reader recording progress without doing the work, a transform writing its output nowhere, a destination rejecting every write, a filter dropping everything. The cluster observes transport, not consequences.

solid answer

~50 s

A broker's signals describe bytes accepted, bytes stored, requests served, and the progress a reader reports. That set is complete for the transport and empty for the work. So four failure shapes leave every number normal: a reader that advances past records while the processing failed and the error was swallowed; a transform stage whose output is written where nothing reads it; a downstream store rejecting every write while the reader keeps moving; and a filter that drops everything after a bad deploy. The blindness is **structural**, not a missing metric — the failing step runs in a process the broker does not execute and writes to systems it does not know about. No amount of extra broker instrumentation reaches it. What does reach it is a count of completed outcomes published by the application, reconciled against inputs, with an alert on the count stopping.

go deeper

for a junior

Hold on to the boundary: the broker sees records arrive, be stored and be handed over, and nothing after that. Failures past the hand-off are somebody else's to report.

for a middle

Be able to name two or three concrete shapes — swallowed errors, output written where nothing reads it, a destination rejecting writes — and say why each leaves the transport numbers normal.

for a senior

Show the diagnosis and the fix together: identify which shape you are in from input rate, reader progress and destination errors, then add an outcome count and a reconciliation so it cannot recur silently.

for a principal

Argue the boundary explicitly in a review, so the organisation stops buying broker monitoring in the belief it covers pipelines, and decides where outcome instrumentation is mandatory.

## Where the cluster's knowledge stops A broker cluster is a transport with storage. Its observable surface is exactly that: records accepted, records retained, records served, and progress reported by readers. Draw a line at the moment a record leaves the broker for a reader — everything on the far side of that line is executed by somebody else's code, in somebody else's process, against systems the broker has never heard of. That line is the whole answer. The failures below are not hard to see; they are **impossible** to see from where the broker stands. ## The four shapes - **Progress without work.** The reader reads a record, the processing throws, something catches and discards the error, and the reader reports being past it. Where a reader owns a stored read position, the position advances; where a broker deletes on acknowledgement, the record is removed. Both look like success. - **Output into the void.** A transform stage — whether it is a library running inside the reader or a separate hosted mover — reads everything and writes its results to a stream nobody reads, or to a destination that silently discards them. Every hop looks alive; the chain is severed at the far end. - **A rejecting destination.** The store the pipeline writes to refuses every write: wrong credentials after a rotation, a column the incoming shape no longer fits, its own volume full. The rejection is on the far side of the broker's boundary and never appears in a broker number. - **A filter that swallowed everything.** A predicate meant to drop a small share now drops all of it after a deploy. Records are read at full rate; almost nothing is emitted. A fifth, quieter case belongs with them: **the right volume of wrong content**. Records keep flowing at the usual rate carrying values that are subtly incorrect. Byte counts cannot distinguish good data from bad, so nothing on either side of the boundary reacts. | Failure shape | What the cluster shows | What actually catches it | |---|---|---| | Progress without work | Reader progress advancing, lag flat and low | Outcome count flat while inputs flow | | Output into the void | Reads and writes both healthy | Reconciling outcomes against inputs at the far end | | Rejecting destination | Nothing unusual at all | The destination's own error rate, or an outcome count | | Filter dropping everything | Reads at full rate | Emitted-record count against read-record count | | Wrong content | Normal throughput | Content-level assertions, not a transport signal | ## Why more broker metrics will not close the gap It is tempting to treat this as an instrumentation backlog — surely one more metric would catch it. It would not, for a simple reason: the broker would have to know what the pipeline is **for**. "Was this record turned into a fulfilled order" is a question with no transport-level answer; the same bytes are a fulfilled order in one system and a discarded duplicate in another. A transport that tried to answer it would have to execute the business logic itself, at which point it is no longer a transport. This is also why the boundary is worth stating plainly in an incident review. The answer to "why did our monitoring not catch this" is not "we had the wrong alert"; it is "we were watching a system that structurally cannot see this class of failure." ## What does catch them 1. **A count of completed outcomes**, emitted by the application at the point the outcome becomes durable. This is the one signal that means what everyone assumed the green dashboard meant. 2. **A reconciliation** of inputs against outcomes over a window, with an allowance for records the pipeline is meant to drop. It catches the filter case and the void case, which a bare outcome count can miss if the pipeline is merely *less* productive rather than silent. 3. **An alert on absence** — a dead-man's-switch style condition that fires when the outcome count stops arriving, because this failure's signature is silence rather than a bad value. 4. **A whole-path measurement** — a synthetic probe record timed from write to finished work — but only if the probe is processed by the real processing step. A probe that is written and read back measures the transport, which was never the part that failed. ## What this is not A single record that cannot be processed and is routed aside has its own well-known handling and its own owner; so does a reader that has genuinely stopped making progress, which the cluster reports honestly as growing lag. The class in this leaf is narrower and nastier: the transport is working, the reader is working as far as the transport can tell, and nothing useful is coming out the other end.

  • Why is a rejecting destination invisible even though writes are failing loudly somewhere?
    They fail in the destination, which is a different system with its own signals. The reader keeps reading and reporting progress, so the broker's view is unchanged. Unless the destination's error rate is collected and watched, or the pipeline publishes an outcome count, the loud failure has no audience.
  • Which of these shapes would a whole-path probe actually catch?
    Only the ones whose path it shares. A probe processed by the real processing step and written to the real destination catches progress-without-work and a rejecting destination. A probe that skips the processing step, or is filtered out before it, catches none of them and produces false confidence.
  • Does a reconciliation need to be exact?
    No, and demanding exactness makes it unusable. Pipelines drop records on purpose, batch across window edges, and retry. Compare over a window with a tolerance and look at the trend; a ratio that collapses from its usual value is the signal, not a mismatch of a few records.

saying these in an interview costs you the question

  • Believes another broker metric would have caught it
  • Assumes a rejecting destination shows up as broker errors
  • Thinks flowing records prove the chain is intact
  • Treats correct throughput as proof of correct content
  • Confuses this with a single unprocessable record being routed aside