skip to content

Across an estate of hundreds of streams, which broker alert conditions should a platform team own centrally, which must each stream owner set, and what happens to a stream whose owner sets none?

level: principalimportance: nice to knowfreq 40%

answer

  1. whose fact is the number
  2. cluster facts against workload facts
  3. one nightly reader defeats one number
  4. platform supplies shape, owner supplies threshold
  5. unset means file, never page

basics

~20 s

Conditions on the cluster's own health are properties of the cluster, so one number serves the estate and the platform team owns it. Conditions on flow are properties of a stream's commitment, so the owner sets the number; an unset stream gets a non-paging default.

solid answer

~50 s

Split the conditions by whose fact the number is. A condition on the cluster's own health has a correct value that depends on the cluster and not on any tenant, so the platform team writes it once for the estate and keeps it. A condition on flow - how far behind a reader may fall, how long a record may take to be handled - has no correct value that anybody but the stream's owner can know, because the same reader lag is routine for a nightly batch reader and unacceptable for a payments reader. The platform's job there is to publish the signal per stream and per reader group, to supply the condition's shape, and to require a number. A stream with no number gets a deliberate default that files rather than pages, plus a record that it is unmonitored - never a guessed page, which trains the estate to ignore the pager.

go deeper

for a junior

The idea to take away is that being behind is not automatically bad - it depends entirely on what the stream feeds. A nightly report and a payment flow cannot share one number.

for a middle

Be able to explain why a cluster-health threshold generalises across an estate while a lag threshold does not: one is a fact about the cluster, the other a fact about a commitment the cluster knows nothing about.

for a senior

Show what the platform has to provide before it can demand a number - the per-stream signal, the history, the condition template - and what a stream with no number should actually do rather than page on a guess.

for a principal

This is the contract you own. Decide where the line sits, what the unset default costs, and how thresholds are reviewed when a stream changes hands - and be honest about which signals a rented cluster will not give you.

## The question behind the question: whose fact is the number? Ownership arguments about alerting usually get stuck on who is on call. The productive cut is different: for each condition, ask whether the correct threshold is derivable from the cluster or only from the workload. | Condition class | Correct value depends on | Owner | If the wrong party owns it | |---|---|---|---| | Cluster health - shortfalls, unserved partitions, coordination-plane state | The cluster's own configuration | Platform team | Hundreds of teams each guessing at a fact that has one answer | | Node capacity - free space, handle headroom, request wait time | The hardware and the node's limits | Platform team | Nobody notices the volume filling until writes stop | | Flow - reader lag, delivery interval, backlog age | The stream's downstream commitment | Stream owner | The platform pages on a number it has no basis for choosing | | Estate hygiene - streams with no reader, no owner, no threshold | The governance rules | Platform team | Orphans accumulate silently | The first two classes are one number each, written once. The third cannot be, and that is the whole argument. ## Why a single estate-wide flow threshold is wrong Take the most common proposal: one reader-lag threshold for every reader group in the estate. - A nightly batch reader is *supposed* to be millions of records behind for six hours. The estate threshold either fires on it every night or is set so high that nothing else in the estate can ever reach it. - A payments reader is in trouble at a few seconds of staleness. The same number that leaves the batch reader alone leaves this one unmonitored entirely. - A stream whose traffic varies tenfold between business hours and the weekend has two normals of its own, so even one stream can defeat a single number. The number is not a technical parameter with an unknown value. It is an expression of a commitment - the report that must be ready by half past six, the payment that must be visible within a few seconds - and only the owner of that commitment holds it. A platform team that sets it anyway is guessing on behalf of a business it cannot see. ## What the platform owes the owner Requiring a number is only fair if setting one is possible. The platform side of the contract has three parts: 1. **Publish the signal, per stream and per reader group**, with a documented meaning and enough history that an owner can see the stream's own normal before choosing a line. Without per-stream history, every owner is inventing a number. 2. **Supply the condition's shape, not just the signal.** The owner supplies the threshold and the deadline; the platform supplies the template that already contains a sensible sustained window, a stance on missing data, and a floor under any trend clause. Letting every team invent the shape as well as the number is how an estate ends up with four hundred subtly broken rules. 3. **Record the justification with the threshold.** A number whose reason is lost cannot be raised, lowered or defended, and it will outlive the service it was chosen for. ## The default for a stream nobody has set a number for There will always be streams whose owner never supplies one. The tempting answer - apply an estate default and page on it - is the worst available, because it generates pages against a number nobody chose, for a team that has not agreed to receive them, and it teaches the estate that pages are noise. A better default has three parts: a condition that files rather than pages, so the stream is not entirely dark; an explicit record that the stream is unmonitored for flow, so nobody believes otherwise during an incident; and the stream's appearance on a list the platform team reviews. Absence of a threshold is a governance finding, not an alerting one, and treating it as such keeps it visible instead of hiding it behind a fake page. ## Keeping it from drifting Thresholds rot in predictable ways: traffic grows past a record-count number, a commitment moves and the deadline does not, a service changes hands and its rules do not. Two cheap habits help. Prefer age-based flow thresholds to count-based ones, because age survives throughput growth and expresses the commitment directly. And review flow thresholds at the moments the estate already notices - when a stream changes owner, and when its traffic shape changes materially. ## What varies across platforms Which signals the platform can publish per stream differs sharply. Some expose lag per reader group and per partition; some expose only a coarse per-queue depth; a rented cluster may publish a fixed set of numbers and nothing else, in which case the flow condition has to be measured by the application rather than by the platform - and the contract has to say so honestly rather than promising a signal that does not exist.

  • An owner asks the platform team to just pick a reader-lag threshold for them. What do you give them?
    The shape, the history and a question - not a number. Show them the stream's own normal over a few weeks and ask what downstream commitment the records feed; the threshold falls out of the answer. If there is genuinely no commitment, that is the finding: the stream may not need a paging condition at all.
  • Why prefer an age-based flow threshold to one expressed in records?
    Age survives growth. A count chosen at today's throughput quietly becomes a different amount of delay as traffic rises, and nobody revisits it. Age also states the thing the business cares about - how stale the oldest unhandled record is - so the justification stays legible to whoever inherits it.
  • Does this split mean the platform team never alerts on a tenant's stream?
    No. The platform keeps estate-hygiene conditions that are about the cluster's wellbeing rather than the tenant's outcome - a stream growing without any reader at all, a stream consuming disproportionate capacity. Those are the platform's facts. What it does not own is how late that tenant's records are allowed to be.

saying these in an interview costs you the question

  • Sets one estate-wide reader-lag threshold and calls it a standard
  • Lets each team invent the condition shape as well as the number
  • Pages on a default threshold that nobody deliberately chose
  • Assumes the platform team can know a stream's business deadline
  • Leaves an ownerless stream paging the platform team indefinitely
  • Keeps count-based thresholds that quietly shrink as traffic grows