Across an estate of many streams, which signal honestly reports that a pipeline is working, who must publish it, and what does requiring it everywhere cost?
answer
- cluster signals are a shared service
- completion is a domain definition
- the platform makes it cheap, not itself
- tier by what silence costs
basics
~20 sThe honest signal is a count of completed business outcomes, published by the team that owns the pipeline, because only that team can define completion. Requiring it estate-wide costs per-pipeline instrumentation, extra stored series, and a definition argument per pipeline.
solid answer
~50 sSplit the contract. The platform team owns and guarantees the cluster's signals — copies behind, unserved partitions, free space, request wait and service time, reader progress — as a shared service for every stream. Nobody but the owning team can publish the signal that matters for the work, because "completed" is a business definition: an order fulfilled, an index updated, a settlement written. So the estate rule is that each production pipeline publishes at least one outcome counter and something that notices when it stops. The costs are real: instrumentation work in every pipeline, more stored series, and an argument per pipeline about what counts as an outcome when records are deliberately filtered or partially processed. Where that is too expensive, tier it — mandatory for pipelines whose silent failure is expensive, and a cheaper periodic reconciliation for the rest.
go deeper
Take away the division of labour: the platform publishes the cluster's numbers, and the team that built the pipeline publishes whether its work is getting done.
Be able to describe the outcome counter concretely — where it is emitted, what it counts, and why an alert on it stopping matters more than an alert on its value.
Argue the reconciliation honestly: a window, a tolerance, and an allowance for records the pipeline drops on purpose, rather than an exact match that will page every night.
Own the trade: which tier each pipeline belongs to, what time to detection you are buying, what the series and instrumentation cost, and where the rule is enforced so it is not optional.
## The signal One signal tells you a pipeline is doing its job: **a count of completed business outcomes**, emitted where the outcome becomes durable. Everything the cluster publishes is upstream of that — necessary, insufficient, and routinely mistaken for it. A usable outcome signal has three properties: - it is **counted at the point of durability**, not when work was started or attempted; - it is **comparable to an input count** over the same window, so a collapse in the ratio is visible even when the count is not zero; - its **absence is itself an alert**, because the signature of this failure class is silence, not a bad value. ## Who can publish it Not the platform team, and this is the part that gets argued. The platform team runs the clusters and can expose every transport signal for every stream from one place. It cannot define completion, because completion is domain knowledge: the same record is a fulfilled order in one pipeline and a discarded duplicate in another. A platform that tried to derive outcomes from reader progress would just be republishing the signal that already lies. So the division is: | Question | Signal | Owner | |---|---|---| | Is the cluster serving? | copies behind, unserved partitions, free space, handler-pool saturation | platform team | | Are records moving? | reader progress, delivery interval on probed streams | platform team | | Did useful work happen? | completed-outcome count, outcomes against inputs | the pipeline's owning team | The platform's real leverage is not to produce the outcome signal but to **make it cheap and standard**: one agreed way to emit the counter, one agreed dashboard shape, one agreed absence alert, so a team adds four lines rather than designing a monitoring strategy. ## What mandating it costs 1. **Instrumentation work per pipeline.** Small individually, real across a hundred pipelines, and it lands on teams who believe their green dashboard already covers them — so the first cost is persuasion. 2. **Stored series.** Every pipeline adds counters, and if they are broken down per stream, per tenant or per outcome type, the series count multiplies. Metric storage is cheap per series and expensive per thousand. 3. **The definition argument.** What counts as an outcome when a record is deliberately filtered? When one input yields three outputs? When a partial success is retried tomorrow? Each pipeline answers this once, and each answer takes a conversation. 4. **Reconciliation tolerance.** An exact input-equals-outcome rule is unusable, because filtering, batching across window edges and retries all break it. Every comparison needs a window and a tolerance, which is a per-pipeline judgment rather than an estate-wide constant. ## Where to spend less Universal coverage is rarely the right call, and saying so is part of the answer. Tier it by what silence costs: - **Pipelines whose silent failure is expensive** — money movement, customer-visible fulfilment, regulatory reporting — get an outcome counter, an absence alert and a whole-path probe that is processed by the real processing step. - **Pipelines whose silent failure is merely embarrassing** get an outcome counter and an absence alert, with reconciliation reviewed rather than alerted. - **Everything else** gets a periodic reconciliation — a scheduled comparison of inputs against outcomes — which is far cheaper than continuous measurement and still bounds how long a silent failure can run. The number worth naming in the decision is not the cost of the metric but the **time to detection**: with cluster signals alone it is however long until a human complains, which is typically hours to days. Each tier buys that number down, and the question a principal is really answering is how much the organisation will pay to shorten it, on which streams. ## Making it stick A policy nobody enforces produces documentation, not signals. The enforceable version attaches to something that already happens: a pipeline is not accepted into production, or a stream is not provisioned for it, until the outcome counter and its absence alert exist. That also fixes ownership at the moment somebody is paying attention, which is the same moment the stream gets a name and an owner — and an unowned pipeline with no outcome signal is precisely the one that will fail silently for a week.
- Why not have the platform derive outcomes from reader progress for every stream?Because reader progress is the signal that lies in exactly this failure class. Deriving from it would produce a metric that looks like an outcome count and reads healthy while a reader advances past records doing nothing. A derived signal cannot be more truthful than its input.
- What is the cheapest thing a team can do if a full outcome counter is out of reach?A scheduled reconciliation: periodically compare records read against rows, files or messages produced at the far end, and review the ratio. It has poor time to detection compared with a live counter, but it bounds how long a silent failure can run, and it needs no change to the running pipeline.
- How do you stop the estate filling with outcome metrics nobody reads?Tie each one to an alert on its absence and to a named owner at provisioning time. A counter with no alert and no owner is storage cost with no detection value; if nobody will be woken or ticketed when it stops, do not create it.
saying these in an interview costs you the question
- Expects the platform team to define completion for every pipeline
- Derives an outcome signal from reader progress and trusts it
- Demands exact input-equals-outcome matching estate-wide
- Mandates the same instrumentation for every pipeline regardless of stakes
- Counts work when it starts rather than when it becomes durable