What must a synthetic probe record share with real traffic before it can settle a dispute between cluster-side and reader-side delivery-time numbers?
answer
- inject a record to time it
- same host writes and observes
- share the real client path
- cover every part of the stream
- silence is itself a signal
basics
~20 sThe same client path, the same cluster, the same acceptance rules and coverage of every part of the stream — written and observed by one host so the timing needs no cross-host subtraction. Anything the probe does not share is a segment it silently excludes.
solid answer
~50 sA probe is a record injected on a schedule purely so its journey can be timed, and it is only evidence for the parts of the path it actually traverses. To settle the argument it must use the same endpoints, authentication and transport as real clients, be accepted under the same durability rules, be of comparable size, and cover every part of the stream — one probe on one partition says nothing about the nine it never touched. It should be written and observed by the same host, so the whole interval is one clock's arithmetic and no clock offset enters the number. The decision that shapes everything else is whether the probe rides its own stream and reader, which measures the path, or the real stream and reader group, which measures what users actually wait through. They answer different questions, and its silence — records stopping — is a signal in its own right.
go deeper
Recall what a probe is: a record written on a schedule only so someone can time how long it takes to be processed, and a stream that carries such records.
Explain why the same host should write and observe the probe, and why a probe that connects differently from real clients is only evidence about its own path.
Design it: per-part coverage, real acceptance rules, real client path, a cadence finer than what you need to see, and a conscious choice between a probe stream and the real stream plus the payload discipline that choice imposes.
Decide how far the probe estate goes. Measuring delivery honestly on every stream costs writers, readers, records and owners forever; measuring only where a commitment exists leaves the rest to cluster-side numbers, and that trade should be stated, not drifted into.
## What a probe is, and why it settles anything A **synthetic probe record** is a record written on a schedule for no purpose except to be timed. A probe writer emits it; a probe reader processes it; the elapsed time is the delivery interval for that journey. It settles a cluster-versus-reader argument for two reasons. First, both endpoints can be observed on **one host**, so the interval is a single clock's subtraction and carries none of the offset that corrupts a reader comparing a travelling timestamp against its own clock. Second, its arrival rate is known in advance, so an absence of probe results is itself informative in a way that a quiet real stream never is — a dead-man's-switch alert on the probe catches the case where the measuring path, not the broker, is what died. ## What the probe must share with real traffic 1. **The same cluster and the same client path** — the same endpoints, the same authentication, the same transport settings. A probe that connects differently proves only that its own privileged path works. 2. **The same acceptance rules** — accepted under the durability the real writers ask for. A probe that is acknowledged sooner than real writes is measuring a shorter path and will look healthy while real writes stall. 3. **Coverage of the whole stream** — a probe on one part of a stream says nothing about the others, and an unserved or slow part is exactly what you are trying to catch. Probes are normally written so that every part receives them. 4. **A comparable record size** — a tiny probe travels through cost structures that large records do not, so it will miss size-driven effects on the write and read path. 5. **A cadence finer than the interval you care about** — if the probe fires hourly you cannot see a five-minute stall, and the resolution of every conclusion is the probe period. 6. **A reader deployed the way real readers are** — the same shape of process in the same place, or the probe measures a network path nobody else uses. ## The choice that decides what the probe means | | Own probe stream and own reader | Injected into the real stream, read by the real reader group | |---|---|---| | Measures | the path's health, independent of any backlog | what a real record actually experiences, backlog included | | Sees a reader that is hours behind | no | yes | | Pollutes business data | no | yes — every real reader must recognise and skip it | | Survives the real reader being down | yes, and keeps reporting | no, it stops with everything else | | Typical use | is the platform delivering at all? | is this pipeline meeting its commitment? | Neither is the right answer on its own. An estate that cares about both usually runs a path probe everywhere and an in-stream probe on the handful of streams with a published commitment, and accepts the payload discipline that the second one imposes on every reader of that stream. ## What a probe cannot tell you - **Nothing about a reader group it does not go through.** A probe on its own stream and its own reader will read a healthy interval while the business reader is hours behind on its unread backlog. - **Nothing about parts of the stream it never touches**, which is why per-part coverage is not a refinement but a requirement. - **Nothing about effects that scale with volume or record size**, because probe traffic is by design negligible. - **Nothing about whether the work was useful.** The probe times the path; whether the real reader's processing has any effect is a different signal entirely. - **Nothing about a tenant whose credentials or quotas differ**, if the probe runs as a platform identity rather than as that tenant. ## Cost, and what it buys The continuous cost is real but modest: a small record per part per period, its copies, its retention, and one reader process per probed stream. The organisational cost is larger — someone owns the probe writer, the probe reader and the records they leave behind, forever. That is why estates usually measure delivery honestly on the streams that carry commitments and settle for cluster-side numbers elsewhere, and why the number of probed streams is a deliberate decision rather than an accident of who asked first. What the probe buys is an argument-ending number: a measurement that cannot be dismissed as the other team's dashboard, does not depend on two hosts agreeing about the time, and reports its own failure by falling silent.
- Should the probe be read by the real reader group rather than a dedicated one?If you want the number users feel, yes — it then queues behind the real unread backlog and includes real processing. The price is that probe records enter business data and every reader of that stream must recognise and skip them. Many estates run a path probe everywhere and an in-stream probe only where a commitment exists.
- How often should the probe run?Often enough that the interval you care about spans several probe periods, because the probe's period is the resolution of every conclusion drawn from it. Continuity matters as much as frequency: a probe that pauses leaves a gap indistinguishable from a healthy quiet period unless its absence is itself watched.
- The probe reports a healthy interval while users report stale data. What are the candidates?The probe is bypassing what is broken. It may ride a stream the affected readers do not, touch only some parts of the stream, connect on a path with different authentication or quotas, or be accepted under weaker rules than the real writers ask for. Each is a segment the probe excluded by construction.
saying these in an interview costs you the question
- Thinks a probe on its own stream reflects the real reader's unread backlog
- Probes one part of a stream and calls the stream measured
- Runs the probe on a client path no real client uses
- Treats an absence of probe results as good news
- Times the probe by subtracting two different hosts' clocks