skip to content

Broker & streaming operations

10 roadmaps259 questionsupdated

Running a broker or streaming platform in production: how many copies a write needs, how long records live, what lag is telling you, what breaks on upgrade. Asked because operating one is its own job.

on this pageshow

guide

overview

~2 min

Operating a message broker or streaming platform is a different job from building applications on one. Application questions ask what a queue, a stream or a delivery guarantee means; operations questions ask what happens to the records when a node dies, a disk fills or an upgrade is half finished. Interviewers probe it because most messaging incidents are not platform defects but a durability setting nobody revisited, a history window that quietly shrank, a reader that stopped moving while every node dashboard stayed green, or a planned change that removed one copy too many. The subject splits into twelve sections. [Copy sets & durability](/topics/cloud-messaging-concepts-durability) and [retention & storage pressure](/topics/cloud-messaging-concepts-retention) decide what the cluster keeps and for how long. [Cluster shape & capacity](/topics/cloud-messaging-concepts-topology) fixes node roles, partition counts, disks and failure domains before traffic arrives; [membership & data movement](/topics/cloud-messaging-concepts-membership) and [upgrades & configuration change](/topics/cloud-messaging-concepts-lifecycle) cover changing that shape while it keeps serving. [Quotas, throttling & fairness](/topics/cloud-messaging-concepts-flow-control), [estate layout & governance](/topics/cloud-messaging-concepts-estate) and [broker security controls](/topics/cloud-messaging-concepts-security-ops) are about sharing clusters across many teams without one of them spending everyone's headroom or access. The reading side has its own section, [consumer operations & lag](/topics/cloud-messaging-concepts-subscriptions), and [operating signals & alerting](/topics/cloud-messaging-concepts-signals) decides which readings deserve a page. Two sections look beyond a single cluster: [cross-cluster continuity](/topics/cloud-messaging-concepts-continuity) for losing one entirely, and [renting a broker](/topics/cloud-messaging-concepts-managed) for what a hosted tier takes off your hands and what it leaves. Start with durability, because every later section assumes you can say how many copies a write needs and what an acknowledgement really promises. Topology and retention come next, since they set the budgets; lag and signals show whether those budgets hold; the change-management sections come last, where earlier ideas are tested while the cluster is in motion. Questions range from a junior asked what a writer can wait for to a principal asked what a failover plan cannot give back.

primer

A small set of ideas carries the whole subject. With them in place, most questions below become applications rather than facts to recall. - **An acknowledgement is a promise with a scope.** What "stored" means depends on how many copies had to accept the write and whether any of them had forced it to persistent media. Copies protect against losing hardware; an operator mistake or a malformed record is replicated just as faithfully as a good one. - **Copies are only as independent as the things they share.** A copy count is a claim about independent failures. Copies placed on one host, one rack or one zone collapse into a single copy against that failure class, so placement matters as much as the number. - **Retention is a replay budget.** Age limits, byte ceilings, latest-value-per-key rules and legal holds together bound how long a record stays readable. Every lagging or rebuilding reader depends on that span, and a traffic surge can shrink it silently. Watch the span you actually have, not the setting you wrote. - **Lag is a question about time.** A backlog matters by how stale the oldest unread record is and by whether reading outpaces arrival. Many incidents reported against the broker turn out to be the reading side stuck, slow or rebalancing. - **Shape is expensive to change.** Node roles, partition counts, disk layout and failure-domain spread are chosen before traffic arrives. Some of those choices only move one way, and moving data between nodes means copying retained records, which costs time and bandwidth. - **Most planned changes pass through a degraded window.** Rolling restarts, node additions, upgrades and credential swaps run for a while with fewer current copies, mixed versions or mixed trust. Operating well means keeping that window short, visible and reversible until the step that cannot be undone, and knowing which step that is. - **A shared cluster fails as one.** Named spaces divide names, grants and quota attachment points; they do not divide disks, network, the upgrade calendar or the on-call rota. Quotas and limits exist so one tenant cannot consume the headroom of all the others. - **A healthy cluster is not a working pipeline.** Node-level readings can be green while nothing is delivered. End-to-end delivery time and reader progress are the readings that catch it.

Replication factor
The number of copies of each partition or queue a cluster keeps on separate nodes; it sets how many node failures a unit can absorb, not which kinds.
Leader copy
The copy of a partition or queue that accepts writes and that the other copies follow; losing it forces a promotion.
Caught-up copy set
The copies currently close enough to the leader to count toward an acknowledgement and to be promoted without losing confirmed writes.
Acknowledgement threshold
How many copies must hold a write before the writer is told it succeeded; each step up trades latency for one fewer way to lose the record.
Partition
An ordered slice of a stream on log-based platforms: the unit of ordering, of placement across nodes and of reader parallelism.
Retention window
The span of history a stream keeps readable, bounded by age, by bytes, by a per-key rule or by a legal requirement.
Log compaction
Keyed retention that keeps at least the newest record for every key and removes superseded ones at no promised time.
Committed read position
The stored marker of how far a reader group has processed a stream, from which it resumes after a restart or reassignment.
Consumer lag
The gap between the newest record in a stream and a reader group's committed position, stated as a record count or as the age of the oldest unread record.
Rebalance
Reassigning a stream's shares among a reader group's members when one joins, leaves or stops proving it is alive.
Failure domain
A set of components that fail together, such as a host, a rack, a power feed or a zone; copies should not share one.
Fencing
Stopping a deposed leader from accepting writes after a new one is chosen, usually by rejecting requests carrying an outdated epoch number.
Client quota
A ceiling on bytes per second or requests per second that a cluster grants one authenticated identity, enforced by slowing or refusing it.
Backpressure
Any way a saturated node pushes cost back to writers: queuing, delaying answers, blocking or refusing, each visible differently to the client.
Cross-cluster mirroring
Asynchronous copying of streams from one cluster to another by a process that reads the source and writes the target, so the target trails.
Recovery point objective
The most data, measured as time, a stream may lose when its cluster is lost; an asynchronous copy's lag sets its floor.

The sections read best as layers: what a cluster promises, how it is built to keep that promise, and how the promise survives change, sharing and loss. **Durability and retention define the promise.** [Copy sets & durability](/topics/cloud-messaging-concepts-durability) says which failures an acknowledged record survives; [retention](/topics/cloud-messaging-concepts-retention) says how long it stays readable. Nearly every other section is a threat to one of those two statements. A disk filling up threatens both at once, which is why storage pressure sits inside retention. **Topology is the promise made physical.** Node roles, partition counts, disk profile and failure-domain spread decide whether the configured copy count means anything and how much throughput the cluster can take. Quotas, throttling and record-size ceilings then share that capacity out, and estate governance decides whether one cluster carries everyone or several carry a few teams each. **The reading side spends the retention budget.** A reader group's lag, its rebalances and its stored positions only make sense against the retention window: a reader behind by more than the window has already lost records, and a reset of position is a choice about which ones to reprocess or skip. **Change management reuses all of it under motion.** [Membership & data movement](/topics/cloud-messaging-concepts-membership) and [upgrades & configuration change](/topics/cloud-messaging-concepts-lifecycle) are the durability questions asked while copies are being moved, restarted or run on mixed releases. Credential rotation in the security section follows a similar staged sequence. **Signals tell you which layer is failing.** Node health, copy-set shrinkage, tail latency, reader lag and end-to-end delivery time each point at a different section; the skill interviewers test is reading them in combination rather than one at a time. **Continuity and renting move the boundary.** Cross-cluster continuity extends the durability promise past a single cluster at the price of an asynchronous gap. A hosted tier takes over the machines, but choosing retention, designing streams, granting access, picking client releases and running readers remain with the team that owns the traffic.

  1. Copy Sets & Durability →

    Every later section assumes you can say how many copies a write needs and what an acknowledgement actually guarantees.

  2. Cluster Shape & Capacity →

    Node roles, partition counts and failure-domain placement decide whether the durability settings hold in practice, and they are hard to change later.

  3. Retention & Storage Pressure →

    The history window is the budget every lagging reader and every recovery spends; learn what shrinks it.

  4. Consumer Operations & Lag →

    Most incidents surface on the reading side first, so learn to tell a stuck reader from a slow one and a backlog from staleness.

  5. Operating Signals & Alerting →

    Ties the earlier sections to the readings that show them failing, and to which of those deserve to wake someone.

  6. Membership & Data Movement →

    Planned node changes are where durability is lost most often; this is the first section where the cluster is in motion.

  • Treating an acknowledgement as proof the record is on disk; on many setups it means accepted into memory on enough copies, and the loss window depends on that difference.

  • Quoting a copy count without saying where the copies live; three copies in one rack or one zone are one copy against that failure.

  • Reporting lag only as a record count; without the age of the oldest unread record nobody can tell whether a reader is seconds or hours behind.

  • Assuming retention is fixed at the configured age when a byte ceiling also applies; a surge can cut the readable history to hours with no error.

  • Adding reader instances past the stream's parallelism ceiling, or past what the downstream system absorbs, and expecting the backlog to drain faster.

  • Restarting the next node before the previous node's copies have caught up, which can leave a partition with no current copy at all.

  • Promoting an out-of-date copy to restore writes without stating that it silently discards records producers were already told were stored.

  • Alerting on a single sample of a bursty signal, so the rota learns to ignore the page that eventually matters.

  • Assuming a hosted broker removes the on-call duty, when retention mistakes, permission changes, client versions and stalled readers still page the owning team.

The same few decisions recur across nearly every section, and naming the one you are making is usually worth more than the setting itself. - **Durability versus write latency.** Each extra copy that must accept a write before the answer adds a round trip and closes off one route to losing it. The right threshold depends on what losing a single record costs the business, not on a platform default. - **Availability versus acknowledged data.** When the caught-up copies are gone, a cluster can wait, unwritable, or promote a stale copy and silently drop confirmed records. Which one a stream should choose is a product decision to make in advance, not during the incident. - **History versus disk and money.** A longer retention window buys replay and recovery room and costs bytes on every copy. On rented tiers retained bytes are billed, multiplied by the copy count. - **Parallelism now versus flexibility later.** More partitions raise the reader ceiling, and each one costs coordination, open files and recovery time; on most platforms the count is easy to raise and hard to lower. - **Sharing versus isolation.** One shared cluster is cheaper to run and puts every team behind the same failure, capacity and upgrade calendar. Dedicated clusters cost more and contain the blast radius. - **Batching throughput versus per-record delay.** Waiting to fill a batch cuts per-request overhead at the node and adds delay to the first record in each batch; at low traffic it is almost all delay.

Several shapes recur across the sections under different names; recognising them is the fastest way to place an unfamiliar question. - **Widen, move, narrow.** Credential rotation, protocol upgrades and many configuration changes first make every node accept both the old and the new, then move clients, then remove the old. Skipping the first step produces intermittent failures rather than a clean break. - **Declare, then copy, then settle.** Adding a node, moving a partition and evacuating a node all start with a plan accepted in seconds and finish only when retained records have been copied and ownership has moved. The plan is not the change. - **One at a time, wait for green.** Rolling restarts and rolling upgrades take one node down, wait until its copies are current again, and only then touch the next. The waiting is the safety, not a formality. - **Reversible first, irreversible last.** Recovering from a full disk, rolling back an upgrade or cutting over to a standby all order their steps so that the one which destroys history or cannot be reverted comes last, after cheaper options are exhausted. - **Measure in time, not in counts.** Lag, retention and recovery point all mislead when stated in records or bytes and become meaningful as seconds or hours a person would notice. - **The quiet failure.** Silent truncation under a byte ceiling, a stale copy promoted, records discarded under saturation: the dangerous cases raise no error, so each needs a reading someone actually watches.

explore

report an issue with this guide →

questions

259 · 12 sections

Why is a messaging client configured with one entry address instead of the address of every node in the cluster?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The entry address only has to reach any one member, because the cluster then describes itself and hands back the per-node addresses later connections use. A hand-written list of every node duplicates membership the cluster already knows, and goes stale.

open as a page

Why does a broker's append path ask more of a volume's sustained throughput than of its seek performance?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Broker storage is dominated by ordered appends, and most reads arrive moments later asking for the same bytes in the same order, so the volume is asked for steady sequential bytes per second rather than fast random seeks.

open as a page

Three copies of one partition sit on three record-serving nodes in the same rack — what does that placement protect against?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Node-level failures only: a disk, a process, one machine. The rack is a single power and network domain, so one rack event takes all three copies at once — three copies against one failure class, one copy against another.

open as a page

In a broker cluster, what work does the metadata role do that a record-serving node does not?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A cluster splits two jobs. Record-serving nodes accept writes and serve reads for the partitions or queues they own. The metadata role holds membership, stream and partition metadata and configuration, and admits every structural change.

open as a page

A stream is split into a fixed number of partitions and a shared queue is not split at all - what caps reader parallelism in each?

level: juniorimportance: must knowfreq 70%
basics
~20 s

On platforms that split a stream into partitions, the count is the parallelism ceiling: each partition is normally served to one reader at a time, so extra readers idle. A shared queue has no such count; competing readers simply take more work.

open as a page

The data volume under a broker node is full — what happens to writes landing on that node, and which streams feel it?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A full data volume is a refusal, not a slowdown: the node stops accepting writes, and it stops accepting them for every stream whose data lives on that volume, not only the one that filled it. Reads usually keep serving.

open as a page

A stream is configured with both a seven-day age bound and a byte ceiling — which one decides when the oldest record is removed?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Whichever binds first. Both bounds are live at the same time, and the oldest records become removable as soon as either the age is reached or the bytes are reached, so a busy stream hits the byte ceiling long before seven days.

open as a page

A team turns on keyed retention to shrink a stream before a deadline — what does keeping only the latest value per key actually promise?

level: middleimportance: must knowfreq 55%
basics
~20 s

Keyed retention promises only that the latest value for every key stays readable. It never states when a superseded value is removed, so it is a guarantee about what survives, not a size or deadline lever.

open as a page

When a stream uses remote offload, which segments leave the data volume, and what does that change about writes?

level: middleimportance: must knowfreq 46%
basics
~20 s

Remote offload moves only closed segments to a remote object store; the segment still being appended to stays on the data volume. History stops being bounded by the local device, while the write path is untouched, so it buys capacity, not speed.

open as a page

Before a broker answers a write, what can the writer be made to wait for, and what does each option cost in write latency?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A write can be answered with no wait at all, once the leader holds it, once a majority of copies hold it, or once every caught-up copy holds it. Each rung up adds a network round trip of latency and removes one way to lose the record.

open as a page

Within one cluster, a stream's copy count goes from two to three - what does the third copy cost in stored bytes and traffic?

level: juniorimportance: must knowfreq 60%
basics
~20 s

The third copy adds a complete extra set of the stream's stored bytes - about fifty per cent more disk, since the base was two - plus one more transfer of every incoming byte, continuously, and a one-off backfill of everything already retained.

open as a page

A broker acknowledges a record as stored, yet a power cut on that machine loses it — why?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An acknowledgement usually means the bytes were accepted into the operating system's file cache, which is volatile memory, not that they were forced onto persistent media. Everything accepted but not yet forced is the loss window a power cut takes.

open as a page

A stream keeps three copies within a single cluster. Which failures does that survive, and which does it not?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Copies survive the loss of whatever they do not share: a broker process, a host, a volume. They do not survive what replication reproduces faithfully - a deleted stream, a bad record, expiry, or losing the whole cluster.

open as a page

A stream's acknowledgement rule is every caught-up copy and its minimum-copies floor equals its copy count; what happens when one copy drops out?

level: middleimportance: must knowfreq 62%
basics
~20 s

Writes to that stream are refused until a third copy is caught up again. Setting the floor equal to the copy count turns the loss of any one copy into a write outage, which is why the floor is normally set one below the count.

open as a page

A cluster node joins a busy broker cluster and reports healthy — does the load on the existing nodes drop yet?

level: juniorimportance: must knowfreq 72%
basics
~20 s

No. A newly joined node holds no records, so it serves nothing and relieves nothing until units of ownership are assigned to it and their records have been copied across. Detached-storage designs are the exception.

open as a page

When the leader of a partition (or queue) dies, what do writers see before another copy is promoted within the same cluster?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Writes to that partition fail rather than block: the client's cached owner map still names the dead node, so sends are rejected until a replacement copy is promoted. Records survive only if the writer's retries outlast that gap.

open as a page

Each cluster node holds an equal number of copies, yet one node serves most reads and writes — why?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Leadership is a second distribution, separate from where copies sit. Where exactly one node serves each partition (or queue), the work follows leadership rather than storage, so perfectly even copies can still leave one node leading — and serving — most units.

open as a page

A reassignment plan moving a partition (or queue) to another cluster node is accepted in a second — why is the move not finished?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A reassignment plan only declares which node should hold which partition (or queue). Accepting it starts ordinary copy traffic that streams the unit's retained records to the new node, and ownership settles only once that copy has caught up.

open as a page

Each cluster node holds copies of partitions (or queues) - why does a maintenance restart take one node at a time, waiting between?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A restart takes that node's copies out of service, and while it is down they fall behind. Stopping the next node before they have caught up removes a second current copy and can leave a partition with none.

open as a page

After an outage, what must a reader group's reading rate exceed before its backlog of unread records starts to shrink?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The arrival rate. Only the surplus between reading rate and arrival rate drains a backlog of unread records, so capacity sized to keep pace in steady state merely holds the backlog at whatever depth the outage left it.

open as a page

When a reader group shares one stream's work, which membership events cause its shares to be reassigned?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Three events: a member joins, a member leaves cleanly, or a member stops proving it is present — by missing its liveness signal or by holding work past its progress deadline. Designs where readers merely compete reassign nothing.

open as a page

A reader group's lag shows two million unread records, so why does that count alone not say whether anyone is waiting?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Unread count states the gap in records; unread age states it in time. Two million records may be seconds old on a fast stream or half a day old on a slow one, so only the age reading shows staleness.

open as a page

What happens to the waiting unread records when an operator moves a reader group's read position forward to the newest record?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Moving a reader group's read position forward abandons every record between the old position and the new one, so nothing processes them. The gap drops to zero at once because work was skipped, not done.

open as a page

A reader group member still holds its share and makes no progress: what separates a stuck member from a merely slow one?

level: juniorimportance: must knowfreq 62%
basics
~10 s

A slow member still completes records, just fewer than arrive, so its read position keeps advancing. A stuck member completes nothing: its position is parked. Only the slow case is answered with capacity.

open as a page

Why does a writer group many records into one batch before sending them to a broker node?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Because a large part of what a request costs a node is fixed, whatever the request carries: a round trip, a queue slot, a handler, per-request bookkeeping. A batch spreads that fixed cost over many records, so the same hardware accepts far more records per second.

open as a page

A broker node cannot process writes as fast as they arrive, with every configured ceiling respected — what can it do?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A saturated node has four responses: queue the work, delay the answer, block the writer, or refuse the request. Each shows up differently at the call site, and the quietest discards records without failing the send.

open as a page

A broker enforces a byte-rate ceiling on one client identity - what exactly is capped, and what does that client see?

level: juniorimportance: must knowfreq 75%
basics
~20 s

A byte-rate ceiling caps how many bytes per second the cluster will handle for one named principal - a client identity, user, application or tenant - not per connection. Over it, that principal is either answered more slowly or refused.

open as a page

Why is a broker node's ceiling on concurrent connections a separate ceiling from its bytes-per-second allowance?

level: juniorimportance: must knowfreq 60%
basics
~20 s

A connection cap counts connections held open, not traffic. Holding one consumes memory and bookkeeping on the node whether or not data flows, so a fleet of near-idle clients can exhaust the cap while the byte rate stays tiny.

open as a page

A single record exceeds the maximum record size a broker node accepts — what happens to that write and to everything else?

level: juniorimportance: must knowfreq 68%
basics
~20 s

The node refuses that one write with a size error and keeps serving everything else on the same connection. Retrying the identical record never succeeds: it is a deterministic check, so the record must shrink or a ceiling must rise.

open as a page

A broker cluster gives each team its own named space for its streams — which properties does that space scope, and which does it not?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A named space scopes three things: the names streams may take, the grants written against them, and the point a quota attaches to. Nothing physical is scoped — disks, memory, network and node failure stay shared across every space on the cluster.

open as a page

A joining team reads a stream's name before any documentation — which segments should the name carry, and why each?

level: juniorimportance: must knowfreq 65%
basics
~20 s

A stream's name should carry the owning domain, its purpose, the environment and a version segment. Each answers a question a joining team would otherwise have to ask a person, and grants and quotas are written against the leading segments.

open as a page

On a cluster where naming a stream in a client call creates it, why is that convenience a governance hazard?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Creation on first use turns every typo into a permanent stream: empty, billed, and indistinguishable from a deliberate one. It also means the stream is born with whatever values the platform supplies and no recorded owning team, rather than values anyone chose.

open as a page

Why is renaming a live stream a migration rather than an edit, and what is written against the old name?

level: middleimportance: must knowfreq 58%
basics
~20 s

Most platforms have no rename operation — the name is the stream's identity — so a rename means standing up a second stream and moving everyone. Grants, quotas, retention settings, mirroring rules, dashboards and alert conditions all name the old string.

open as a page

A stream has taken no writes for thirty days, so what further evidence does a decommission need before calling it unused?

level: middleimportance: must knowfreq 58%
basics
~20 s

Absence of writes is one line of evidence out of three. A decommission also needs no reader — nothing attached, no stored reading position advancing — and, where records stay readable after delivery, no replay, all observed over a stated observation window and put to the owner record.

open as a page

What does a single broker grant have to name before the cluster can decide whether a request is allowed?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A grant names four things: the principal making the request, the operation attempted, the named resource it is attempted on, and whether that combination is permitted or refused. Leave any one loose and the rule decides more than intended.

open as a page

When a client opens a connection to a broker, what families of proof can it present, and what does each one establish?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A client proves itself with a shared secret through a salted challenge-response password exchange, with a client certificate, with a directory-issued or issuer-signed short-lived token, or with delegated platform identity where it holds no secret at all.

open as a page

The volume beneath a broker cluster is encrypted at the storage layer — which reader does that stop, and which does it not?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Volume-level encryption beneath the broker stops bytes that leave the cluster: a removed drive, a copied volume image, replaced hardware. It stops nobody who connects, because the storage layer decrypts beneath the broker process, which then serves plaintext to any principal holding a grant.

open as a page

A broker cluster encrypts client connections but not the node-to-node hop — what is exposed there?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Whole record payloads, their keys and the cluster's coordination traffic. Where copies are kept by nodes forwarding records to each other, that hop carries every record again for each extra copy, continuously, for as long as the cluster runs.

open as a page

Beyond reading and writing records, which operations does a broker authorize separately, and why does that separation matter?

level: middleimportance: must knowfreq 60%
basics
~20 s

Brokers separate far more than reading and writing: creating a stream, deleting it, changing its settings, listing or describing names, recording a read position, and cluster-wide administration are distinct rights. Teams that grant only two end up granting all of them.

open as a page

A cluster's health list shows one partition with no node currently serving it. What happens to writes and reads for that partition?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Writes and reads for that partition fail: with no node holding the job of serving it, there is nothing to accept or answer requests. The stream's other partitions keep working, so the outage is a slice, not the whole cluster.

open as a page

A stream's delivery interval is four minutes while its writes are acknowledged in six milliseconds — what is the interval measuring?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The delivery interval covers the whole path: from a write being accepted to a reader finishing the work on that record. Acknowledgement time covers only the first hop, so the wait in the stream and the reader's own processing are what make up the four minutes.

open as a page

A node's mean publish latency is 9 ms while its p99 is 700 ms — which number does a writing client feel, and why?

level: juniorimportance: must knowfreq 70%
basics
~20 s

The p99. Most requests a node serves are cheap, so the mean tracks them and stays flat, while the tail is where a writer actually waits — and one writer in a hundred waiting 700 ms is what a customer reports.

open as a page

When a messaging cluster reports every signal it publishes as normal, what does that actually guarantee about the pipeline's work?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A green cluster dashboard guarantees only that the broker accepted writes, stored them, served reads and recorded that readers moved forward. It says nothing about whether any record was understood, transformed, or turned into useful work downstream.

open as a page

An alert on a reader group's record lag fires on a single sample and pages nightly on write bursts; what does a usable condition need instead?

level: middleimportance: must knowfreq 62%
basics
~20 s

A usable condition names three things: which signal it reads, the threshold that counts as a breach, and a sustained window the breach must hold across. Size that window longer than one full write-and-drain cycle of the flow it watches.

open as a page

After a broker cluster moves to a newer release, why does an application's years-old client library keep working?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Because the wire version is settled per connection, not per cluster. At the connect step each side says what it can speak, the connection uses a shape both understand, and the upgraded nodes still answer in that older shape.

open as a page

A value has a cluster-wide default and a per-stream override on one stream — which value does that stream use, and which do the rest?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The override wins for that one stream; every stream or queue without an override follows the cluster-wide default. The effective value is therefore a per-object answer, and editing the default leaves overridden objects exactly where they were.

open as a page

Why is a new broker release rolled out one broker node at a time, and what runs side by side while it is?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Taking one broker node at a time keeps every other member serving, so the cluster stays available through the change. The price is a mixed-version window: until the last node is done, two broker releases are live in the same cluster.

open as a page

Why does one edited cluster-wide default change a running broker node's behaviour immediately while another waits for that node to restart?

level: middleimportance: must knowfreq 57%
basics
~20 s

Settings split into live-applying values, which a running broker node re-reads and acts on at once, and restart-only values, which the process reads while starting and holds for its lifetime. A restart-only edit is not ignored — it is pending until the next restart.

open as a page

During a two-pass roll, why is the agreed internal version raised only after the last broker node carries the new binaries?

level: middleimportance: must knowfreq 60%
basics
~20 s

Because the agreed internal version fixes how members speak to each other. Held at the old level, a new binary keeps talking the way an old one expects, so any node can still be reverted. Raising it early would leave old members unable to follow their peers.

open as a page

A standby cluster is kept fed by an ongoing copy and carries no writers — what happens when an operator switches onto it?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Writers are pointed at the standby, readers restarted there, and only one site keeps accepting writes. The standby holds only what the copier had carried, and a reader's stored position from the source names a different record there.

open as a page

Your only continuity plan for a live stream is last night's file backup of the broker data volumes. What has that already cost you by morning?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A file backup fixes a stream at the instant it was taken, so every record written since is gone, along with every reader's progress. Streams keep moving while files do not, which is why the gap is counted in hours.

open as a page

Why is a cross-cluster copy of a stream always behind the source cluster that feeds it?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A cross-cluster copier is an ordinary client of both clusters: the source stores a record and answers the writer before the copier has even read it. The record therefore exists on the source first and on the target some time later.

open as a page

Your stream is copied asynchronously to a second cluster — why can the recovery point you state for it never be smaller than the copy lag?

level: juniorimportance: must knowfreq 70%
basics
~20 s

The asynchronous copy hop sets the floor. Records acknowledged on the source cluster but not yet carried to the target exist in one place only, so losing the source loses them — the recovery point is at least the copy lag.

open as a page

Two sites accept writes to a stream of the same name, with a cross-cluster copier running each way — what stops a record circulating forever?

level: middleimportance: must knowfreq 52%
basics
~20 s

An origin stamp: a marker on the record, or an origin-qualified stream name, that tells the copier running the other way this record came from elsewhere. Without one, each copier treats the other's output as new local traffic and the loop never ends.

open as a page

Moving from a self-run broker cluster to a hosted tier, which broker settings does the rental most commonly withhold?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Hosted broker tiers usually withhold the durability and security internals: how many replica copies a stream keeps, the acknowledgement floor, flush behaviour, and pluggable identity or decision hooks. Stream design, permissions, client versions and the spend stay yours.

open as a page

Your hosted broker carried almost no records all weekend yet still billed heavily — which broker quantities are charged for existing rather than flowing?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Much of a hosted broker's bill is standing rather than usage-driven: purchased throughput capacity reserved ahead of the traffic, the streams, partitions or queues that merely exist, and retained bytes still held, multiplied by their replica copies.

open as a page

A team moves its broker cluster to a hosted tier - which operating duties does the rental take over, and which never transfer?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A rental takes over the machines: replacing dead nodes, patching the engine, keeping the endpoint reachable, growing capacity inside the shape you bought. Stream design, retention choice, permissions, client versions, the reading side and the spend stay yours.

open as a page

On a rented broker tier, which ceilings will the provider not raise however you ask, and what classes do they fall into?

level: middleimportance: must knowfreq 52%
basics
~20 s

A hosted broker's hard caps fall into five classes: how many streams or queues the cluster admits, how many partitions, how many concurrent connections and subscriptions, how large one message may be, and how long history is kept at all.

open as a page

A cluster that was cheap on machines you own becomes expensive on a rented tier at the same traffic — which design choices explain that?

level: seniorimportance: must knowfreq 56%
basics
~20 s

Habits that are free on owned hardware are billed dimensions once rented: many small streams and a fine-grained partition layout, a generous retention policy multiplied by replica copies, huge populations of tiny records inflating billed request counts, and readers placed across a charged boundary.

open as a page