skip to content

The Managed Operating Boundary

What a hosted offering genuinely takes over - dead nodes, patching, copy placement, capacity - and what never transfers. Asked because the duties left behind are the ones that page you at night.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

4

A team moves its broker cluster to a hosted tier - which operating duties does the rental take over, and which never transfer?

level: juniorimportance: must knowfreq 68%

answer

  1. machines transfer, meaning does not
  2. contractible without knowing your data
  3. design, retention, permissions, clients, readers, spend
  4. a duty was removed, not an outage

basics

~20 s

A rental takes over the machines: replacing dead nodes, patching the engine, keeping the endpoint reachable, growing capacity inside the shape you bought. Stream design, retention choice, permissions, client versions, the reading side and the spend stay yours.

solid answer

~50 s

Renting a broker buys you out of infrastructure work, not out of operating the messaging system. The provider owns the fleet: a failed node is replaced, the engine is patched, the address clients dial stays reachable, storage and hardware faults are theirs, and capacity is added inside the purchase shape you chose. What it cannot take over is anything that needs to know what your data means - how many streams or queues exist and how they are named, how long records are kept, who may connect and what they may do, which client versions your services run, whether the reading side is keeping up, and what the whole thing costs. These are the `residual duties`, and they are the ones that page you. The honest summary is that a hosted tier removed a class of work, not a class of outage.

go deeper

for a junior

Memorise the residual list in the order it bites: stream design and naming, retention choice, permissions, client versions, the reading side, the spend. Being able to say "renting removed a duty, not an outage" out loud is most of the mark here.

for a middle

Explain why the line falls where it does: a provider can only take on duties it can define without knowing what your data means. Then name the split duties - node repair and capacity growth - where the provider performs the work but you own detecting and living with it.

for a senior

Show it from an incident you can describe: a healthy cluster by the provider's definition while nothing useful was being delivered, and the residual duty that caused it. Interviewers want the residual work named without being prompted for it.

for a principal

Frame it as an ownership question rather than a feature comparison. Every residual duty needs a named owner before the move, and the transferred duties still need detection and an escalation path on your side even though the repair capability is gone.

A **hosted broker** is a broker or streaming cluster that somebody else operates and you buy as a service. The interview question is never "is managed good" - it is whether you can draw the line between what the rental genuinely performs and what silently stayed on your side of it. ## The duties a rental genuinely performs These are the duties definable without knowing anything about your application, which is exactly why a provider can sell them: - **Machine and node lifecycle.** A node that dies is replaced. Where the design keeps per-node replica copies of a stream's data, the copies that lived on the dead node are re-placed onto its replacement; on designs where storage is shared rather than replicated per record, there is nothing to re-place and the replacement simply re-attaches. Replacing failed hardware is the duty every hosted offering sells, though how visible that repair is to you differs. - **Patching the engine.** Security and bug-fix releases of the broker software are applied for you, within the set of releases the purchase tier offers. - **Keeping the endpoint reachable.** The address clients dial, its certificates, and whatever spreads connections across nodes behind it are the provider's to keep answering. - **Storage and failure domains.** Disks, the machines under them, and the spreading of copies across independently-failing units, where the tier offers that. - **Capacity growth inside the shape you bought.** A provisioned-capacity tier adds pre-bought units of capacity or nodes when you ask; a consumption-priced tier absorbs a traffic rise without being asked. Either way the provider performs the growth - but the shape being grown is one you chose. - **Maintenance.** The provider schedules a maintenance window and restarts nodes inside it. ## The residual duties - what no rental transfers 1. **Stream design and naming.** How many streams or queues exist, how they are split, what they are called, and which of those names other teams now depend on. 2. **Retention choice.** How long a record survives is a policy decision about your data and your regulator. A provider can enforce it; it cannot choose it. 3. **Permissions.** Who may connect and what each principal may do to which stream. 4. **Client versions and client behaviour.** The services that produce and read are yours, and so is every upgrade of them. 5. **The reading side.** On platforms where readers own a stored position, a reader that falls behind accumulates unread records; on destructive-read designs the same duty shows up as an unacknowledged backlog that stops draining. Either way the provider watches its cluster, not your readers. 6. **The spend.** The bill follows from choices you made, and nobody else will notice it rising. ## Why the line falls exactly there A provider can only assume a duty it can define in a contract without knowing your business. "Replace a node within N minutes" is contractible. "Keep the order-events stream useful to the fulfilment team" is not. Everything that requires knowing what your data means, who needs it, and how long it matters stays with the people who know those things. | Duty | Who performs it on a hosted tier | |---|---| | Replacing a dead node, re-placing the copies it held | Provider | | Patching the broker engine | Provider | | Keeping the client endpoint reachable | Provider | | Growing capacity | Provider performs it, inside a shape you chose | | Deciding how many streams and what they are named | You | | Deciding how long records are kept | You | | Granting and reviewing permissions | You | | Upgrading producing and reading services | You | | Noticing the reading side has stopped keeping up | You | | The bill | You | ## The consequence teams miss Some duties are not transferred or retained but **split**, and those are where surprises live. The provider repairs a dead node, but you are still the one who notices that publishing slowed while it was being repaired, and you are still the one who tells the affected teams. The provider grows capacity, but only into the shape you chose up front - on platforms that split a stream into parts, that shape caps how much parallel reading is possible at all, and on shared-queue designs the shape question is instead how many queues exist and who competes on each. Said in one line: **managed removed a duty, not an outage.** The cluster can be perfectly healthy by the provider's definition while nothing your business cares about is being delivered - and every reason for that sits in the residual list above.

  • Which duties are split between you and the provider rather than cleanly owned by one side?
    Node repair and capacity growth. The provider performs both, but you own the detection and the consequences: you notice that publishing slowed during a repair and tell the affected teams, and you chose the shape that capacity is being grown into. Treating a split duty as fully transferred is how teams end up with nobody watching the half that stayed.
  • Does a rental take over deciding how many copies of a stream exist?
    No - it takes over placing and repairing them. Whether you may even choose the number is a separate question about which settings the purchase tier exposes, and tiers differ: some let you set it per stream, some fix it for the whole cluster. What is consistent is that keeping the chosen copies healthy is the provider's work.
  • Why is "the provider is responsible for availability" an incomplete answer here?
    Because it names the provider's availability target for its cluster, which is not the same thing as your messages being delivered. A cluster can be inside its target while a reader is stalled, a permission is wrong or a stream was never created - all residual duties, all yours, and none of them visible as provider downtime.

Renting a flat: the landlord fixes the boiler, replaces the roof and repaints the hallway. Nobody rents you out of deciding what goes in the fridge, who has a key, or what the shopping costs.

saying these in an interview costs you the question

  • Says a hosted tier means there is nothing left to operate
  • Believes the provider chooses retention or creates streams for you
  • Thinks permissions are handled because the endpoint is authenticated
  • Assumes client upgrades come with the rented engine upgrade
  • Calls capacity growth fully transferred, ignoring the shape you chose
  • Treats the provider's availability target as a delivery guarantee for your messages
open as a page

After moving to a hosted broker, a team proposes halving its broker on-call rota - which incident classes still page it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Every incident whose cause is a residual duty: a stream that was never created or was misnamed, records aged out too early, a permission change that cut off a client, a client version mismatch, readers falling behind, and a bill nobody watched. Renting changed who repairs a node, not who is paged when delivery stops.

open as a page

A hosted broker's provider restarts your nodes during a maintenance window it schedules - what does that reveal about the rental boundary?

level: middleimportance: should knowfreq 46%

basics

~20 s

It shows the rental took the repair work and the scheduling control together. The provider decides when nodes go away; your residual duty is that the workload survives a node leaving at a moment you did not pick.

open as a page

Across an estate of hosted brokers, how do you decide which residual duties need a named owner and a rota?

level: principalimportance: should knowfreq 42%

basics

~20 s

Write the residual duties out, and give each one an owner judged on two axes: can it cause an outage on its own, and does it need someone awake or only someone accountable. Duties the provider performs still need detection and escalation on your side, never repair.

open as a page