skip to content

Across an estate of hosted brokers, how do you decide which residual duties need a named owner and a rota?

level: principalimportance: should knowfreq 42%

answer

  1. enumerate from the boundary, not the org chart
  2. two axes: stops delivery, and how fast
  3. detection and escalation, never repair
  4. central where failure crosses teams
  5. unnamed duty is assumed gone

basics

~20 s

Write the residual duties out, and give each one an owner judged on two axes: can it cause an outage on its own, and does it need someone awake or only someone accountable. Duties the provider performs still need detection and escalation on your side, never repair.

solid answer

~50 s

Start from the list rather than from the org chart: stream design and naming, retention choice, permissions and credential lifetimes, client versions, the reading side, and the spend. For each, ask two questions - can it take delivery down by itself, and does it fail suddenly or slowly. Duties that fail suddenly and stop delivery need rota coverage, not just an owner; duties that degrade slowly, like spend or an ageing permission model, need an accountable owner and a recurring review instead. Then handle the transferred duties separately: you keep **detection and escalation** for them but have no repair capability at all, so what you staff is the ability to tell quickly whether a symptom is theirs or yours. The estate-level rule that keeps this honest is that a duty without a name is assumed gone, and assumed-gone duties are found during incidents.

go deeper

for a junior

The idea to take from this is that every duty a rental does not cover needs a person's name against it. Knowing which duties those are matters more at this stage than deciding who gets them.

for a middle

Be able to sort the residual duties by whether they fail suddenly and whether they can stop delivery on their own - that sort is what decides between rota coverage and a periodic review.

for a senior

Demonstrate the transferred-duty half: detection, disambiguation and escalation are what you staff for a cluster you cannot repair, and answering "theirs or ours" fast is the capability worth building.

for a principal

Own the method and its trap: staffing sized by the rented fraction is the wrong arithmetic, and shared ownership of a residual duty is no ownership. Decide deliberately which duties are uniform across the estate and which stay with the consuming team.

At a single-team scale you can carry the residual duties informally. Across an estate of rented brokers - many teams, many purchase tiers, one platform group - informality is how duties evaporate. This is a judgment call with no single right answer, and an interviewer is listening for a method rather than an answer. ## Step 1: enumerate, do not infer Write down the **residual duties** explicitly, because they are the whole surviving surface: - stream and queue design, and the names other teams now depend on; - the retention choice for each stream; - who may connect, what they may do, and when credentials expire; - which client versions the producing and reading services run; - whether the reading side is keeping up; - the spend. Inferring this list from what people currently do reproduces the gap you are trying to close. Derive it from the boundary instead: anything that needs knowledge of your data, your teams or your business was never contractible and therefore stayed. ## Step 2: classify each duty on two axes | Duty | Can it stop delivery alone? | Fails suddenly or slowly? | What it needs | |---|---|---|---| | Stream existence and naming | Yes | Suddenly, at deploy | Owner + change review | | Retention choice | Yes, for replay and late readers | Slowly, discovered late | Owner + periodic review | | Permissions and credential expiry | Yes | Suddenly, at rotation | Owner + rota coverage | | Client versions | Yes | Suddenly, after an engine upgrade | Owner + upgrade tracking | | Reading side keeping up | Yes | Both | Rota coverage | | Spend | No | Slowly | Accountable owner + review | The axes matter because they produce different staffing. A duty that fails suddenly **and** stops delivery needs someone reachable out of hours. A duty that degrades slowly needs an accountable owner and a recurring review; putting it on a rota buys nothing, and putting a sudden-failure duty on a quarterly review buys an incident. ## Step 3: staff the transferred duties differently The duties the provider performs still need something from you, just not repair: 1. **Detection.** You must be able to see that something is wrong before the provider tells you, because the first notice often arrives from your own users. 2. **Disambiguation.** The most valuable capability in a rented estate is answering "is this theirs or ours?" quickly, since both produce the same symptom of data not arriving. 3. **Escalation.** A named path, known before the incident, with the evidence the provider will ask for already collected. What you deliberately do **not** staff is repair capability for transferred duties. Keeping node-rebuild skills sharp for a cluster you cannot log into is training for a job you no longer have. ## Step 4: decide centrally what must be uniform Across an estate, some residual duties are better held once than per team: - **Naming and design conventions** are worth central ownership because their failure crosses team boundaries. - **Permission models and credential lifetimes** benefit from one pattern, because the failure mode is the same everywhere and the review is otherwise never done. - **Client version tracking** is central because a provider's engine upgrade affects every tenant of that tier at once. - **Spend** should be visible per paying scope but owned somewhere that can say no. - **The reading side** stays with the team that owns the consumer, because nobody else knows what "behind" means for that workload. ## The trap to name out loud The standard failure is **budgeting the staffing by the fraction of work rented**. The rented fraction is large and the surviving fraction is the one that pages, so the arithmetic is simply the wrong arithmetic. A second, subtler trap is assuming that because a duty is shared across many teams it is therefore owned by all of them - shared ownership of a residual duty is reliably no ownership at all. The position worth stating: *a residual duty without a name is assumed gone, and assumed-gone duties are discovered during incidents.* Everything above is just a method for making sure every name exists before that happens.

  • Which residual duties are better owned centrally than by each consuming team?
    Naming and design conventions, the permission model and credential lifetimes, and client version tracking - their failures cross team boundaries or hit every tenant of a purchase tier at once. The reading side stays with the team owning the consumer, because only they know what falling behind means for that workload.
  • What do you staff for the duties the provider did take over?
    Detection, disambiguation and escalation. You need to see the problem before the provider tells you, answer quickly whether the cause is theirs or yours, and have a named escalation path with the evidence already gathered. You deliberately do not staff repair, because that capability is gone by design.
  • Why is shared ownership of a residual duty usually worse than assigning it badly?
    Because a badly-assigned duty still has someone who will be asked about it, while a shared one has nobody who believes the question is theirs. In an estate, shared ownership of stream naming or permission review reliably means it is reviewed by no one until an incident forces it.

saying these in an interview costs you the question

  • Sizes the surviving staffing by how much work was rented away
  • Assumes shared ownership across teams counts as ownership
  • Keeps node-repair drills for a cluster nobody can log into
  • Puts slowly-degrading duties like spend on an out-of-hours rota
  • Leaves escalation paths to be discovered during the first incident
  • Treats an unowned duty as an absent one rather than a hidden one