skip to content

questions

3

Expanding a live cluster is approved with a date attached: what decides whether added capacity earns work at once or only after a deliberate move?

level: middleimportance: should knowfreq 52%

answer

  1. shape decides, not machine count
  2. pulled work against owned work
  3. shared queue absorbs on contact
  4. split stream waits for assignment
  5. approving a move, not a machine

basics

~20 s

Consumption shape decides. Competing consumers pulling from a shared queue absorb new capacity as soon as they reach it; a stream split into fixed parts earns a new broker node nothing until parts are assigned to it.

solid answer

~50 s

The shape of consumption decides, and it changes what you are approving. Where consumers compete for messages on a shared queue, work is taken by whoever pulls next, so capacity that becomes reachable is capacity that gets used — the expansion is finished when the join is. Where a stream is split into fixed parts and each part has an owner, a broker node that has joined earns work only once parts are deliberately assigned to it; the expansion is not finished at the join, it is a decision to start a data move, and it should be approved as one. Some designs sit outside both cases: where storage is shared rather than owned per broker node, a new member can serve without anything being copied. Name the shape before you promise relief by a date.

go deeper

for a junior

Recall that a cluster gaining a machine and a cluster gaining throughput are not the same event. Be able to say that on some designs a newly joined broker node has to be given work before it does any.

for a middle

Explain the discriminator: work pulled from a shared queue is taken by whoever pulls next, while a stream split into owned parts gives a new broker node work only when ownership moves to it.

for a senior

Show that you translate the shape into a date. Say what you would promise the requester, and refuse to promise relief on the day a machine joins when the relief depends on an assignment nobody has scheduled.

for a principal

Argue about what is actually being approved. An expansion whose benefit needs a data move is a decision to run that move, and your review process should force the requester to ask for the whole thing rather than the easy half.

## Why the schedule, not the hardware, is the answer An expansion request almost always arrives with a date on it: headroom before a seasonal peak, before a large tenant is onboarded, before a volume that is filling reaches its ceiling. Approving it is therefore a promise that relief arrives by that date. The binding constraint on that promise is rarely how fast a machine can be provisioned or how long an installation takes. It is **how the cluster hands work to a broker node that has just arrived**, and that differs sharply by design. This is the single most useful thing to establish before an approval, because it decides whether the request in front of you is a small change or the front half of a large one. ## Two consumption shapes answer "when" differently **Competing consumers on a shared queue.** Where work is pulled — several consumers reading the same queue, each taking the next message — nothing has to allocate a share to anybody. Whoever can pull next does the work, so capacity that becomes reachable is capacity in use, with no further step and nobody to schedule. This is the shape in which an expansion genuinely is complete when the join is. **A stream split into fixed parts.** Where the stream is divided into parts and each part has an owner, **ownership is the unit of work**. A broker node that has joined is a member and a candidate owner, and it earns work only once parts are deliberately assigned to it. That assignment is a separate act, with its own duration and its own risk. Until it happens the cluster has one more machine and exactly the distribution of work it had before. **Shared storage under the nodes.** Some designs keep the data off the broker nodes entirely. There a new member can begin serving without anything being copied onto it, which shrinks the distance between joining and earning work — though routing still has to send it something. | Shape | What decides who works | When added capacity earns work | What the approval really covers | | --- | --- | --- | --- | | Competing consumers, shared queue | Who pulls next | As soon as clients can reach it | The join, and reachability | | Split stream, fixed parts | Who owns each part | Only after parts are assigned | Starting an assignment, approved as such | | Shared storage beneath the members | Routing | On joining, subject to routing | The join, and how clients are steered | ## What this changes about the approval 1. **State the shape before you state the date.** "Capacity in place on Tuesday" and "capacity carrying work on Tuesday" are the same sentence in one shape and two different sentences in another. 2. **If the benefit needs an assignment, you are approving the assignment.** An expansion whose relief arrives only after parts are moved is a decision to start that move. It must be approved as the longer, riskier thing it is, not as the short thing the request described. 3. **Decide in advance what happens if the second half never happens.** A cluster that has gained a machine and not gained a distribution is a cluster paying for a machine. That can be a perfectly acceptable end state — but it should be a chosen one, not a surprise discovered in an invoice. ## The caveats that bite - A queue hosted by a single broker node is not helped by another broker node until that queue sits somewhere else; the shared-queue claim is about work being **pulled rather than assigned**, not a promise that every queue benefits from every machine. - Where the constraint is on the reading side rather than in the cluster, more broker nodes change nothing that matters, and the request is aimed at the wrong half of the system. - Clients usually learn about members through discovery, so "reachable" can trail "joined" by a refresh interval. Small, but it belongs in an honest date. - Extra machines never help when the ceiling being hit is a per-object one rather than an aggregate one. ## How to tell which shape you are in You rarely need documentation for this; the operational behaviour answers it. - Ask what happens to the distribution of work when a broker node is lost. If the survivors simply keep pulling, work is pulled. If something has to decide new owners, work is owned. - Ask whether the cluster's own view names an owner per unit of work. Where an owner exists, a new machine is not an owner until something makes it one. - Ask what was actually done the last time the cluster grew. If the answer includes a planned step that ran for hours and was watched, the benefit of the next expansion will arrive the same way. ## What a strong answer sounds like A strong answer refuses a single rule and names the discriminator: **work that is pulled absorbs capacity on contact; work that is owned absorbs capacity only when ownership moves.** It then draws the consequence for the approval — that in the second case the request understates what is being asked for — and it says out loud that the two halves may be approved separately, the join now and the assignment on a stated date, as long as nobody is told that relief has arrived when only a machine has.

  • The cluster gained a broker node on Friday and nothing improved by Monday. What should be established first?
    Whether anything was supposed to change on its own. Where a stream is split into owned parts, nothing moves until parts are assigned, so an unchanged Monday is the expected result of a join with no assignment behind it rather than a fault. Establish the shape, and whether the second half was ever requested, before anyone investigates the machine.
  • Why is "we added fifty per cent more broker nodes" a poor thing to report back to the requester?
    It reports the input rather than the relief. The requester asked for headroom by a date; broker nodes are what was spent. Report the date the added capacity begins carrying work and what that date depends on — reachability and client discovery in one shape, an approved and completed assignment in the other.

saying these in an interview costs you the question

  • Assumes adding a broker node always relieves pressure immediately
  • Promises relief by a date without naming the consumption shape
  • Treats the join and the data move as one small approval
  • Believes every design copies data onto a new member automatically
  • Claims a queue hosted on one broker node is helped by any extra machine
  • Reports machines added instead of the date capacity carries work
open as a page

A team wants to expand a live cluster during a scheduled change window at 3 a.m. What does the quiet hour actually reduce, and what does it not?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A quiet hour reduces exposure, not reversibility: fewer clients affected, more spare headroom. A step that cannot be undone at noon cannot be undone at 3 a.m., so approval still turns on the half-finished state.

open as a page

Who should be allowed to approve expanding a live cluster under traffic, and on what stated evidence should they refuse?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

The cluster's operating owner together with the owner of the affected workloads — not the requester alone. Refuse unless waiting carries a named, dated risk, someone can say when the capacity earns work, and the half-finished state has an answer.

open as a page