skip to content

Who should be allowed to approve expanding a live cluster under traffic, and on what stated evidence should they refuse?

level: principalimportance: nice to knowfreq 36%

answer

  1. capacity is a decision, not a chore
  2. waiting must carry a dated risk
  3. say when capacity earns work
  4. half-finished state needs an owner
  5. two signatures, cluster and workloads

basics

~20 s

The cluster's operating owner together with the owner of the affected workloads — not the requester alone. Refuse unless waiting carries a named, dated risk, someone can say when the capacity earns work, and the half-finished state has an answer.

solid answer

~50 s

Treat it as a decision with a gate rather than a chore. Three things must be on the request before anyone signs: a **named, dated risk of waiting** — what runs out, when, and what happens then; a **statement of when the added capacity actually carries work**, which is at once in some shapes and only after an assignment in others; and an **answer for the half-finished state** if the work is interrupted. The pen belongs jointly to whoever operates the cluster and whoever answers for the workloads on it, because the second party owns the consequence; the requesting team alone is not enough, and a lone on-call engineer at four in the morning is not either. Refuse on any missing gate, and hand back a deferral with the cost of waiting written down and a review date, so the refusal is a decision rather than an obstruction.

go deeper

for a junior

Know that adding capacity to a running cluster is a change somebody has to approve, and that the team asking for it is not by itself the team that approves it.

for a middle

Be able to say what a complete request contains: what runs out and when, what the added capacity will carry, and when it actually starts carrying it.

for a senior

Show the refusal. Turn down a request with no dated risk behind it or no answer for a half-finished state, and hand back a deferral with the cost of waiting written down.

for a principal

Set the standing rule: who may approve a change made under traffic, what evidence is mandatory, and how an accepted risk is recorded so that a deferral has an owner and a review date.

## Why this needs a standing rule at all Adding capacity feels benign. Nothing is being deleted, nobody is being downgraded, and the request usually arrives with good intentions and a spreadsheet. That is exactly why it slips through with less scrutiny than a version change, and why organisations end up with clusters that grew on a Tuesday afternoon because somebody felt uneasy about a graph. A standing rule costs one page and removes the argument from the moment, which is the only moment at which it cannot be had calmly. It also protects the requester: a refusal against a written gate is a refusal they can answer, rather than a matter of who was in the room. ## The three gates | Gate | Evidence that satisfies it | What its absence looks like | | --- | --- | --- | | A dated risk of waiting | What runs out, on what date, and what happens on that date | "Headroom feels low" with no date and no consequence | | A schedule for the benefit | When the added capacity begins carrying work, and what that depends on | A machine count presented as if it were relief | | An answer for the partial state | What is true if the work is interrupted, and who may call a stop | A start time and an estimate, and nothing else | The second gate is the one most often missed, and it is the one that changes the size of the decision. Where consumers compete for work on a shared queue, capacity that becomes reachable is capacity in use, and the request is as small as it looks. Where a stream is split into owned parts, a joined member earns nothing until parts are assigned to it — so the request is really asking for that assignment, and it must be reviewed as the longer, heavier thing it is rather than as the join. ## Who holds the pen Two signatures, for two different consequences. - **The cluster's operating owner** answers for the thing being changed: whether the work can be performed safely now, who will watch it, and who may stop it. - **The owner of the affected workloads** answers for the consequence: whether a degraded or partly redistributed cluster is acceptable during the work, and on whose behalf that is being agreed. Three tempting shortcuts are all wrong for the same reason — the person signing does not carry the consequence: 1. **The requesting team alone.** They own the deadline, not the cluster. Self-approval on capacity is how an estate grows without anybody able to say why. 2. **The on-call engineer, mid-incident.** Pressure work under traffic is exactly what an on-call rotation should not be authorised to invent alone. Emergency authority should exist, but it should be explicit, narrow, and reviewed the next morning. 3. **Whoever happens to hold administrative access.** Access is a capability, not a mandate. The two must be separated deliberately, because they drift together by default. ## What a refusal sounds like A good refusal is specific and short. It names the missing gate, not a mood. "There is no date on the risk of waiting, so there is nothing to weigh the work against — come back with what runs out and when." Or: "the relief you need arrives only after parts are assigned, which is a different and longer piece of work; ask for that, with its duration." Or: "there is no answer for what stands if this is interrupted, so nobody can say what they are accepting." None of these is a judgement about competence, and all of them can be satisfied within a day, which is why the rule makes people faster rather than slower once they know it. ## Deferral is a decision, and it is costed The standing rule has to make deferral a first-class outcome, or approvers will approve in order to avoid looking obstructive. - Name the cost of waiting in the same terms as the risk: what is more likely, by when, and how bad. - Attach a review date, so that deferral is not a way of never deciding. - Have the cost **accepted by someone who can carry it** — usually the workload owner, sometimes a level above. An accepted risk has an owner; an unspoken one has a victim. - Record it where the next person will find it. The most common failure is not a bad deferral but an invisible one, rediscovered three months later by whoever is on call when the date arrives. ## The estate view At the scale of an estate, the interesting question stops being any one cluster. It becomes: how many of these decisions should be reaching a human at all? Capacity requests that arrive without dated risks are a sign that nothing is forecasting; expansions whose benefit needs an assignment nobody scheduled are a sign that the request template asks the wrong question. Both are fixable in the template rather than in the meeting — which is the principal-level version of this answer, and the one that scales past the first dozen clusters.

  • Should an on-call engineer ever be able to expand a cluster under traffic without those signatures?
    Yes, but only under an explicit, narrow emergency mandate: named circumstances, a bound on what may be done, and a review the following morning. The point is that the exception is written down in advance. Emergency authority invented during an incident is indistinguishable from nobody being in charge.
  • How do you stop a costed deferral from quietly becoming a decision never to act?
    Give it the two things a decision has and a delay does not: an owner who accepted the cost, and a date on which it is looked at again. Record both where the next person on call will see them, because the usual failure is not a wrong deferral but an invisible one.

saying these in an interview costs you the question

  • Lets the requesting team approve its own capacity change
  • Expands because headroom feels low, with no dated risk
  • Treats "we can always add more later" as the plan
  • Gives a lone on-call engineer authority to expand at 4 a.m.
  • Counts deferral as free rather than costed and owned
  • Accepts "it should be quick" in place of a schedule for the benefit