skip to content

questions

3

A shared cluster's overall numbers look merely busy while one tenant reports slow writes — how do you find which owner is responsible?

level: middleimportance: must knowfreq 60%

answer

  1. aggregates hide the distribution
  2. break usage down per owner
  3. loudest in bytes is not costliest
  4. retries inflate the victim after the fact
  5. identities must be distinct beforehand

basics

~20 s

Break usage down per owner on the affected nodes, over the exact interval of the complaint, and on several axes at once — bytes, request rate, request-time share, connections. Cluster-wide totals cannot separate owners, and the tenant reporting the incident is rarely the cause.

solid answer

~50 s

Aggregate numbers only ever say "busy", which is indistinguishable from healthy growth, so attribution needs usage broken down by quota principal — the client identity, application or tenant a request is charged to — restricted to the nodes involved and to the minutes around the complaint. Compare several axes, because the largest owner by bytes is often not the costliest one: a high rate of small requests, or reads of far-behind records, can dominate request-handling capacity while looking modest in byte terms. Anchor on the moment the symptom started and look for the owner whose profile changed then, not the one that is simply biggest. Two traps recur: a victim's retry storm inflates its own numbers after the fact and looks like a cause, and if several applications connect under one shared identity they collapse into a single unattributable blob.

go deeper

for a junior

Recall that a busy cluster chart is a sum, and a sum cannot say which owner grew. Finding the cause means looking at usage per owner, not at the cluster as a whole.

for a middle

Explain the mechanics: the exact interval around the symptom, the specific nodes involved, a breakdown per quota principal, and comparison on bytes, request rate and request-time share rather than one axis.

for a senior

Demonstrate the judgment that separates a cause from a reaction — a victim's retries inflate its numbers after the onset — and that you check attribution is even possible before promising to find the culprit.

for a principal

Treat attributability as a platform requirement, not an incident skill: one identity per application, retained per-owner baselines, and visibility weighed when choosing between running the cluster and renting it.

## Why the cluster-wide view cannot answer this A total is a sum over owners, and summing is exactly the operation that destroys the information you need. Aggregate bytes in, aggregate request rate and aggregate utilisation on a node all rise identically whether ten tenants each grew ten percent or one tenant tripled. "Busy" is therefore not a diagnosis; it is the same picture a healthily growing estate produces. **Contention attribution** — tying an observed slowdown on shared nodes to the owner who caused it — is the step that has to happen before any lever can be chosen, and it is the genuinely difficult part of this whole subject. It is difficult for a structural reason: the tenant that reports the incident is almost never the tenant that caused it. The victim has the tight deadline; the cause has the batch job. So the complaint tells you *when* and *where*, and nothing at all about *who*. ## What attribution actually requires 1. **The right time window.** Not the last hour and not a daily roll-up — the minutes on either side of the symptom's onset. A daily average hides a twenty-minute burst completely. 2. **The right machines.** Contention is local to the nodes both tenants' traffic lands on. An owner that is enormous on machines your complainant never touches is irrelevant to this incident. 3. **Per-owner breakdown.** Usage attributed to a quota principal — a client identity, an application, a team — rather than to the node as a whole. 4. **Several axes at once.** Bytes per second, requests per second, request-time share and connections held. Comparing owners on one axis produces a confident wrong answer roughly as often as a right one. ## The traps - **The loudest is not the costliest.** A tenant sending a very high rate of small requests, or one reading far-behind history that has to come off disk rather than out of memory, can occupy most of a node's request-handling capacity while sitting mid-table on a byte-rate chart. - **Retry storms look like causes.** A client that is being answered slowly or refused will retry, so its request rate climbs *after* the incident begins. Ranking owners by request rate during the incident can therefore promote a victim to prime suspect. Rank by what changed first, not by what is highest now. - **Identity collapse.** If every application in an organisation connects with the same credential or identity, per-owner numbers exist but all say the same thing. This is the single most common reason attribution is impossible, and it can only be fixed before the incident, never during it. - **No baseline.** Without a sense of what each owner normally consumes, "this tenant is using a lot" is unfalsifiable. The useful comparison is against that owner's own ordinary day. ## What varies between platforms This is where honest answers differ, and an interviewer listening carefully will notice: | Attribution granularity available | What you can conclude | |---|---| | Per client identity or application | Direct attribution to an owner | | Per connection or per source address only | Attribution to a fleet, then a mapping exercise to reach an owner | | Per namespace or per stream only | Attribution to a workload, which may be shared by several owners | | Aggregate per cluster only, as with some rented offerings | No attribution from the platform; you must infer from the client side | Because of that last row, attribution sometimes has to be reconstructed from what clients report about their own send and receive rates, correlated across teams. That is slower and less reliable, and it is a real cost of not being able to see inside the thing you rented. ## Making attribution possible before you need it The work that makes an incident tractable is all done in advance: - Give each application a distinct identity, and keep a mapping from identity to owning team. - Retain per-owner usage at a resolution fine enough to see a burst, for long enough to investigate one after the fact. - Keep an ordinary-day baseline per owner, so a deviation is visible as a deviation. - Record which workloads share which machines, so the search can be narrowed to the ones that actually interact. ## After attribution Only once an owner is named does the question of what to do become answerable, because every available response is aimed at a specific owner — a bound on what it may take, a change to what it sends, or, at the far end, moving it off this hardware entirely. Reaching for a lever before attribution means choosing a victim at random, and reaching for more capacity instead means paying to absorb a burst you have not yet understood and that will grow again.

  • Why is ranking owners by request rate during the incident misleading?
    Because being slowed or refused makes a client retry. A victim's request rate therefore climbs after the symptom starts, and can exceed the causing tenant's. Rank by whose profile changed at the onset rather than by whose number is highest mid-incident, and check whether a spike is original or reactive before acting on it.
  • What makes attribution impossible no matter how good your telemetry is?
    Identity collapse: several applications connecting under one shared identity. Every per-owner number then describes the same blob, and no amount of resolution separates them. The fix is administrative and has to precede the incident — one identity per application, with a recorded mapping to the owning team.
  • What can you do when a rented cluster exposes no per-owner usage at all?
    Reconstruct it from the client side: ask each team what their applications were sending and receiving over the interval, and correlate those against the symptom's onset. It is slower, incomplete and depends on teams having their own telemetry, which is itself a reason to weigh visibility when choosing what to rent.

saying these in an interview costs you the question

  • Thinks the cluster-wide totals identify the responsible owner
  • Blames the tenant that filed the incident report
  • Ranks owners on byte rate alone and stops there
  • Mistakes a victim's retry storm for the original burst
  • Averages over a day and misses a twenty-minute burst
  • Assumes every platform exposes per-identity usage
open as a page

On a broker cluster shared by several teams, why do one team's writes slow down during another team's nightly burst?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Shared nodes mean shared finite resources. A tenant does not get its own slice of the machine, it gets a turn, so one owner's burst consumes the network, disk, memory and request-handling threads that every other tenant on those nodes is waiting for.

open as a page

After a bursting tenant is bounded on a shared cluster, why is the contention only capped and charged rather than removed?

level: seniorimportance: should knowfreq 46%

basics

~20 s

While the hardware stays shared, every available lever is a fairness lever, not an isolation one. Bounding an owner limits how much of the pool it can take and moves most of the cost onto it, but the resources are still shared, so interference is bounded rather than eliminated.

open as a page