On a broker cluster shared by several teams, why do one team's writes slow down during another team's nightly burst?
answer
- shared machines, shared fate
- no per-tenant slice by default
- bytes, requests, handler time, connections
- the quiet neighbour complains first
basics
~20 sShared nodes mean shared finite resources. A tenant does not get its own slice of the machine, it gets a turn, so one owner's burst consumes the network, disk, memory and request-handling threads that every other tenant on those nodes is waiting for.
solid answer
~50 sA broker cluster is a small set of machines, and by default almost nothing on them is divided per tenant. Network bandwidth, disk throughput, the memory a node uses to serve recent records, and the finite pool of threads that actually do request work are all pooled. When one owner's traffic grows, it does not eat into capacity reserved for it — there is no such reservation — so every other request on those nodes queues behind it and is answered more slowly. The burst can be expensive on several independent axes: bytes per second, requests per second, how long each request occupies a handler, and how many connections are held. The tenant causing it usually sees nothing wrong, because it is being served; the complaint comes from a quieter neighbour with a tighter latency budget.
go deeper
Recall the core fact: several owners share the same machines, and there is no per-tenant reservation by default, so one owner's burst is paid for by everyone else on those nodes.
Explain the independent cost axes — bytes per second, requests per second, how long each request occupies a handler thread, connections held — and why a tenant can be small on one and dominant on another.
Show that you start from the burst's timestamp rather than from the complainant, and that you expect the causing tenant to look healthy while a quieter neighbour with a tight deadline raises the incident.
Frame it as a coupling cost: a shared cluster couples unrelated teams' release and batch schedules, and that coupling is what the organisation is buying in exchange for the efficiency of sharing hardware.
## What "shared nodes" actually means A broker cluster is a finite set of machines. Every tenant that connects to it — every team, application or environment — pushes its traffic through the same network interfaces, the same disks, the same memory and the same bounded set of threads that perform request work. Almost none of that is partitioned per tenant unless someone deliberately partitions it. **A tenant does not get a slice of the machine; it gets a turn.** That single fact is the whole mechanism behind the symptom. When one owner's traffic grows, it is not consuming "its own" capacity, because on shared hardware no such capacity exists. It consumes the pool, and every other request waits longer for the same pool. This is why a shared cluster is, operationally, a shared outage: the blast radius of one team's bad afternoon is every team on those nodes. ## The axes a burst is expensive on It is worth being precise, because these move independently and a tenant can be enormous on one and invisible on another: - **Bytes per second**, written or read. The most obvious axis and the one everybody looks at first. - **Requests per second.** Many tiny requests cost far more per byte than a few large ones: each one has to be received, parsed, dispatched and answered. A tenant can be modest in bytes and ruinous in requests. - **Request-time share** — how long each request occupies a handler thread. A reader sweeping far-behind historical records is dramatically more expensive per request than one reading records that arrived seconds ago, because on platforms that serve recent data from memory the old records have to come off disk instead. - **Connections held.** Connections are not free; a client fleet that opens one per instance, or per request, occupies node resources before it sends anything. - **Cache and disk pressure.** A heavy write burst can displace what other tenants were reading cheaply from memory, turning their cheap reads into disk reads. That cost can outlive the burst itself. ## Whose symptom shows up first Counter-intuitively, almost never the bursting tenant's. It is being served — perhaps a little more slowly, which nobody notices inside a batch job with no latency budget. The team that raises the incident is typically the quiet one whose request-response path has a tight deadline. | Who | What they observe | |---|---| | The bursting tenant | Its work completes; timings look unremarkable for a batch workload | | A latency-sensitive neighbour | Slower answers, timeouts, client-side retries, a paging alert | | The operator's cluster-wide view | "Busy" — which looks identical to healthy growth | This asymmetry is the reason the incident channel and the cause are almost never the same team, and it is why "who complained" is a poor starting point for diagnosis. ## Why "we didn't change anything" is usually true The complaining team generally is telling the truth: nothing on their side changed. A shared cluster couples release schedules, batch windows and traffic shapes that are otherwise completely unrelated. A backfill someone scheduled overnight, a new reader started by another department, or a fleet that was scaled out for unrelated reasons all land on the same machines. From inside any one tenant, the cluster simply got slower for no reason — which is exactly what shared hardware feels like from the inside. ## The three questions this opens Naming the mechanism is the easy half. What follows is a sequence: 1. **Is the slowdown real and local to these machines?** A cluster that is healthy overall can still have specific nodes under pressure. 2. **Which owner's traffic changed at the same moment?** This is contention attribution, and it is the genuinely hard step, because the totals only say "busy". 3. **Which axis did it change on?** The answer decides which lever is even applicable. ## What this question does not settle How much of the cluster any one owner is allowed to take, how that allowance is expressed and enforced, and what a node does once it simply cannot keep up are each their own subject. So is the larger structural call — whether these workloads should be sharing hardware at all — which is an estate decision rather than something you reach for during an incident. What belongs here is narrower and prior to all of them: on shared nodes, one owner's cost is paid by its neighbours, and until you can say *which* owner, no lever can be chosen at all.
- Why can a tenant that sends very few bytes still be the expensive one?Because bytes are only one axis. A tenant sending a very high rate of tiny requests pays per-request overhead on every one of them, and a tenant reading far-behind historical records occupies a handler thread for far longer per request than one reading freshly written records. Either can dominate a node's request-handling capacity while looking small on a byte-rate chart.
- Does the bursting tenant experience anything at all?Usually very little, which is why it does not self-report. It is being served; its own requests are competing too, so its timings may drift slightly, but a batch workload has no deadline to miss. The effect concentrates on neighbours with tight latency budgets, which is exactly the population least likely to have caused it.
A shared road at rush hour. Nobody owns a lane; one fleet of lorries joining the traffic slows every other driver, and the lorry drivers arrive fine and never hear about it.
saying these in an interview costs you the question
- Assumes each tenant gets its own reserved slice of a node
- Believes only byte volume can make a tenant expensive
- Starts the investigation from whoever filed the incident
- Treats a busy cluster-wide chart as proof of healthy growth
- Thinks adding nodes is the first response to contention