skip to content

A move stays inside its bandwidth ceiling, yet reads on unrelated partitions on the same cluster nodes slow down — where is the contention?

level: seniorimportance: should knowfreq 52%

answer

  1. the ceiling caps bytes, not the work
  2. cold reads on the source node
  3. warm data displaced by a long scan
  4. sustained write on the destination
  5. cut concurrency before lowering bandwidth

basics

~20 s

A bandwidth ceiling caps migration bytes on the wire, not the work that produces them. The source node must read cold historical records off disk, which competes with live reads and can evict the warm data they were being served from; the destination pays a sustained write.

solid answer

~50 s

The copy is not just bytes on a link. To feed it, the source node reads records that are old — the historical part of the unit of ownership, not the recent tail that live readers touch — so it drives disk reads that a bandwidth ceiling on the wire does not account for. On platforms that serve recent reads out of the host's file cache, that long sequential scan of cold records also evicts the warm working set, so ordinary reads that used to be served from memory start hitting storage: the visible symptom is read latency on units that are not moving. At the destination the copy is a sustained sequential write competing with that node's own live writes and its flush path, plus the processing cost of checksums and any re-compression. The fixes are fewer units in flight, scheduling against real headroom, and choosing source copies that are not also serving as leaders where the design permits it.

go deeper

for a junior

The point to carry away is that copying records costs disk work at both ends, not just network. A move can look well within its allowance and still slow the cluster down.

for a middle

Explain the mechanism: the source reads cold historical records to feed the transfer, which competes with live reads and can displace what they were being served from, while the destination absorbs a sustained write.

for a senior

Show the diagnosis and the order of remedies: check link utilisation against the ceiling, confirm disk saturation, cut the number of units in flight, choose a source copy that is not serving clients, and only then touch the bandwidth number.

for a principal

The recurring question is whether the estate should keep paying this contention at all. Price the operational cost of every move against a storage design in which ownership changes without a transfer, and set a standard for when migrations may run.

## The ceiling covers one resource; the move consumes several A **move throttle** is a ceiling on bandwidth. What else it accounts for varies by platform, and the gap between "bytes on the wire" and "everything the transfer costs" is where this symptom lives. A reassignment that is comfortably inside its ceiling can still be the reason production is slow, because the bytes have to come from somewhere and land somewhere. ## Where the rest of the cost lands - **Source disk reads.** To send records, the source node reads them. Those records are the *historical* part of the unit of ownership — a partition, or on queue-shaped brokers a queue — not the recent tail that live readers are consuming. That is a large sequential read that was not happening before the move started. - **Cache displacement.** On platforms that serve recent reads out of the host's file cache, streaming gigabytes of cold records through that cache evicts the warm working set. Ordinary reads that were memory-speed become disk reads, and the latency change shows up on units that are not being moved at all. Platforms that manage their own read cache inside the process show a milder version of the same effect. - **Destination write path.** The receiving node is taking a sustained sequential write on top of its own live traffic, including whatever flushing or durability work that write path does. - **Processing cost.** Checksum verification, and on some designs decompression and re-compression, cost processor time at both ends. - **Second-order replication.** The destination is also a node with its own duties; the extra load can push *its* other copies behind their leaders, widening catch-up distance elsewhere. ## Reading the symptom | What you see | Where to look first | |---|---| | Read latency up on units not being moved | Source node disk utilisation and cache hit behaviour | | Write acknowledgement times up on the destination | Destination disk write utilisation and flush behaviour | | Catch-up distance widening on unrelated copies | Both nodes' saturation, not the ceiling | | Link utilisation well under the ceiling | Confirms the bottleneck is storage or processing, not bandwidth | ## What to do about it 1. **Cut concurrency first.** Fewer units in flight means fewer source disks being scanned at once. This helps more than lowering the ceiling, because it removes whole read streams rather than slowing all of them. 2. **Pick the source deliberately** where the design allows a copy that is not currently serving as leader to be the source; that keeps the extra reads off the node answering client traffic. 3. **Schedule against real headroom.** Move during the hours when the working set is small and the disks are idle, and accept that large moves outlast quiet periods, so keep a dial you can turn when traffic returns. 4. **Lower the ceiling as a second step**, keeping it above the unit's incoming write rate so the move still converges. 5. **Re-measure rather than assume.** If link utilisation is far below the ceiling while disks are saturated, bandwidth was never the constraint and lowering it further only extends the intermediate state. ## The mistake to avoid The tempting diagnosis is that migration traffic is being prioritised over client traffic somewhere in the network. That is almost never it. The source node keeps serving its clients for the whole transfer — that is the design — and what has changed is that the same hardware is now also doing a long cold read it was not doing before. Treating it as a scheduling or priority problem sends you looking at the wrong layer. ## Where designs differ Some platforms account for the read side and throttle it too, some throttle only the transfer, and some expose separate ceilings for sending and receiving. Where the read cache lives differs as well, which changes how sharply the displacement effect bites. And on designs with **detached storage**, where a unit's records sit on shared or remote storage, there is no long cold read at all — the new owner simply starts serving from the same place, paying a cold-start penalty on its own caches rather than a transfer. Name the assumption you are making when you answer, because the shape of the contention follows directly from where the platform keeps the data.

  • How would you confirm that storage, not bandwidth, is the constraint?
    Compare the two directly: if link utilisation for the migration sits well below its ceiling while the source node's disk read utilisation is at or near saturation, the ceiling is not what is binding. Falling transfer rate with flat link usage points the same way. Lowering the ceiling further in that state buys nothing and lengthens the intermediate mapping.
  • Why does the slowdown appear on units that are not part of the plan?
    Because the contended resources are per node, not per unit. Every unit served by that node shares its disks, its cache and its network interface, so a long cold read on behalf of one unit's transfer is felt by all of them. That is also why picking which node sources the copy matters as much as how fast it is allowed to go.

saying these in an interview costs you the question

  • Blames network priority rather than storage contention.
  • Assumes a ceiling on bandwidth caps every resource the move uses.
  • Believes the source node stops serving clients while it is a copy source.
  • Thinks only the units named in the plan can be affected.
  • Restarts the source node to 'clear' it, discarding warm caches and making reads worse.
  • Lowers the ceiling again when link utilisation was already far below it.