skip to content

A reassignment's copy traffic shares disks and links with live traffic — what is the move's bandwidth ceiling for, and what breaks at each extreme?

level: middleimportance: must knowfreq 62%

answer

  1. one dial, two real failure modes
  2. copy traffic competes with production
  3. too low leaves it half-migrated
  4. size against peak headroom, not average
  5. reduce concurrency before raising the number

basics

~20 s

The ceiling caps how much bandwidth a migration's copy traffic may consume, because that traffic competes with production. Set it too high and live reads and writes slow down; set it too low and the cluster sits half-migrated indefinitely, holding copies on both node sets.

solid answer

~50 s

Executing a reassignment plan is ordinary copy traffic over the same disks and links that serve live reads and writes, so the **move throttle** — the bandwidth ceiling on that migration traffic — is the one dial between two real failure modes. Too high, and the copy takes whatever capacity exists: write acknowledgement times climb, ordinary reads start hitting disk, and copies of unrelated units begin drifting behind their leaders. Too low, and catch-up distance closes so slowly that the move never converges against live writes; the estate stays in its intermediate mapping for days, paying disk on both the old and the new node sets and blocking later plans. The honest way to set it is against the headroom available at the busiest hour, not the average, and then to move fewer units at a time rather than reaching for a bigger number.

go deeper

for a junior

Know that a migration's copying competes with live traffic and that operators cap how fast it may go. The word to have is a bandwidth ceiling on the move, and the reason is that the disks and links are shared.

for a middle

Explain both extremes concretely: too high degrades live reads and writes and thins the caught-up set; too low leaves the cluster half-migrated for days and can stop the move converging against incoming writes.

for a senior

Demonstrate operating it: size against peak headroom, leave margin for a node failure mid-move, cut concurrency before raising the number, and watch remaining catch-up to distinguish slow from stuck.

for a principal

Make it a standard rather than a judgement call per incident: what headroom a migration may consume, who may raise it, and whether the estate's growth pattern justifies a storage design in which moves cost no bandwidth at all.

## What the ceiling is actually protecting Executing a reassignment plan is not a special kind of traffic. The destination node reads records from a node that already holds the unit of ownership — a partition, or on queue-shaped brokers a queue — over the same network links, off the same disks and through the same host cache that live producers and readers are using. A cluster running at seventy per cent of its capacity has thirty per cent spare, and a migration will happily take all of it and then some. The **move throttle** is the ceiling on how much bandwidth that migration may consume. It exists because the alternative is an operator choosing between never migrating and degrading production every time they do, and both failure modes are real; this is not a knob where one direction is safe. ## Setting it too high When the ceiling is above the headroom that actually exists, the symptoms are ordinary saturation symptoms and they arrive on units that are not being moved at all: - Acknowledgement times for writes climb, because the disks answering those writes are also feeding the copy. - Reads that were being served warm start hitting storage, on platforms where recent records are served out of the host's cache. - Copies of *other* units begin drifting behind their leaders, so the caught-up set thins across the cluster. - In the worst case it cascades: a thinner caught-up set means a write that waits for it waits longer, which pushes latency up further. ## Setting it too low The low side is quieter, which is exactly why it is worse to diagnose. Nothing alerts. The move simply does not finish: - If the allowed copy rate sits below the rate at which new records arrive on the unit, catch-up distance grows rather than shrinks and the move cannot converge at all. - Even when it can converge, a ceiling a fifth of what is available turns a six-hour move into a multi-day one, and the estate sits in the intermediate mapping for the whole of it. - That intermediate state is not free: both the original and the target node sets hold copies, so disk is committed at both ends and monitoring shows more copies than the plan's final count. - Later maintenance queues behind it — a second plan on the same units is generally refused or deferred while one is in flight. ## Choosing a number | Input | How to get it | |---|---| | Available headroom | Measure disk and link utilisation at the *busiest* hour, not the daily average. | | Safety margin | Leave room for a node failure during the move, which itself generates copy traffic. | | Concurrency | Decide how many units move at once first; the ceiling then divides across them. | | Convergence floor | The rate must exceed the unit's incoming write rate, or nothing finishes. | The order matters. Operators reach for a bigger ceiling when the right move is usually fewer units in flight: a ceiling expressed per node or per link is multiplied by every concurrent transfer that touches that node, so concurrency, not the number itself, is what most often causes the surprise. ## What the ceiling does not cover A bandwidth ceiling is a ceiling on bandwidth. What it covers beyond that varies by design, and the difference shows up as contention the operator thought they had capped: the disk reads that feed the transfer, the flush path on the destination writing a sustained stream, and the processing cost of checksumming or re-compressing records may all sit outside it. A move can be comfortably inside its ceiling and still be the reason production is slow. ## Where designs differ Some platforms let the ceiling be changed while a plan is in flight, so the usual operating pattern is to start conservative and raise it while watching production; others fix it when the plan is submitted. Some apply it per node, some per pair of nodes, some to the whole migration. And on designs with **detached storage**, where a unit's records live on shared or remote storage, there is no migration bandwidth to cap, because changing owner moves no bytes. Say which model you are assuming when you answer, and say that you would find out which one you were operating before choosing a number.

  • What is the first thing you would change if a move is hurting production — the ceiling or something else?
    Reduce how many units are in flight before touching the ceiling. A ceiling is usually applied per node or per link, so every concurrent transfer through the same node multiplies its effect; cutting concurrency lowers real consumption without also slowing the units that are nearly done. Lowering the ceiling as well is fine, provided the result still exceeds the incoming write rate, or the remaining moves stop converging.
  • How do you tell a throttled move from a stuck one?
    Look at the trend of remaining catch-up, not at elapsed time. A throttled move shows catch-up distance falling steadily but slowly, and you can predict a finish time from the slope. A stuck move shows it flat or growing, which means the allowed copy rate is at or below the unit's write rate; no amount of waiting fixes that, only more bandwidth or less load.
  • Why not simply run every reassignment during a quiet period with no ceiling at all?
    Because large moves routinely outlast the quiet period, and an unthrottled transfer that is still running when traffic returns is the worst case: full-rate copy traffic against peak load, with no dial left to turn. A ceiling sized for peak headroom, raised during the quiet hours and lowered before traffic returns, gives the same speed with a way out.

saying these in an interview costs you the question

  • Treats a high ceiling as harmless because copying is background work.
  • Sees only the fast-move risk and not the never-finishes risk.
  • Sizes the ceiling against average utilisation rather than peak.
  • Assumes the ceiling covers every resource the copy consumes.
  • Raises the ceiling instead of reducing how many units move at once.
  • Thinks a stalled move eventually resolves itself if left alone.