Would you standardise your broker estate on volumes local to each record-serving node or on network-attached volumes?
answer
- speed against survival
- storage that outlives the machine
- a network hop inside the write path
- one standard for the whole estate
basics
~20 sNeither wins outright. Local volumes give the highest and least variable throughput but die with the node; network-attached volumes survive the node and can be reattached, paying a network hop in the write path and a throughput allowance you must buy.
solid answer
~50 sThis is an estate-level standard, not a per-cluster preference, because it is embedded in the node shape, the machine types you buy and the recovery procedure your operators know. **Local storage** sits inside the record-serving node: the best sustained throughput, the tightest service-time distribution, and no network in the write path — but when the node is gone, so is the data it held, so a replacement starts empty. **Network-attached storage** is reached over the network and detachable: the stored data outlives the node and a replacement can be brought up against the same volume, at the cost of a network hop inside the append path, more variable service times and a throughput allowance bought separately from size. The honest answer names the workload: a latency promise measured in single-digit milliseconds pushes toward local, while an estate optimised for cheap, fast node replacement pushes toward attached — and some platforms sidestep the question by keeping the primary store in a remote service.
go deeper
The idea to hold is simple: storage inside the machine is fastest but disappears with the machine, while storage reached over the network survives it. Everything else in this discussion follows from that one difference.
Be able to state both costs rather than one. Local storage means a replacement node starts empty; attached storage means a network hop in the write path and a throughput allowance that is bought separately from size.
Show that you judge attached storage on its variability rather than its average, and that you know what refilling a replacement node actually costs in your own estate before calling either option safe.
Own it as a standard with a written rationale: the write-latency promise it supports, the node-replacement cost it assumes, whether storage must scale apart from compute, and what rebuilding the estate would cost if the answer changes.
## What the choice actually is Strip away vendor framing and there are two shapes. In the first, the storage is physically part of the record-serving node: the node and its data are one object, and losing the node loses both. In the second, the storage is a separate object reached over the network and attached to a node for as long as that node needs it: the node and its data can be separated, and either can be replaced without the other. This is a *character* decision, not a sizing one. How large the volume must be is arithmetic from the incoming rate, the record size and how long records are kept, and it is the same arithmetic either way. ## What local volumes buy and cost - **Buy:** the highest sustained throughput available to the node, the tightest service-time distribution, and no network between the append path and the media. Nothing between the broker and the device schedules against it. - **Buy:** no separate throughput purchase to get wrong, and no per-volume allowance to be throttled by. - **Cost:** the data does not survive the node. Replacing a node means a new one that starts empty and must be refilled from the cluster's own redundancy. - **Cost:** the storage is tied to the machine type, so storage and compute are bought together whether or not they scale together. ## What network-attached volumes buy and cost - **Buy:** the stored data outlives the node. A replacement can be attached to the same volume, so node loss and data loss are separate events. - **Buy:** storage and compute are sized independently, and the volume can often be grown without rebuilding anything. - **Cost:** a network hop inside the append path. The average is usually acceptable; it is the variability — the occasional long service time from the network or from the provider's own scheduling — that lands on producers' high percentiles. - **Cost:** throughput becomes a purchased allowance, separate from size, and one more thing to get wrong quietly. ## Side by side | dimension | local volume | network-attached volume | |---|---|---| | sustained throughput | typically the highest available | bounded by a purchased allowance | | service-time variability | lowest | higher, and partly outside your control | | survives loss of the node | no | yes, and can be reattached | | storage sized apart from compute | no | yes | | failure modes to reason about | the device and the machine | the device, the machine and the network path | ## A third shape Some platforms sidestep the comparison by keeping the primary store in a remote service and treating whatever is on the node as a working area rather than the record of truth. That changes what you are deciding: not which media a node owns, but how much recent data is kept near the compute and what the remote service's own rate and failure behaviour are. When such a platform is on the table, the local-against-attached framing is the wrong question to ask of it, and pretending otherwise produces a comparison that flatters whichever option you already preferred. ## Deciding it for an estate 1. **Name the write-latency promise.** If the product commits to a tight upper bound on write acceptance, the variability of an attached volume is the thing most likely to break it, and local storage becomes the default. 2. **Name what node replacement costs today.** If refilling a replacement node from its peers is routine and fast, losing data with a node is an inconvenience; if it takes hours and saturates the cluster, keeping the volume matters much more. 3. **Decide whether storage and compute must scale apart.** Estates where record volume grows far faster than processing are much better served by storage that is bought separately. 4. **Write down the required sustained rate.** Whichever shape is chosen, that figure is what stops somebody later buying the cheap tier for a cluster that needed the expensive one. ## Why it is expensive to revisit The choice propagates into the machine types purchased, the node shape, the capacity model, the recovery procedure operators have rehearsed and the promises the platform team has made about write latency. Reversing it means rebuilding every record-serving node in the estate and re-validating those promises. That is why it is worth an explicit decision with a written rationale, rather than being inherited from whatever the first cluster happened to run on — and why a principal is expected to name the workload the answer depends on rather than declaring one shape universally correct.
- When does a network-attached volume's variability matter more than its average speed?When producers are latency-sensitive. The average is usually fine; it is the occasional long service time, from the network or the provider's own scheduling, that lands on the producer's high percentiles. A workload that must promise a tight upper bound on write acceptance is the one that pushes back toward local media.
- What makes this choice expensive to reverse later?It is embedded in the machine types bought, the node shape, the capacity model, the recovery procedure operators have rehearsed and the write-latency promises already made. Changing it means rebuilding every record-serving node and re-validating those promises, so it is decided once and lived with for years.
- Does storage that survives the node reduce how much redundancy the cluster needs?No. A volume outliving its node protects against one failure mode; it does nothing about a corrupted volume, an unavailable storage service, or a failure domain going down with everything in it. Treating durable storage as a substitute for the cluster's own redundancy is how single-copy estates get built by accident.
saying these in an interview costs you the question
- Picks local volumes for speed without saying what a lost node then costs.
- Assumes a network-attached volume is always slower than a local one.
- Treats storage that survives the node as a substitute for cluster redundancy.
- Believes the choice can be revisited cheaply once the estate is running.
- Thinks throughput on an attached volume scales automatically with its size.