Why does a broker's append path ask more of a volume's sustained throughput than of its seek performance?
answer
- the write only ever goes at the end
- reads follow the writes by seconds
- bytes per second, not accesses per second
- predictable worst case beats a good average
basics
~20 sBroker storage is dominated by ordered appends, and most reads arrive moments later asking for the same bytes in the same order, so the volume is asked for steady sequential bytes per second rather than fast random seeks.
solid answer
~50 sA broker's write path is almost entirely one-directional: a record arrives, it is placed after the one before it, and it is not updated in place afterwards. Readers that are keeping up then ask for the bytes appended seconds earlier, in the order they were appended, so the same region is touched twice in quick succession. The useful specification is therefore `sustained bytes per second` plus a predictable worst-case write time — not the scattered-access figure an index-driven store would care about, because nothing has to be located by key before it can be written. How much storage work this actually creates does vary: a platform that retains records writes everything down and keeps it, while a broker that removes a record once it is acknowledged may keep a queue that is being drained largely in memory and lean on the volume only when readers fall behind or the process restarts.
go deeper
Remember the shape: records are added at the end and not changed later, and whoever is reading is usually just behind. That is why steady bytes per second matters more than how fast the media can jump around.
Be able to say why sustained throughput and small-random-access figures are independent numbers, and name what makes a broker's pattern less sequential in practice: many independently-appended units on one volume, a reader sweeping old records, and a neighbour workload.
Show that you judge storage on its worst case, not its average, because producers wait on individual writes. Be ready to describe how you would confirm the volume actually sustains the rate you need rather than trusting an advertised figure.
The angle to own is that this is a purchasing and standardisation decision: which storage character every record-serving node in the estate gets, what write-latency promise that supports, and what it costs to change once the fleet is built on it.
## What a broker's storage actually sees A broker's durable store is written the same way nearly every time: a record arrives, it is placed after the record that arrived before it, and it is never modified afterwards. Nothing is updated in place, nothing is shuffled to keep a sort order, and nothing has to be located by key before it can be written. That pattern is **the append path**, and it is the single most important fact about the storage a broker wants under it. Contrast it with the store a transactional database keeps. There, a write may have to consult an index, find a page somewhere in the middle of a large file, modify it and write it back, and a read may land anywhere in a data set much larger than memory. That workload is dominated by how quickly the media can move to an arbitrary location, and by how many small independent accesses it can complete per second. ## Throughput and seek are two different numbers A volume is described by at least three figures, and they do not move together: - **Sustained throughput** — the bytes per second it can absorb or deliver while load keeps arriving. This is what an append path consumes. - **Small random accesses per second** — how many independent, scattered operations it can complete. This is what an index-driven store consumes. - **The service-time distribution** — not only the average time a write takes, but the slow tail. Producers wait on individual writes, so the high percentiles are felt directly while the average hides them. A volume can be strong on one of these and ordinary on another, which is why "a fast disk" is not a specification. For broker storage the specification is narrower: can it absorb the sustained write rate, and does it do so with a predictable worst case? | what the store does | what it asks of the volume | the figure that decides it | |---|---|---| | appends in arrival order, never updates in place | absorb a continuous stream of bytes | sustained write throughput | | serves a reader that is keeping up | return bytes written seconds ago, in order | throughput, and often memory instead | | serves a reader that has fallen far behind | sweep older records forward, in order | sustained read throughput | | looks records up by key at random | move to arbitrary locations repeatedly | small random accesses per second | ## Where the reads come from The second half of the pattern matters as much as the first. On a broker, most reads arrive moments after the write and ask for the same bytes in the same order. Two things follow. First, the region being read was touched seconds ago, so it may still be in memory and never reach the media at all. Second, even when it does reach the media, the read is itself a sweep forward rather than a hunt. The pattern degrades in three recognisable ways, and each is worth naming when you talk about storage: 1. **Many independently-appended units on one volume.** Each partition or queue is appended to on its own, so a record-serving node holding hundreds of them interleaves hundreds of append streams onto a single volume. From the volume's point of view that begins to look less sequential, which is why the per-node count of independently-appended units belongs in the storage conversation. 2. **A reader that has fallen far behind.** The records it wants are no longer in memory, so it pulls a long sweep of older data while new records are still arriving, and reads and writes now share the same throughput. 3. **A volume shared with something else.** Another workload's scattered access pattern is interleaved with the broker's sequential one, and neither gets the device's best behaviour. ## What varies between platforms The append shape is common across this product class, but how much storage work it creates is not: - Platforms that **retain** records for a period write everything down and keep it, so the volume carries the full incoming rate continuously and serves re-reads afterwards. - Brokers that **remove a record once it is acknowledged** may keep a queue that is being drained almost entirely in memory, touching the volume mainly to survive a restart or when readers fall behind. The same storage choice can stay invisible for months and then dominate an incident. - Some designs keep the primary store in a **remote service** rather than on the node at all, which moves the same question from a device's characteristics to a network service's. ## What this settles, and what it does not It settles the *character* of the storage you put under a record-serving node: throughput-oriented, comfortable with large sequential transfers, judged on its worst case rather than its average. It does not settle **how much** storage is needed, which is an arithmetic question about rate, record size and how long records are kept. It does not settle whether an acknowledged write has actually reached durable media, which is a durability question with its own answer. And it says nothing about what happens when the volume runs out of space.
- Does the append path stay sequential when one record-serving node holds many partitions or queues?Not entirely. Each one is appended to independently, so a node holding hundreds of them interleaves hundreds of append streams onto the same volume, and from the volume's side that looks progressively less like one ordered stream. This is why the per-node count of independently-appended units is part of the storage conversation, and why one volume behaves differently from several.
- Why does a predictable worst-case write time matter more than an excellent average?Producers wait for individual writes to be accepted, so they feel the slow ones. A volume that is quick on average but occasionally stalls for hundreds of milliseconds produces a long tail on the producer's own latency distribution. The average hides the stall; the high percentiles are what the application and its callers actually experience.
Appending records is like adding lines to the bottom of a long roll of paper rather than filing cards into a cabinet. The pen never has to find a place, it just keeps moving, and whoever is reading is usually only a few lines behind. The roll rewards a fast, steady hand; the filing cabinet rewards a quick index finger. Buying storage for a broker means buying the steady hand.
saying these in an interview costs you the question
- Says broker storage needs a high random-access figure because it is a database.
- Claims the append path only writes the volume and never reads it.
- Assumes fast media removes the need to think about the write path at all.
- Treats a volume's advertised peak rate as the rate it will sustain under load.
- Believes every broker lays records down the same way, so one storage choice fits all.