skip to content

Producers see write latency climb while the volume shows spare space and no errors — which storage decision should you suspect?

level: seniorimportance: should knowfreq 52%

answer

  1. the space is fine, the speed is not
  2. size and throughput are separate purchases
  3. a neighbour on the same device
  4. a flat knee against a moving one

basics

~20 s

Suspect how the volume's speed was provisioned: a throughput allowance bought separately from size, or a volume shared with another workload. The append path is being made to wait, and waiting is not an error, a fault or a space alarm.

solid answer

~50 s

Space and speed are separate purchases on most network-attached storage, and a volume can be half empty while its throughput allowance is fully consumed. Two shapes produce this symptom. First, the allowance was derived from the volume's size or left at a default, so the append path tops out at a fixed figure and every byte beyond it waits. Second, the volume is shared — with another workload, another record-serving node, or another tenant on the same device — so the throughput actually available moves with somebody else's schedule. Either way the storage view stays green, because a throttled volume is not a failing volume, and the delay surfaces at the application as producer write latency. On a broker that removes records once they are acknowledged and is keeping up, the same misconfiguration can hide for a long time and only appear once readers fall behind.

go deeper

for a junior

The takeaway is that storage can be slow without being broken or full. If writes are getting slower and the volume reports nothing wrong, how fast it is allowed to go is still worth checking.

for a middle

Explain that size and throughput are separate dimensions on network-attached storage, and that a throttled volume appears as waiting rather than as an error, which is why ordinary storage monitoring stays green throughout.

for a senior

Show the diagnosis: correlate the broker's own write rate against observed write service time, look for a flat knee against a moving one, and compare peer record-serving nodes before touching any broker setting.

for a principal

The angle to own is procurement: what sustained rate every broker volume in the estate is required to support, whether broker storage may ever be shared, and who is accountable when the cheaper tier is chosen for a cluster with a latency promise.

## The symptom, and why it points at storage The reported problem is almost always at the application: producers are waiting longer to have writes accepted, and the effect grows with load. Meanwhile every storage view is reassuring — the volume has room, reports no faults and raises no alarm. That combination is characteristic rather than mysterious. A volume being made to wait is not a volume that is failing, and most storage monitoring is built to notice faults and fullness, not to notice that a healthy device is being held to a figure. The reason this lands on producers specifically is the append path. Records arrive continuously and have to be written out at the rate they arrive; when the volume cannot absorb that rate, the unwritten work accumulates and the acceptance of each write takes longer. The application feels it as latency because that is the only place it can appear. ## Provisioning shape one: throughput bought apart from size On most network-attached storage, how many bytes per second a volume may move is a separate dimension from how many bytes it may hold, and the two are often bought together in a way that hides the coupling — a default allowance derived from the requested size, or a tier chosen for its price per stored byte. The consequence for a broker: - A volume sized generously for retention can still have a modest throughput allowance. - The allowance is usually a hard figure, so behaviour is *flat*: fine below it, and degrading immediately above it. - Growing the volume to go faster sometimes works and sometimes does nothing, depending on whether the allowance is tied to size at all. ## Provisioning shape two: the volume is shared The other shape is contention. The broker's volume may be shared with a batch job, with logs and telemetry written by the same host, with a second record-serving node placed on the same underlying device, or with another tenant entirely. Here the symptom is *not* flat: the same broker load is fast at one hour and slow at another, because the available throughput moves with a schedule you do not control. ## Why it does not look like a storage problem | what the storage view reports | what the application experiences | |---|---| | space used well under the volume's size | producer write latency climbing with load | | no device faults, no degraded state | the slow tail growing faster than the average | | service times "within specification" for the tier | timeouts and retries at the client, under load | | throughput at its allowance, reported as normal | acceptance of each write taking noticeably longer | ## Telling the two apart 1. **Plot the broker's own write rate against the observed write service time over several days.** A hard allowance shows a knee at the same figure every time. Contention shows no consistent knee. 2. **Check whether the slow periods correlate with the broker's load or with the clock.** Correlation with the clock, or with another team's schedule, points at a neighbour. 3. **Compare record-serving nodes.** If one node is slow and its peers carrying similar load are not, look at what else that node's volume is shared with rather than at the broker's settings. ## What varies between platforms - A platform that **retains** records writes the full incoming rate to the volume continuously, so a throughput shortfall is visible almost immediately under load. - A broker that **removes records on acknowledgement** and is keeping up may barely touch the volume at steady state, so the same shortfall stays hidden until readers fall behind or the process restarts and has to read back. - A design whose primary store is a **remote service** replaces the per-volume allowance with a service's own rate and error behaviour — the shape of the problem is the same, the place you look is not. ## What this is not It is not a volume that has filled up, which has its own symptom and its own remedy; the stem rules that out explicitly. It is not a durability setting, which decides what the write waits *for* rather than how fast the media can absorb it. It is not a question of how large the volume should be, which is arithmetic from the incoming rate. And it is not, at this stage, an alerting question — deciding what to alert on when service times rise is a separate discipline. What it is, is a storage *character* decision made before traffic arrived and never revisited: the volume was chosen by price and size, and nobody wrote down what sustained rate the append path was going to need from it.

  • How would you separate a throughput allowance from a noisy neighbour on the same volume?
    An allowance is flat: the volume tops out at the same figure every time and service times rise the moment demand passes it. A neighbour is not flat — the effective ceiling moves with someone else's schedule, so identical broker load is fast one hour and slow the next. Correlating the broker's own write rate against observed service time over several days usually separates them.
  • Why does giving the record-serving node more memory not fix this?
    Memory helps reads, not the append. Records still have to be written out at the rate they arrive, so if the volume cannot absorb that rate the unwritten work grows regardless of how much is cached. More memory only delays the point at which the write path becomes the constraint, and hides it in the meantime.
  • Does moving to a larger volume reliably make it faster?
    Only where the throughput allowance is derived from size, which is common but not universal. Where speed is a separately purchased dimension, a larger volume is simply a larger volume. Check how the allowance is determined before buying capacity as a performance fix, or you will pay for space and still wait.

saying these in an interview costs you the question

  • Assumes a volume that is not full cannot be the problem.
  • Buys a larger volume for speed without checking how the allowance is set.
  • Blames the network or the clients before examining the write path.
  • Believes a throughput allowance surfaces as an error rather than as waiting.
  • Thinks a shared volume is safe because each workload has its own directory.