skip to content

A caller pipelines a million writes to a volatile tier in one run — where does the memory go, and what bounds a safe run size?

level: seniorimportance: should knowfreq 47%

answer

  1. unread replies sit somewhere
  2. both ends hold them
  3. bounded by memory, not the protocol
  4. size the chunk in bytes
  5. chunk and drain, then repeat

basics

~20 s

Unread replies pile up in the reply buffer at each end — the server's for what it has produced, the caller's for what it has received and not read. Run size is bounded by memory, not by the protocol, so size the chunk in bytes of expected reply.

solid answer

~50 s

A run is only fast because the caller is not reading yet, and that is exactly what makes it expensive: every reply the server has produced and the caller has not drained has to be held somewhere. It is held in **the reply buffer at each end**. Nothing in the protocol caps a run, so memory does, and the honest bound is bytes of expected reply rather than a count of operations — ten thousand counter updates and a thousand whole-collection reads are wildly different bills at the same operation count. What happens when the bound is crossed **varies by store**: some cap what one connection's unread replies may hold and close the connection, others simply consume memory until the tier's own ceiling behaviour takes over. Chunk the work into runs you have sized, and drain each one before sending the next.

go deeper

for a junior

The idea to hold on to: replies you have not read still have to be stored somewhere, at both ends of the connection. That is why you cannot send unlimited operations in one go.

for a middle

Explain the buffering at each end and why that makes memory, not the protocol, the limit on a run. Then give the sizing rule in bytes of expected reply and show the chunk-and-drain loop.

for a senior

Bring the operational signature: which symptom you would see, that stores differ between dropping the connection and consuming memory, and that a long run also holds a pooled connection and loses every outcome at once if it breaks.

for a principal

Turn it into a capacity number. The bound is per connection, so multiply by concurrency to get the memory the estate actually holds in flight, and publish the figure so teams are not each guessing a run size.

## Why a run costs memory at all The mechanism is defined by not waiting: the caller writes operation after operation and reads nothing back until it has sent them. That is where the latency saving comes from, and it is also the bill. Every reply the server has produced and the caller has not yet consumed exists somewhere in memory. There are three places, and it is worth naming all three: - **The caller's outgoing buffer** — the operations written but not yet flushed onto the network. - **The server's reply buffer for that connection** — replies the server has produced because it executed the operations, which it cannot hand over until the caller reads. - **The caller's incoming reply buffer** — replies that have arrived and not been parsed and matched to their operations. A million operations in one run means all three are being asked to hold a share of a million replies. The protocol imposes no limit on any of it; memory does. ## Sizing in bytes, not in operations The common mistake is an estate-wide rule of the form "batch in runs of one thousand". Operation count is the wrong unit because reply size varies by orders of magnitude: | Run | Operations | Rough reply size each | Unread reply bytes | |---|---|---|---| | counter updates | 10,000 | tens of bytes | under a megabyte | | small value reads | 10,000 | a few hundred bytes | a few megabytes | | whole-collection reads | 1,000 | tens of kilobytes or more | tens of megabytes | The third row has one tenth the operations and the largest bill by far. The rule that survives contact with real workloads is: **bound a run by the bytes of reply it is expected to produce**, pick a figure you are willing to hold at both ends on every concurrent connection doing this, and derive the operation count from it per call site. The multiplier matters too. A bound is per run, but the memory is per run **times concurrent connections doing it**. Fifty workers each holding a ten-megabyte run is half a gigabyte of reply data in flight across the tier. ## What happens when you cross the line This is where stores in this class genuinely differ, and asserting one behaviour as the model is wrong: - Some stores cap what a single connection's unread replies may occupy and **close the connection** when it is exceeded — protecting the tier, at the cost of the run's outcomes. - Others simply hold the data and let it count against the component's own memory ceiling, so the symptom is the tier's ceiling behaviour firing rather than a connection error. - Network flow control also participates: once the socket buffers fill, the server may stop making progress on that connection, which converts a memory problem into a stall. The operational signature differs accordingly — a dropped connection in one place, a memory alert in another — which is why a run size that has been safe for years against one tier is not evidence that it is safe against a different one. ## The other costs of a long run 1. **All outcomes arrive at the end.** A run's replies are useful only once drained, so a connection failure mid-run loses the outcome of every operation in it at once, while their effects remain applied. 2. **The connection slot is held for the whole run.** From the first send to the last reply read, that pooled connection is unavailable to anyone else, which matters more the larger the run. 3. **The server is doing your work for a while.** A very long run is a lot of consecutive execution; how much that delays other callers depends on the execution model — a store that executes one operation at a time and a store that serves requests from a thread pool spread the impact differently — so the size that is polite on one is not automatically polite on another. ## The practical shape Chunk and drain: split the million writes into runs sized by expected reply bytes, send a chunk, read its replies, check them, send the next. This keeps the latency win — one round-trip time per chunk instead of per operation, so a million operations in chunks of a thousand still costs a thousand trips instead of a million — while keeping the amount of unread data bounded and known. If the workload allows, run a few chunked streams concurrently rather than one enormous run: the same throughput with a fraction of the memory in flight, and a failure costs you one chunk's outcomes rather than the lot.

  • Does chunking into runs of a thousand give up the latency benefit?
    Almost none of it. A million operations sent one at a time pay a million round-trip times; in chunks of a thousand they pay a thousand. The remaining cost is three orders of magnitude below the loop, and in exchange the unread reply data is bounded and a failure costs one chunk's outcomes instead of the whole job.
  • Why is a fixed operation count a poor estate-wide rule for run size?
    Because it ignores reply size. A thousand counter updates and a thousand whole-collection reads are the same count and differ by orders of magnitude in bytes held at both ends. A rule expressed in expected reply bytes adapts to the call site, and the operation count falls out of it.
  • Fifty workers each send a ten-megabyte run at the same time. What is the real number?
    Around half a gigabyte of reply data in flight, held across the server's per-connection reply buffers and the callers' own. Run size is a per-connection bound, so the figure that matters for capacity is the bound multiplied by the number of connections doing it concurrently.

saying these in an interview costs you the question

  • Sizes a run by operation count alone
  • Believes replies are buffered only on the server side
  • Assumes every store pauses rather than dropping the connection
  • Thinks a bigger run is always faster
  • Forgets the bound is per connection, not per service
  • Treats a run size proven on one tier as proven everywhere