skip to content

A deep hold keeps a log-shipping pipeline alive, but archived lines now arrive nearly an hour late — why?

level: middleimportance: should knowfreq 52%

answer

  1. capacity buys delay, not speed
  2. throughput is set elsewhere
  3. depth over drain rate
  4. Little's Law rearranged for wait
  5. size from the freshness budget

basics

~20 s

Queue depth is waiting time. By Little's Law a steady depth divided by the drain rate is how long every arriving item waits, so a hold deep enough to survive a mismatch also delays each line by exactly that much, in order.

solid answer

~40 s

Adding capacity never raises throughput — the slowest stage still sets that — so a hold that stays deep is not catching up, it is storing a backlog. Little's Law gives the cost directly: with a steady depth `L` and a drain rate of `lambda` items per second, the average wait is `L / lambda`. Forty million retained lines draining at 11,500 per second is about 58 minutes, and because the hold delivers in arrival order, the freshest line waits behind everything older. The result is a pipeline that is technically lossless and operationally useless: alerts computed downstream describe the state of the system an hour ago. That is why a hold's capacity should be derived from the freshness the data needs, not from how much memory happens to be spare.

code

pseudocode · 7 lines
pseudocode
drainRate = 11500          // lines per second the archival writer sustains
depth     = 40000000       // lines currently retained in the hold

deliveryAgeSeconds = depth / drainRate    // 3478, about 58 minutes

// sizing the other way, from a 60 second freshness budget:
maxDepth = 60 * drainRate                 // 690000 lines, whatever memory allows

go deeper

for a junior

Hold on to one sentence: items waiting in a queue are items being delayed. A pipeline that never loses anything can still be useless if everything it delivers is old.

for a middle

Be able to turn a depth and a drain rate into a waiting time and back again, and to say why capacity cannot raise throughput when a later stage is the bottleneck.

for a senior

Show that you measure delivery age, not only depth, and that you would derive a stage's capacity from the freshness consumers need, then check it against memory rather than starting from memory.

for a principal

The trade-off you own is which failure the organisation prefers: silent staleness that keeps every item, or a visible, bounded loss that keeps the data current. Say which the data is worth and why.

## Depth is time, not just memory The first thing an engineer notices about a growing hold is that it consumes memory. The second, and the one that decides whether the pipeline is still doing its job, is that **depth is latency**. An item entering a hold that already contains `L` items, served in order at `lambda` items per second, waits `L / lambda` before it is even handed to the next stage. That relation is Little's Law: for a stable system, `L = lambda x W`, where `L` is the average number of items in the system, `lambda` the average arrival (and, in steady state, departure) rate, and `W` the average time an item spends inside. Rearranged, `W = L / lambda`. It assumes nothing about the arrival pattern or the service distribution, which is exactly why it is usable in an interview and on call: measure two of the three, get the third. ## Working the numbers Take the log-shipping pipeline whose archival writer sustains **11,500 lines per second**, with a hold sitting at a steady **40,000,000 lines**: - `W = 40,000,000 / 11,500 = 3,478 seconds`, about **58 minutes**. - Every line written to the archive right now was read from the application about an hour ago. - Halving the depth halves the delay; doubling the drain rate also halves it. Adding capacity changes neither — it only permits a deeper backlog. The same arithmetic run backwards is how the capacity should have been chosen. If downstream consumers need lines no more than **60 seconds** old, then at 11,500 lines per second the hold may not exceed about **690,000 lines** — and that is the capacity to declare, regardless of how much memory would have fitted. ## What the hold did and did not buy | Property | Before adding capacity | After adding a deep hold | |---|---|---| | Throughput | Set by the slowest stage | Unchanged, still the slowest stage | | Memory | Small | Depth times item size | | Delivery age | Roughly the write time | Depth divided by drain rate | | Burst tolerance | Poor | Good, up to the capacity | | Failure under sustained mismatch | Fast and loud | Slow, silent, and stale | The honest summary is that capacity converts a **burst** into a **delay**. That is a genuine win when the burst is bounded and the data tolerates the delay. It is not a win when the mismatch is sustained, because then the depth never drains and the delay is permanent. ## Order makes it worse A hold that preserves arrival order means the newest and most valuable line is behind every older line. For a log-shipping pipeline this inverts the value of the data: the lines an operator wants during an incident are the last ones produced, and those are the last ones delivered. A deep in-order hold therefore delivers **everything stale**, not "most things fresh and a few things late". This is also the reason "just make it bigger" degrades so quietly. The pipeline reports no loss and no error. Its throughput matches its input. Only a measurement of **delivery age** — the time between an item being read and being written — shows the damage, and that measurement usually does not exist until someone asks why a dashboard disagrees with reality. ## Sizing from the freshness budget A capacity that has been thought about satisfies two bounds at once, and the smaller one wins: 1. **The memory bound.** `capacity x bytes per item` must fit in the memory the process can lose without dying. 2. **The age bound.** `capacity / drain rate` must be under the delivery-age budget the consumers actually need. For most pipelines the age bound is far tighter than the memory bound, which is the surprise: the hold you can afford to hold is much larger than the hold you can afford to wait for. When the two bounds conflict with reality — the age bound is small and the mismatch is sustained — the answer is not more capacity but a change to the rates, or an explicit decision about which items may be sacrificed. ## The interview register A candidate who says "we added a buffer and the errors stopped" has described a symptom moving, not a problem solved. The answer that lands names Little's Law or its plain-language equivalent, computes a delivery age from a depth and a drain rate, states that capacity changes latency rather than throughput, and closes on how the capacity should have been chosen: from the freshness the data needs, checked against memory, not the other way round.

  • Does Little's Law still apply while the hold is growing rather than steady?
    It applies to averages over a stable period, so a growing backlog breaks the assumption. Use it on the current depth as a snapshot lower bound instead: an item entering now waits at least depth divided by drain rate, and longer if the depth keeps rising while it queues.
  • Downstream only needs the most recent state, not every line. What does that change?
    It changes what a full hold should sacrifice, not the arithmetic. If only recency matters, retaining a long ordered backlog is pure cost, and a policy that keeps the newest is a better fit. The four-way choice of what to sacrifice is its own subject; the point here is that depth is what makes that choice necessary.
  • Why does delivery age need its own metric when depth is already exported?
    Depth alone cannot be read without the drain rate, which varies. Ten thousand items is a second of waiting or an hour depending on the consumer. Exporting the age between read and write states the harm directly and survives a change in either rate.

The length of a single-file queue tells you how long you will wait, not how fast the counter serves. Joining a longer line never speeds the counter up.

saying these in an interview costs you the question

  • Adding capacity raises throughput as well as absorbing bursts
  • A lossless pipeline is a healthy pipeline whatever the delay
  • Depth matters only for memory, not for latency
  • A deep in-order hold still delivers recent items promptly
  • Capacity should be set from spare memory rather than a freshness budget