Your log-shipping service's memory climbs steadily for hours and then dies; how do you show it is a rate mismatch, not a leak?
answer
- rates first, memory last
- drop input below drain rate
- gap integral equals retained count
- same slope after every restart
- intended owner, many identical items
basics
~20 sCompare rates and run a drain test. A rate mismatch retains items that belong in a queue and recovers when the input drops below the drain rate; a leak retains objects nothing needs and recovers only on restart.
solid answer
~50 sStart from the rates rather than the memory. Export lines read per second and lines written per second: a sustained gap between them is the mismatch, and the integral of that gap over the window should equal the retained item count — if those two numbers agree, the diagnosis is done. Confirm with a drain test: cut the input below the archival writer's rate and watch. A backlog flattens and then falls as the queue drains; a leak keeps climbing because nothing releases what it holds. Two more signals: the retained set is one owner holding many identical pipeline items, and growth is linear in elapsed time rather than in a particular code path's call count. The repair is not a smaller allocation — it is a bound on the retaining stage so the same condition fails early and visibly.
code
pseudocode · 9 lines// evidence, gathered over a four hour window
readRate = 12000 // lines per second into the pipeline
writeRate = 11500 // lines per second out to the archive
deficit = readRate - writeRate // 500 lines per second
predictedDepth = deficit * 4 * 3600 // 7200000 lines
// if measured retained items ~= predictedDepth, the backlog explains the memory
// then drop readRate below writeRate and confirm depth fallsgo deeper
Know that a queue holding work is not the same thing as a leak, even though both show rising memory, and that the difference is whether the memory can come back on its own.
Explain the arithmetic link: the gap between arrival and drain rates, integrated over time, should equal the retained item count. If it does, the backlog explains the memory.
Run the diagnosis in order: rates first, drain test second, retained-object ownership third. Then repair with a bound plus a rate change, and leave depth and delivery-age alerting behind so the next occurrence is triaged in minutes.
Decide what the service should do when the rates cannot be reconciled. Keeping everything means staleness and eventual collapse; bounding means an explicit, visible loss. That call belongs to whoever owns what the data is for.
## Two shapes of growth Both a leak and an unbounded hold look the same from the outside: memory rises steadily, the process dies, a restart buys another few hours. They are different problems and the diagnosis has to separate them before anyone reaches for a fix. | | Object leak | Rate mismatch in a hold | |---|---|---| | What is retained | Objects nothing will use again | Items still waiting to be processed | | Why they are retained | An owner keeps a reference by mistake | Retention is the configured behaviour | | Grows with | Operations performed on a path | Elapsed time, at arrival minus drain rate | | Effect of reducing input | Growth slows but never reverses | Depth falls as the queue drains | | Correct repair | Release the reference | Bound the stage and change a rate | The distinguishing property is **reversibility**. A queue's contents are supposed to leave. A leak's are not. ## The drain test The single most decisive experiment: reduce the input rate below the drain rate and watch memory for a few minutes. 1. Cut the reader's rate, pause it, or route a fraction of the input elsewhere. 2. Watch retained memory and, if it is exported, the hold's depth. 3. A mismatch shows a plateau followed by a decline as the archival writer clears the backlog. The decline rate is the writer's spare capacity, and it is a second measurement of the same number. 4. A leak shows a plateau at best. Nothing comes back, because nothing is waiting to be consumed. If you can run this test, you rarely need anything else. The reason to know the other signals is that on many systems you cannot throttle production input on demand. ## Confirming with rates, not with memory size Memory is the last observable in this chain and the least informative. The rates are the first: - Export **lines read per second** at the source and **lines written per second** at the archival writer. - Plot them together. A leak leaves the two curves on top of each other; a mismatch shows a persistent gap. - Integrate the gap over the incident window. At a deficit of 500 lines per second over four hours that is 7.2 million lines. If the retained item count is near that number, the retained memory is explained entirely by the backlog and there is nothing else to find. This is also the evidence that survives a restart. Growth that resumes at the **same slope** every time, from a clean process, in proportion to the rate gap, is structural. A leak triggered by a rare input would restart with a different slope depending on whether that input recurred. ## Where the items are held Inspecting what dominates retained memory adds a third, independent signal. A mismatch shows **one owner holding many equal-shaped items of the pipeline's own type**, in a queue-shaped structure, all of them reachable by design. A leak more often shows an owner that has no business retaining anything — a registry, a cache with no eviction, a listener list that only ever grows — and the retained objects are frequently unrelated to the stage that allocated them. Be careful with one trap: item type alone proves nothing, because a leak of pipeline items is entirely possible. It is the combination of an intended owner, a rate gap that accounts for the count, and reversibility under the drain test that closes the case. ## The stall case A writer that is fast on average can still ratchet. Suppose the archival writer stalls for **10 minutes each hour** while the reader runs at 12,000 lines per second. Over the hour, 43.2 million lines arrive and must all be written in the remaining 3,000 seconds, so the writer needs `43,200,000 / 3,000 = 14,400 lines per second` of real capacity to end each hour empty. At 12,000 per second — the reader's own rate, which sounds sufficient — it ends every hour with a residue, and the residue compounds. "The average is fine" is therefore not a defence; the question is whether the backlog built during each stall is cleared before the next one starts. ## What the fix is and is not Raising the memory limit moves the failure later and makes it worse, because a larger backlog also means a larger delivery age and more data lost when it finally dies. The repair has two parts: 1. **Bound the retaining stage** so the mismatch produces an early, explicit signal at the stage boundary instead of an allocation failure hours later. What should happen when that bound is reached is a separate design decision with its own trade-offs. 2. **Close the rate gap**: make the archival writer faster, reduce what is read, or accept that some items will not be kept. And the durable part of the fix is the instrumentation: depth and delivery age per retaining stage, alerted on a depth that never returns to baseline. That signal turns this four-hour mystery into a two-minute triage the next time it happens.
- The writer is fast on average but stalls for ten minutes every hour. Does that change the diagnosis?No, it sharpens it. With 12,000 lines per second arriving, an hour brings 43.2 million lines that must clear in the 3,000 working seconds, needing 14,400 per second. A writer matching only the reader's average rate leaves a residue after every stall, and the residue compounds hour after hour.
- You cannot throttle production input to run the drain test. What is the next best evidence?Read and write rates plotted together over the incident, with the integral of their gap compared against the retained item count. Agreement between those two independently measured numbers is strong, and it is reinforced by growth that resumes at the same slope after every restart.
- What should have alerted before the process died?Depth and delivery age per retaining stage, with the alert on depth that fails to return to its baseline across successive periods rather than on any absolute value. Bursts legitimately push depth up; only a backlog that never recovers indicates the rates themselves are wrong.
saying these in an interview costs you the question
- Steadily rising memory always means a leak
- Raising the memory limit is a fix rather than a delay
- Retained items of the pipeline's own type prove it is not a leak
- A writer whose average rate matches the reader cannot build a backlog
- Restarting on a schedule is an acceptable resolution