How long must you watch a freshly restarted long-lived service before a rising post-collection floor entitles you to call it retention?
answer
- restarting empties everything on purpose
- caches and pools fill legitimately
- the slowest filler sets the start
- compare the same hour across days
- warm-up flattens, retention does not
basics
~20 sLong enough that every legitimate slow fill has plateaued — bounded caches, pools, lazily built structures — and then across several complete traffic cycles, comparing floors at the same phase of each cycle. Retention does not plateau; warm-up does.
solid answer
~40 sA restarted service's floor is *supposed* to rise for a while. Caches fill toward their eviction ceiling, connection and worker pools grow to their steady count, lazily initialised structures get built on first use. All of that raises the live set legitimately and then stops. So the observation window has two requirements. First, it must start after the slowest legitimate filling process has settled, which is set by things like a cache's eviction horizon rather than by the clock. Second, it must span several complete traffic cycles, with floors compared at the **same phase** of each cycle — a busy afternoon holds more in flight than a quiet night, and comparing across phases manufactures a trend that is not there. The decisive distinction is shape: warm-up decelerates into a plateau, retention keeps its slope.
go deeper
Understand that memory climbing right after a restart is normal: caches and pools are filling up to the size they are meant to be.
Explain which structures legitimately raise the live set during warm-up, and why the distinguishing feature of retention is that the climb never flattens.
Show the operating discipline: derive the warm-up cut-off from the slowest filling structure, compare floors at the same phase across days, and state the verdict as a rate with the window it was measured over.
Own the consequence for release cadence: if everything redeploys daily, no instance lives long enough to reveal a slow leak, so decide deliberately how long-running evidence gets collected.
## Why a rising floor is expected at first Restarting a service empties every structure it maintains. What follows is a genuine increase in the live set that has nothing to do with a defect: - **Caches fill.** A cache bounded by entry count or by an eviction horizon climbs toward that bound at whatever rate distinct keys arrive. If the horizon is long, the climb is long. - **Pools grow.** Connection pools, worker pools and buffer pools expand toward their configured or demand-driven steady size, and each element is reachable by design. - **Lazily built structures appear.** Anything constructed on first use — compiled patterns, prepared statements, per-endpoint metadata, internal maps keyed by whatever the traffic actually contains — enters the live set only when the traffic first exercises that path. - **Rare paths arrive late.** A structure built only by a nightly job or by one seldom-used endpoint raises the floor days after start-up, and it is not retention. The naive conclusion — "we restarted it four hours ago and memory is already up 300 MB, it is leaking" — is the single most common false positive in this material. ## The two requirements on the window 1. **Start after the slowest legitimate fill has settled.** The right number is not a habit like "skip the first ten minutes"; it is derived from the system. Ask what the longest-lived thing the service caches is, and how long a key has to be absent before it is evicted. The warm-up segment ends when that process has reached its bound, which may be minutes or may be most of a day. 2. **Span several complete traffic cycles, comparing like phases.** Most services have a daily rhythm, and the live set genuinely differs between peak and trough because in-flight work, per-connection state and per-session state are all reachable. Comparing this afternoon's floor with last night's produces a slope out of nothing. Compare the same hour across days. ## Deciding between fill and retention The test is the second derivative, stated in plain terms: **does the climb flatten?** | observation across the window | reading | |---|---| | rises, then decelerates into a level plateau | warm-up or cache fill, now at steady state | | rises at a roughly constant rate that survives several cycles | retention | | rises, plateaus, then rises again after a deploy or a new traffic pattern | a new steady state, then watch again | | rises and falls with the daily rhythm around a level mean | in-flight and per-session state, not retention | A practical routine on a fourteen-day chart: drop the warm-up segment, take one floor at the same hour on each remaining day, and fit a line. If the slope is inside the day-to-day variation you already see in the flat region, you are not entitled to call it retention yet; if the line holds its slope across a week, you are. Reporting it as a rate — "the floor is climbing about 90 MB a day, so from 1.2 GB it reached 2.1 GB in thirteen days" — is what makes the claim checkable by someone else. ## Complications worth naming - **Retention proportional to traffic.** If a service retains a little per connection, and connection counts follow the daily rhythm, the floor rises in a staircase whose steps match the load pattern. Same-phase comparison across days is exactly what separates that from healthy diurnal variation: the healthy one returns to the same nightly floor, the retaining one does not. - **Deploys reset the clock.** A deploy restarts warm-up, and a fleet that deploys daily may never run long enough to expose a slow leak. That is a real reason a leak survives for months, and a reason to keep one long-running instance out of the deploy rotation when investigating. - **A bounded cache still counts as memory.** Concluding "it plateaued, so there is nothing to do" can be wrong for capacity reasons even when it is right about retention; the plateau level itself may be too high for the room available. - **Do not restart to test.** A restart destroys the very series you need, and buys only as much time as the retention takes to rebuild. The short answer an interviewer wants: you are entitled to call it retention when the climb has outlived every process that legitimately fills memory, has survived several full traffic cycles measured at the same phase, and has not decelerated.
- The floor climbs during the working day and returns to the same level every night. Retention or not?Not retention. Returning to the same nightly level means nothing is accumulating across cycles; the daytime rise is in-flight work, per-connection and per-session state that is released as load falls. The signal to watch is the trough-to-trough trend. If successive nightly floors start creeping up while the daily shape stays the same, that is retention riding on top of a healthy rhythm.
- Your fleet redeploys every day, so nothing runs longer than 24 hours. How do you detect a slow leak at all?Keep at least one instance out of the deploy rotation long enough to accumulate history, or reconstruct a series across instances by comparing floors at equal age since start — every instance at hour twelve, then hour twenty-four. Frequent deploys do not remove a leak; they hide it by restarting before the floor gets high enough to hurt.
saying these in an interview costs you the question
- Declares a leak from a few hours of post-restart growth
- Compares a busy afternoon floor with a quiet overnight floor
- Assumes warm-up is always over within minutes of start-up
- Treats a cache filling toward its bound as a defect
- Restarts the service to test the theory, destroying the series