For a fleet of long-lived services, what would you page a human on: peak footprint, percentage of the limit, or the trend of the post-collection floor?
answer
- separate detection from urgency
- peaks fire on healthy behaviour
- percent of limit arrives too late
- normalise slope to remaining headroom
- trend opens a ticket, exhaustion pages
basics
~20 sUse two signals with different urgencies: the post-collection floor's slope opens a ticket days ahead because it detects retention early, and proximity to the limit pages, because it means failure is imminent. Peak footprint alone pages on healthy churn and should not.
solid answer
~40 sSeparate **detection** from **urgency**. The floor trend — slope of post-collection floors measured at steady state, expressed as a fraction of remaining headroom per day — is the only one of the three that identifies retention while there is still time to act, so it should open work, not wake someone. Proximity to the limit is the page: it fires late and says little about cause, but when it fires, failure is close. Peak footprint is the worst primary signal, because a healthy high-churn service legitimately peaks near whatever room it is given; alerting on it pages on normal behaviour and trains the team to ignore memory alerts. Across a fleet the threshold has to be normalised and the warm-up and deploy windows suppressed, or every rollout pages.
go deeper
Know that memory alerts usually come in two flavours: one that says the service is nearly out of room now, and one that says it is slowly heading there.
Explain why an alert on peak footprint fires on ordinary high-allocation behaviour, and why the post-collection floor is the reading that reflects a real problem.
Show the operational detail: normalised thresholds, suppression of warm-up and deploy windows, a persistence requirement across cycles, and a verdict expressed as projected days of headroom.
Own the trade explicitly — how much false-ticket volume the organisation will accept for earlier detection, how restart policy and deploy cadence affect whether any signal can accumulate, and what the scheme still cannot catch.
## The three candidate signals, honestly compared | signal | what it detects | lead time | false-positive source | |---|---|---|---| | peak footprint | how much room the runtime is allowed to fill | none, and often meaningless | healthy churn, a larger limit, a traffic spike | | fraction of the limit in use | imminent exhaustion | minutes to hours | a service deliberately sized tight | | slope of the post-collection floor | retention itself | days | warm-up, deploys, load-pattern changes | Peak footprint fails as a primary signal for a structural reason: a collector will generally use the room it is given before collecting, so a perfectly healthy service with a high allocation rate sits near its peak much of the time. An alert on that condition is an alert on normal operation, and its real cost is not the page but the habituation — a team that has learned memory alerts are noise will also ignore the one that matters. Fraction-of-limit is a genuine signal, but it answers a different question. It fires when the outcome is nearly determined, and it says nothing about whether the cause is retention, a workload change, or a service that was always sized close to its ceiling and is behaving exactly as designed. As the only signal, it guarantees that every slow leak is discovered at the worst possible moment. The floor slope is the only one that detects the mechanism. Because retention is by definition a growing reachable set, and the floor is the reachable set, a fit over floors is a direct measurement of the thing you care about, with days of warning in hand. ## Designing the floor-trend signal for a fleet A fleet makes a single absolute threshold useless — "floor climbing more than 100 MB a day" is catastrophic for a small service and irrelevant for a large one. The design that survives contact: 1. **Normalise the slope.** Express it as remaining headroom consumed per day, or equivalently as projected days until the floor reaches the limit. A fleet-wide rule of the form "fewer than N days of projected headroom" compares services of different sizes fairly. 2. **Derive the threshold from measured noise.** Fit slopes across services already believed healthy, look at the spread, and put the threshold outside it. A threshold chosen by intuition either never fires or fires constantly. 3. **Suppress the windows where the signal is known to lie.** Exclude the period after a restart or a deploy, when legitimate fill is under way, and require the window to span whole traffic cycles with like phases compared. 4. **Require persistence.** Demand that the slope hold across several cycles before it counts. This costs nothing — the signal already has days of lead time — and removes most of the remaining false positives. ## Routing: what each signal is allowed to do - The **floor trend** opens a ticket with the measured rate and the projected date. Paging on a condition whose consequence is a week away is how a fleet teaches people to silence pages. - **Proximity to the limit** pages, because the window for action is short. - **Peak footprint**, if kept at all, is a dashboard line for capacity conversations, not an alert. Two further decisions belong to whoever owns this: - **Automatic restarts are a mitigation and an evidence destroyer.** A fleet that recycles instances on memory pressure will be stable and will never fix a leak, because every series is truncated before it becomes conclusive. If you adopt it, keep at least one long-running instance outside the policy so the evidence exists. - **Deploy cadence interacts with detectability.** If nothing runs longer than the time a leak needs to become visible, the signal cannot fire. Either keep long-running instances, or reconstruct a series across instances by comparing floors at equal age since start. ## What makes this a judgment call rather than a fact There is no universally right threshold, and the cost asymmetry is what decides it. A fleet where memory exhaustion means a brief restart with no data loss can afford a lax threshold and a late page; one where exhaustion means a long recovery or a user-visible outage should trade more false tickets for earlier detection. The defensible answer names both signals, assigns each a routing consistent with its lead time, states how the threshold was derived from observed variance, and admits what the scheme still cannot catch — a leak faster than the persistence requirement, and a service that redeploys before its own signal can accumulate.
- Your fleet already restarts instances automatically under memory pressure. What does that do to this scheme?It keeps the fleet up and makes the leak undiagnosable. Every restart truncates the floor series before the slope becomes conclusive, so the trend signal never accumulates enough history to fire. Keep the mitigation, but exclude a small number of instances from it and alert on those, so somewhere in the fleet the evidence is allowed to reach the threshold.
- How would you turn a measured floor slope into the number a ticket is triaged on?Project it: take the remaining headroom between the current floor and the point at which the service fails, divide by the slope, and report days remaining. A floor rising about 90 MB a day with roughly 1.4 GB of room left is about two weeks out. Days-to-exhaustion is comparable across a fleet and orders the queue by itself, which a raw megabytes-per-day figure cannot do.
- What does this alerting scheme still fail to catch?Retention faster than the persistence requirement — a service that exhausts memory in hours will hit the limit before a multi-cycle trend has confirmed anything, which is what the proximity page is for. It also misses anything on services that redeploy more often than the detection window, and it says nothing about a steady live set that is simply too large for the room available.
saying these in an interview costs you the question
- Alerts on peak footprint and calls healthy churn an incident
- Uses percentage of the limit as the only memory signal
- Pages a human on a trend whose consequence is a week away
- Applies one absolute megabytes-per-day threshold across a whole fleet
- Treats automatic restarts as a fix rather than a mitigation
- Leaves deploy and warm-up windows unsuppressed, then silences the alert