A per-IP counter hash table's p99 spikes periodically while mean latency stays flat — how do you confirm resizing is the cause?
answer
- the mean hides it, the tail shows it
- graph capacity next to latency
- measure the growth path directly
- are the gaps constant or stretching
- tables grow but rarely shrink back
basics
~20 sExport the table's entry count and capacity, and emit a timed event on every resize. Each latency spike should land on a capacity doubling, and the gaps between spikes should roughly double as the table grows.
solid answer
~50 sStart by proving correlation rather than arguing it. Export the table's entry count and capacity as gauges, and log each resize with its duration and the number of entries moved. Then overlay the resize timestamps on the p99 chart: if every spike sits on a capacity doubling, and the spike duration scales with the entries moved, you are done. Two further signatures make the case to a skeptical reviewer. First, the interval between spikes roughly doubles each time — the sawtooth stretches out, it does not repeat on a fixed period, which rules out anything driven by a timer such as a scheduled flush or a cron. Second, the spikes end once distinct client addresses stop arriving, because tables grow but almost never shrink. Mean latency stays flat throughout because one operation in hundreds of thousands is slow — exactly the shape a tail metric exposes and an average hides.
go deeper
Know that growing a hash table copies every entry, so one operation can be far slower than the rest, and that an average latency chart will not show it.
Explain why the spikes get rarer and larger as the table grows, and name the instrumentation that proves it: entry count, capacity, and a timed event on the growth path.
Run the correlation before proposing a fix, rule out timer-driven and reclamation causes using the doubling interval, and know that the spikes ending on their own is itself a finding about key-space saturation.
Decide what this deserves: a warm-up exclusion in the SLO, a pre-size, a split table, or a redesign because the underlying defect is an unbounded key space. Be able to defend doing nothing when the transient is genuinely bounded.
## The shape you are looking at A rate limiter keyed by client address inserts a new entry the first time it sees each address. As distinct addresses accumulate, the table crosses its load threshold and doubles, re-placing every stored entry on the unlucky insertion that triggered it. That one operation runs for as long as it takes to walk hundreds of thousands of entries; every other operation that hits the same table behind a lock waits for it. The result on a dashboard is a **sawtooth in the tail metric and nothing at all in the mean**. If one operation in 500,000 takes 40 ms while the rest take 200 ns, the average moves by well under a microsecond — invisible — while p99.9 and max jump by orders of magnitude. "Tail moves, mean doesn't" is itself evidence: it says a tiny number of operations are enormously slow, which is the signature of a rare bulk operation rather than of general slowness. ## Confirming it, not asserting it The goal is a correlation an SRE can check without trusting your intuition. 1. **Export the table's state.** Entry count and capacity as gauges. Capacity is a step function; every step should line up with a spike. 2. **Time the resize itself.** Wrap the growth path so it emits an event carrying duration, entries moved, and old/new capacity. This turns the hypothesis into a measured fact and gives you the spike's magnitude in advance. 3. **Overlay and check three things.** Do spikes coincide with capacity steps? Does spike magnitude scale with entries moved — each spike roughly twice the previous one? Do the gaps between spikes roughly double? 4. **Check which operations are slow.** Only paths that can trigger growth should spike. If read-only lookups spike identically, growth is not your culprit. 5. **Watch for the ending.** Once the population of distinct addresses saturates, growth stops and so do the spikes — without any deploy. That behaviour is very hard to explain with any other cause. ## Ruling out the usual suspects The doubling interval is the discriminator worth leading with, because the competing explanations mostly have *fixed* periods: - **A scheduled task** — a flush, a metrics roll-up, a cron — repeats on a constant interval. Resize spikes do not; they get rarer geometrically. - **Memory reclamation pauses** correlate with allocation volume and are often multi-threaded and process-wide, hitting operations that never touch this table. A resize is single-structure and single-threaded. Note that resizing *feeds* reclamation pressure by discarding a large array each time, so the two can appear together — measure the resize event itself rather than inferring from allocation graphs. - **Lock contention** rises with concurrency, not with table size, and shows up as a broad p99 lift rather than isolated spikes. During a resize the effect is real but it is a *consequence* of the transfer, not an independent cause. - **A downstream dependency** shows correlated latency at that dependency's own metrics; the table's capacity gauge would show no step. ## Explaining the sawtooth to a skeptic The question an SRE will ask is "why is it getting less frequent but worse?" The answer is one sentence: each resize doubles the capacity, so it takes twice as long to fill again, and there are twice as many entries to move when it does. Frequency halves, magnitude doubles, total work per unit of growth stays flat. The chart looks alarming precisely because it is asymptotically well-behaved — that is the amortized bound rendered as a picture. ## Then decide whether to fix it Diagnosis is not automatically a work item. If the address population saturates within minutes of a deploy, the spikes are a startup transient and the correct action may be to exclude the warm-up window from the SLO. If the spikes are recurrent because the key space genuinely churns — addresses arriving and never being reused — the table is growing without bound and the real defect is the missing expiry, not the rehash. If they breach the budget in steady state, the mitigations are pre-sizing from a known address population, splitting the counters across several fixed sub-tables so each transfer moves a fraction of the entries, or bounding the table with eviction. Choosing among those is a design conversation; confirming the cause with the capacity gauge is the part that must come first.
- Why does the interval between spikes double while the spike itself gets worse?Each resize doubles capacity, so twice as many new entries must arrive before the next threshold is crossed, and twice as many entries must be moved when it is. Frequency halves as magnitude doubles — that is the amortized bound drawn as a chart. A fixed-period spike pattern points at a timer-driven cause instead.
- How would you distinguish resize pauses from memory-reclamation pauses?Resize pauses are local to the operation that triggers growth and line up exactly with capacity steps; reclamation pauses correlate with overall allocation rate and typically hit unrelated operations too. Because a resize discards a large array it also *causes* reclamation pressure, so measure the growth path directly rather than inferring from allocation graphs.
- The spikes stop after twenty minutes without any deploy. What does that tell you?The distinct-key population saturated, so the table reached a capacity it never has to exceed and stopped growing. Almost no implementation shrinks on its own, so once growth ends the spikes end permanently. It also reframes the problem as a startup transient, which may warrant excluding the warm-up window from the SLO rather than changing the structure.
- The address population never saturates — spikes keep recurring. What is the real defect?Unbounded growth, not rehashing. A table that accumulates a fresh key for every new client and never releases them will exhaust memory regardless of how the resize is implemented. The fix is expiry or eviction so the entry count reaches a steady state; tuning the growth path only makes the memory exhaustion arrive more quietly.
saying these in an interview costs you the question
- Blames memory reclamation without measuring the growth path
- Expects resize spikes on a fixed periodic schedule
- Looks only at mean latency and sees nothing wrong
- Assumes the table shrinks back when entries are removed
- Jumps to a mitigation before correlating spikes with capacity steps