A Go API's per-client rate-limiter map grows without bound as callers come and go. How do you fix it?
answer
- a cache with no eviction policy
- who chooses the keys here?
- entries need a timestamp and a sweeper
- evicting early gives the burst back
- deleting keys does not shrink a map
basics
~20 sEvery unseen key adds an entry nothing removes, and on a public tier the key space is caller-controlled. Stamp entries with a lastSeen time and let a janitor goroutine delete ones idle longer than a full bucket refill.
solid answer
~40 sThe map is a cache with no eviction policy, so its size tracks distinct keys ever seen, not active clients. On a free-tier public API that count is caller-controlled: unauthenticated per-address keys, multiplied again by a key that combines client and route. Make entries expire — record a `lastSeen` timestamp on every hit, and run one janitor goroutine on a `time.NewTicker` that takes the same mutex and deletes entries idle longer than a bucket takes to refill completely, stopping when the server does. That threshold is correctness, not memory: evicting a caller who is still throttled hands them a fresh full burst. If the key space is hostile, cap the map with a fixed-size LRU instead. Confirm the diagnosis on a heap profile, where the entries dominate `inuse_space`.
code
go · 23 linesfunc (l *limiters) sweep(idle time.Duration) {
cutoff := time.Now().Add(-idle)
l.mu.Lock()
defer l.mu.Unlock()
for k, v := range l.m {
if v.lastSeen.Before(cutoff) {
delete(l.m, k)
}
}
}
func (l *limiters) janitor(stop <-chan struct{}, every, idle time.Duration) {
t := time.NewTicker(every)
defer t.Stop()
for {
select {
case <-t.C:
l.sweep(idle)
case <-stop:
return
}
}
}go deeper
Understand that a map entry stays alive as long as the map does — the garbage collector cannot remove something the map still references. Know that something has to delete keys explicitly.
Explain the mechanics: a lastSeen timestamp updated under the same lock as the lookup, one janitor goroutine on a ticker, deleting during a range being legal, and stopping that goroutine cleanly.
Show that you would confirm it on a heap profile's inuse_space before changing code, and that you understand the expiry threshold is a limit-correctness decision, not just a memory dial. Knowing that a Go map does not shrink on delete is what separates a real fix from a partial one.
Own the cardinality itself: whether anonymous callers get their own bucket at all, whether the key space is caller-controlled, and what fixed memory budget the limiter gets per replica so this can never be the reason the service dies.
## Why it grows A per-client limiter map is a cache whose insert path is driven by strangers. Every key the middleware has never seen allocates an entry, and nothing in the request path ever removes one. Its steady-state size is therefore the count of distinct keys ever observed over the process's lifetime, which for a long-lived server is unrelated to how many clients are active right now. On a free public tier the growth is worse than linear in customers, for three reasons. The key is often derived from the caller's address rather than a credential, and addresses are plentiful — an IPv6 client effectively has as many as it wants. If a proxy header is trusted for the key, the caller writes the key directly. And a key that combines client and route multiplies the cardinality by the size of your route set, so a scanner walking your API creates an entry per path it touches. Each entry is small — a struct, a limiter, a map slot — which is exactly why this is a slow leak rather than an immediate crash. It shows up as memory that climbs for days, flattens only when the process restarts, and eventually hits the container limit. ## Confirming it Take a heap profile — `/debug/pprof/heap` if `net/http/pprof` is registered, or a profile written from a load test — and read `inuse_space`, the live-memory view rather than the cumulative `alloc_space` one. A leaking limiter map is unmistakable there: the allocation sites are the entry constructor and the map's own growth, and the retained size climbs between two profiles taken hours apart. Comparing two profiles is what turns "memory is high" into "this map is why". ## The expiry fix Stamp each entry with a `lastSeen` time on every request that touches it — you are already holding the lock for the lookup, so it costs nothing. Then run one janitor goroutine that ticks on a `time.NewTicker`, takes the same mutex, ranges the map and deletes entries older than the cutoff. Deleting keys during a `range` over a map is explicitly allowed in Go, so the sweep is a plain loop. Give the janitor a stop channel and select on it, so the goroutine ends with the server rather than outliving it in tests. The threshold is a correctness decision, not a memory one. If you evict a caller who is currently being throttled, the next request recreates the entry with a full bucket — you have handed them their burst back, and a client that paces itself just above your eviction window is never limited at all. So the idle threshold must be at least the time it takes an empty bucket to refill to its full burst, plus margin. Sweep frequency is the memory-versus-CPU dial and can be much longer than the tick you would guess: sweeping every few minutes is usually plenty, since the whole point is to bound growth, not to be exact. ## The bounded-size fix Expiry bounds the map by *idle time*, which still lets a burst of a million distinct keys in one minute cost you a million entries. If the key space is caller-controlled, bound it by *size* instead: a fixed-capacity LRU, sized from the memory you are willing to spend, evicting the least recently used entry on insert. Callers that fall out of the cache land on a shared fallback limiter that is stricter than the per-client one, which is a defensible answer for anonymous traffic. Reducing cardinality is the other half: require an API key so the key space equals your customer list, normalise addresses with `net.SplitHostPort` and group IPv6 by prefix, and key on the client alone rather than client-plus-route unless you actually need per-route quotas. ## The trap after the fix A team frequently deletes the stale keys, watches the map's length drop, and finds resident memory unchanged. Go's map does not return its bucket memory when you delete keys — a map that once held ten million entries keeps a table sized for ten million even when it holds ten. The values are collected, so most of the bytes come back, but the table does not shrink. After a genuine cardinality spike, the remedy is to rebuild: allocate a fresh map, copy the surviving entries in, and let the old one be collected. Doing that inside the janitor when the live count is a small fraction of the map's capacity is a few lines and closes the loop. Finally, none of this is a distributed limit. Each process has its own map and its own buckets, so the same expiry and sizing decisions repeat per replica, and the memory you are budgeting is per replica too.
- Why is evicting entries too eagerly a correctness bug rather than just churn?An evicted caller's next request recreates the entry with a full bucket, so the burst is handed back. A client that paces itself just above the eviction window is then never limited. The idle threshold has to be at least as long as it takes an empty bucket to refill to its full burst, with margin.
- Idle expiry still lets a burst of a million distinct keys cost a million entries. What bounds it regardless of traffic?A fixed-capacity LRU sized from the memory you are willing to spend: insertion evicts the least recently used entry, so memory is a constant you choose rather than a function of the caller's imagination. Callers not in the cache fall back to a shared, stricter limiter. Pair it with lower cardinality — an API key instead of an address, client rather than client-plus-route.
- You deleted millions of stale keys, the map's length dropped, and resident memory did not. Why?Go's map does not shrink its bucket table when keys are deleted; it was sized for the peak and stays there. The entry values are collected, so most bytes are reclaimed, but the table is not. After a large cardinality spike, rebuild the map — allocate a fresh one, copy the live entries in, and drop the reference to the old one.
- How do you keep the janitor from becoming its own leak?One janitor goroutine for the whole map, never one per client, and give it a stop channel it selects on alongside the ticker so it exits when the server shuts down. Call `Stop` on the ticker with a defer. Long-lived tests otherwise accumulate janitors, which shows up as a rising goroutine count in the goroutine profile.
It is a guest list that only ever gains names. Without an ink that fades, the book grows until the shelf breaks — and tearing out a page too early lets that guest walk in fresh.
saying these in an interview costs you the question
- Expects the garbage collector to reclaim entries still in the map
- Sweeps on every request instead of on a ticker
- Picks an idle threshold shorter than a full bucket refill
- Keys the map on an unnormalised caller-supplied header
- Starts one expiry goroutine per client entry
- Assumes deleting keys returns the map's memory