skip to content

Besides deleting a key when it is next accessed, Redis runs a background cycle over keys that carry a time-to-live. Walk through how that cycle decides how much work to do, and why its guarantee is probabilistic rather than exact.

level: middleimportance: must knowfreq 52%

answer

  1. 20 samples, repeat if >25% expired
  2. Only the expires table is sampled
  3. Fast pass ~1ms; slow pass at hz, ~25% CPU
  4. expired_time_cap_reached_count = behind
  5. LATENCY expire-cycle, never SLOWLOG

basics

~20 s

It repeatedly samples about 20 random keys from the table of keys with a TTL, deletes the expired ones, and loops again if more than 25% of the sample was expired. Time-boxed per run, so it bounds stale keys statistically, not exactly.

solid answer

~50 s

The cycle runs in two flavours. A **fast** pass runs before each event-loop iteration with a budget of about a millisecond; a **slow** pass runs from the server cron at `hz` times per second (default 10) and may use up to about 25% of CPU time. Each pass, per database, samples roughly 20 random keys from the *expires* table — only keys that have a TTL, so the cost is independent of total keyspace size — deletes those past their deadline, and if more than 25% of the sample was expired it immediately repeats; otherwise it moves on. That loop-while-hot rule is what makes the effort proportional to how much garbage exists. The consequence is a statistical bound: at steady state, roughly under a quarter of keys with a TTL are expired-but-not-yet-freed. Aggressiveness is tunable via `hz` and `active-expire-effort` (1–10). Observability: `expired_keys` and `expired_time_cap_reached_count` in INFO stats, and the `expire-cycle` event in LATENCY LATEST.

code

text · 11 lines
text
# redis.conf
hz 10                      # cron beats per second (slow pass)
active-expire-effort 1     # 1..10, higher = more aggressive, more CPU
lazyfree-lazy-expire no    # yes = free big values on a background thread

# observe
INFO stats
#  expired_keys:184213
#  expired_time_cap_reached_count:0   # >0 and rising = cycle is saturated
LATENCY LATEST
#  1) 1) "expire-cycle"  ... max latency ms

go deeper

for a junior

Know that a background job samples keys with a TTL and deletes expired ones, and that this is why cold keys eventually disappear.

for a middle

Give the numbers and the feedback rule: ~20 samples, repeat while over 25% expired, fast pass versus cron-driven slow pass, expires table only.

for a senior

Add the operational read: expired_time_cap_reached_count, the expire-cycle latency event, replication traffic from propagated DELs, and lazyfree-lazy-expire for large values.

for a principal

Discuss it as a control loop: bounded CPU share versus memory-waste bound, and when you would rather change the workload's TTL distribution than crank effort.

## Why sampling at all Redis executes commands on one thread. An exact sweep would mean either a per-key timer or an ordered structure of deadlines that must be walked to the current time — both add work proportional to the number of expiring keys at unpredictable moments, on the thread that also serves user traffic. Instead Redis uses randomized sampling with a feedback rule, which makes the cost proportional to how much actual garbage exists and hard-caps the time spent per run. ## The algorithm The cycle only ever looks at the *expires* table — the per-database table of keys that carry a TTL. That matters: a database with 50 million keys but 1,000 TTLs costs the same as one with 1,000 keys total. Per iteration, per database: 1. Sample about 20 random keys from the expires table. 2. Delete the ones whose deadline has passed; each deletion propagates a DEL (or UNLINK when `lazyfree-lazy-expire yes`) to replicas and the AOF, and fires the `expired` notification. 3. If more than 25% of that sample was expired, assume the table is still dirty and immediately repeat; otherwise stop for this database. 4. Stop early whenever the run's time budget is exhausted. The 25% rule is the whole design in one line: while garbage density is high the cycle keeps grinding; once density falls below a quarter it yields the thread back to clients. ## Fast and slow passes The **slow** pass runs from `serverCron`, i.e. `hz` times a second (default 10). It is allowed a percentage of CPU time — on the order of 25% of the cron interval — so it can do real work but cannot monopolize the loop. The **fast** pass runs in the event loop's pre-sleep step with a much smaller budget, roughly a millisecond, and with a minimum gap between runs; its job is to keep latency-sensitive reclamation snappy between cron beats. When a run hits its cap rather than reaching the density threshold, Redis increments `expired_time_cap_reached_count` — a direct signal that expiry work is arriving faster than the cycle is permitted to clear it. ## The probabilistic guarantee Because the loop continues while more than a quarter of samples are expired, the steady-state expectation is that fewer than ~25% of keys carrying a TTL are expired-but-unreclaimed at any instant. This is a bound on *memory waste*, not on correctness: correctness comes from the access-path check, which never serves an expired value. It is also an expectation, not a hard limit — a burst of simultaneous deadlines will exceed it briefly. ## Tuning knobs and what they cost - `hz` (default 10) raises cron frequency: more frequent, smaller expiry passes, at some baseline CPU. - `active-expire-effort` (1–10, default 1) scales the sampling size, the density threshold and the time budget. Raising it reclaims memory sooner at the price of more CPU stolen from command processing and higher tail latency. - `lazyfree-lazy-expire yes` moves the *freeing* of large collection values to a background thread, so expiring a hash with millions of fields does not block the loop — the unlink from the keyspace is still done inline. ## Diagnosing it `INFO stats` gives `expired_keys` (a rate to graph) and `expired_time_cap_reached_count` (chronic saturation if it climbs). `LATENCY LATEST` reports an `expire-cycle` event when a pass exceeded the latency threshold — essential because expiry is not a command, so it never appears in SLOWLOG. If memory sits higher than the TTL profile predicts while `expired_time_cap_reached_count` grows, the cycle is behind: raise effort, raise `hz`, or reduce the number of simultaneously expiring keys.

  • Hundreds of thousands of keys share the same deadline and all expire in the same second. What happens inside the server?
    Samples come back almost entirely expired, so the cycle keeps looping until it hits its time cap, and expired_time_cap_reached_count rises. Each deletion also emits a DEL to replicas and the AOF, so replication traffic spikes. If the values are large collections, freeing them inline blocks the single thread, which shows up as a latency spike under the expire-cycle event rather than in SLOWLOG; enabling lazyfree-lazy-expire moves that freeing off the loop.
  • Does the cycle get slower as the total keyspace grows?
    No. It samples only from the table of keys that carry a TTL, so its cost tracks the number of volatile keys and how many of them are stale, not the total key count. A huge keyspace with few TTLs is cheap to police.

saying these in an interview costs you the question

  • Describing the cycle as scanning the whole keyspace
  • Saying it guarantees expired keys are gone within a fixed time
  • Believing raising active-expire-effort is free
  • Expecting expiry pauses to appear in SLOWLOG
  • Confusing this cycle with the maxmemory eviction sampler

context