skip to content

You run Redis as a cache with the allkeys-lfu policy and want eviction to follow the real access pattern rather than yesterday's. How do lfu-decay-time, lfu-log-factor and maxmemory-samples interact, and how would you tune them for a workload?

level: principalimportance: nice to knowfreq 24%

answer

  1. decay-time = minutes per one point off the counter
  2. decay applied lazily on lookup or sampling
  3. log-factor = dynamic range, decay-time = time horizon
  4. samples = decision quality vs CPU
  5. tune from OBJECT FREQ distribution, one knob at a time

basics

~20 s

lfu-decay-time (minutes) is how fast an untouched counter decays so once-hot keys can fall out; lfu-log-factor sets how many accesses fit on the 0-255 scale; maxmemory-samples sets how many candidates each eviction round inspects. Tune by sampling OBJECT FREQ on known hot and cold keys.

solid answer

~60 s

Without decay, LFU would remember popularity forever and a key that was hot last week would outrank today's traffic permanently. `lfu-decay-time` sets the number of minutes of idleness that subtract one point from the counter, using the 16-bit last-decay timestamp in the object header; the decay is applied lazily, when the key is looked up or sampled for eviction, not by a sweeper. The three settings do different jobs: - `lfu-log-factor` sets the *dynamic range* - how many real accesses map onto 0-255. Too low and every hot key saturates and becomes indistinguishable; too high and everything sits near the initial value 5. - `lfu-decay-time` sets the *time horizon* - how quickly popularity is forgotten. Short means the cache tracks recent traffic and behaves closer to recency; long means it favours long-run popularity. - `maxmemory-samples` sets *decision quality and CPU* per eviction round. Tune empirically: sample `OBJECT FREQ` across known-hot and known-cold keys. If both bands look alike, the scale or the horizon is wrong. Watch hit ratio and `evicted_keys`, and change one knob at a time.

code

text · 9 lines
text
redis-cli INFO stats | grep -E 'keyspace_(hits|misses)|evicted_keys'
redis-cli CONFIG GET lfu-log-factor lfu-decay-time maxmemory-samples

# check separation between known-hot and known-cold keys
redis-cli OBJECT FREQ product:hot:1
redis-cli OBJECT FREQ product:cold:99991

# widen the scale if everything saturates at 255
redis-cli CONFIG SET lfu-log-factor 20

go deeper

for a junior

Know the defaults and their meanings: counters decay over time so old popularity fades, and the sample size affects how carefully a victim is chosen.

for a middle

Separate the three roles - range, horizon, decision quality - and explain that decay is lazy and driven by a per-key timestamp in minutes.

for a senior

Drive a measurement loop: baseline hit ratio and eviction rate, sample OBJECT FREQ for separation, change one knob, and watch latency when raising the sample size.

for a principal

Argue from workload shape and cost: choose the horizon from how fast popularity moves in the domain, state when tuning is the wrong lever versus capacity or partitioning, and defend the defaults when no measured problem exists.

## What decay is for A pure frequency counter has no memory of time. A key hammered during a campaign last month would keep a high counter forever and survive eviction while today's genuinely hot keys are discarded. Redis solves this by decaying counters. The object header under LFU holds a 16-bit last-decay time in minutes alongside the 8-bit counter. `lfu-decay-time` is the number of minutes of elapsed time that removes one point from the counter; the default is 1. The decay is computed lazily: when a key is looked up, or when eviction sampling scores it, Redis computes elapsed minutes since the stored timestamp, divides by `lfu-decay-time`, subtracts that many points (floored at 0) and rewrites the timestamp. Nothing scans the keyspace to age counters, which keeps the cost proportional to traffic rather than to key count. A consequence worth naming: because decay is lazy, a cold key's stored counter is stale until something looks at it - and eviction scoring is one of the things that looks at it, so the decay is applied exactly when it matters. ## Reading the three knobs as three separate axes **Dynamic range - `lfu-log-factor` (default 10).** It controls how many real accesses map onto the 0-255 counter. Raise it and the counter climbs more slowly, so keys with very high traffic remain distinguishable from merely busy ones. Lower it and low-traffic keys separate faster but everything popular saturates at 255 and becomes a tie, which silently degrades eviction to "random among the saturated". **Time horizon - `lfu-decay-time` (default 1 minute).** Short horizons make popularity evaporate quickly, so the policy tracks current traffic and starts to resemble recency-based eviction. Long horizons favour keys that are popular over hours or days and resist eviction during short traffic anomalies, at the price of holding stale winners. **Decision quality - `maxmemory-samples` (default 5).** More samples per eviction round means each victim is closer to the true minimum-frequency key, paid for in CPU on the single command-execution thread while eviction is active. ## A tuning method 1. Establish the baseline: `INFO stats` for `keyspace_hits`, `keyspace_misses`, `evicted_keys`, and `INFO memory` for `used_memory` versus `maxmemory`. Hit ratio is the outcome metric; everything else is a lever. 2. Sample the counter distribution. Take a list of keys you know are hot and a list you know are cold and run `OBJECT FREQ` on each. You want two visibly separated bands. 3. Diagnose from the distribution. Everything at or near 255 means the scale is too compressed - raise `lfu-log-factor`. Everything clustered near 5 means counters are not climbing for this traffic rate or are being decayed away - lower the factor or lengthen `lfu-decay-time`. Genuinely hot keys drifting down between bursts means the horizon is too short. 4. Only then touch `maxmemory-samples`, and only if eviction decisions still look poor with a healthy counter distribution. Raise it to 10 and watch latency percentiles; if eviction rate is high, the CPU is not free. 5. Change one knob at a time and give the cache long enough to reach a new steady state - counters take real traffic to redistribute. ## Judgment an interviewer is listening for Say out loud that the defaults are good for most caches and that tuning is warranted only when you can show a problem: a hit ratio below what the working set justifies, or a counter distribution with no separation. Say that the workload shape decides the horizon - a session or feed cache with fast-moving popularity wants a short decay, a catalogue cache with stable long-tail popularity wants a long one. Mention the failure mode of a permanently full instance: eviction runs inline on ordinary commands, so eviction pressure is a latency problem before it is a hit-ratio problem, and the real fix may be more memory or a smaller working set rather than any of these three values.

  • Is the LFU counter decay performed by a background task?
    No. It is lazy: the elapsed time since the stored 16-bit decay timestamp is turned into a number of points subtracted whenever the key is looked up or scored during eviction sampling. That keeps the cost proportional to traffic instead of to the number of keys, and it means eviction always sees a freshly decayed value for the candidates it inspects.
  • What symptom tells you lfu-log-factor is set too low?
    Sampled OBJECT FREQ values pile up at or near 255 across keys with very different real traffic. Once counters saturate they tie, and eviction can no longer distinguish a merely busy key from an extremely hot one, so victim choice degrades toward arbitrary among the saturated set. Raising the factor stretches the same 0-255 range over more accesses and restores separation.
  • When is the right answer to none of these knobs?
    When the instance sits permanently at maxmemory with a high eviction rate. Then eviction CPU is being paid inline on ordinary commands and the working set simply does not fit, so tuning changes which keys are lost rather than whether keys are lost. The real fixes are more memory, sharding the keyspace, or shrinking values and TTLs.

Log factor is the scale on the chart, decay time is how long the chart's memory runs, and sample size is how many data points you glance at before deciding what to cut.

saying these in an interview costs you the question

  • Believing a background thread periodically decays every key's counter
  • Treating lfu-log-factor and lfu-decay-time as interchangeable ways to make eviction 'more aggressive'
  • Tuning several settings at once and attributing any improvement to the last change
  • Assuming higher maxmemory-samples is free because eviction feels rare
  • Ignoring that saturation at 255 makes counters tie and eviction effectively arbitrary

context