skip to content

A teammate reports that the Redis 'expired' notification for a key with a 10-second TTL sometimes arrives minutes late. Why does that happen, and what determines the moment the event is actually published?

level: seniorimportance: should knowfreq 34%

answer

  1. event fires at deletion, not at TTL elapse
  2. touch-driven or sampled by the periodic cycle
  3. logically gone at TTL, physically later
  4. huge expires set + untouched key = long tail
  5. value already gone; no ordering guarantee

basics

~20 s

The event is published when the key is actually deleted, not when its TTL logically elapses. Deletion happens either when something touches the key or when the background expiry cycle happens to sample it, so an untouched key in a large keyspace can wait far beyond its TTL.

solid answer

~50 s

`expired` is a deletion event, not a timer. Redis does not schedule a callback per TTL; a key whose TTL has passed is logically gone (reads return nothing) but physically still present until it is removed, and the notification is published at removal. Removal happens on two paths: when a command touches the key, which deletes it immediately, or when the periodic expiry cycle samples it. That cycle works on random samples of keys carrying TTLs and is bounded so it never monopolises the server, so with a huge expires set and low traffic on the specific key, the gap between TTL elapsing and deletion can be seconds or minutes. Operational consequences: never treat `expired` as a scheduler tick; never assume ordering between two keys' expiry events; and remember the value is already gone when the event fires, so a subscriber cannot read what was lost. In a replicated setup the primary drives the deletion and propagates it, so timing follows the primary, not the replica.

go deeper

for a junior

Know that the event fires when the key is actually deleted, which can be later than the TTL, and that reads already see the key as gone.

for a middle

Explain both deletion paths - touch-driven and bounded periodic sampling - and why the lag grows with the number of TTL-bearing keys.

for a senior

Diagnose it: correlate lag with expiry pressure and key traffic, state that ordering and punctuality are not guaranteed, and redirect the requirement to a due-time queue if timeliness matters.

for a principal

Rule on whether the architecture should depend on expiry events at all, and set the boundary between 'data must be unavailable at T' (TTL is right) and 'work must run at T' (a scheduler is right).

## The event is a side effect of deletion The crucial mental model is that `expired` is emitted by the code path that removes the key, not by a timer attached to the TTL. Redis has no per-key timer wheel. Once a key's TTL has passed, the key is treated as absent by every command - a `GET` returns nil, an `EXISTS` returns 0 - but the entry still occupies memory until something deletes it, and only that deletion publishes the notification. Two things trigger the deletion: a client command that reaches the key, which removes it on the spot, or the server's periodic expiry work, which examines samples of the keys that carry TTLs. That periodic work is deliberately bounded in effort per cycle so that expiry never turns into a stall, which means it does not guarantee that any particular expired key is found within any particular window. A key nobody touches, in a database with a very large number of TTL-bearing keys, may sit past its deadline for a long time before it is sampled - and its notification is late by exactly that amount. ## Why the lateness is variable The delay is a function of how much of the keyspace carries TTLs, how many of those are already past their deadline, and whether anything reads the key. Under light expiry pressure the delay is typically small; when a huge batch of keys expires at the same instant, they drain over time rather than all at once, and the notifications trickle in accordingly. The upshot is that the delay distribution has a long tail, and any design that treats the notification as punctual will misbehave under exactly the conditions that matter - load and large keyspaces. ## What the event does and does not tell you - It tells you the key is gone *now*. It does not tell you when the TTL elapsed; there is no timestamp in the message. - It carries no value. The data was removed before the message was published, so a subscriber that reacts by reading the key finds nothing. Designs needing the old payload must keep a companion record, for example a mirror key with a longer TTL or a row in a durable store keyed by the same identifier. - Ordering between different keys is not meaningful. Two keys whose TTLs elapse a second apart can produce events in the opposite order depending on which is touched or sampled first. ## Replication and cluster placement Expiry is driven by the primary: replicas do not independently delete data that has expired, they wait for the primary to propagate the deletion, so that primary and replica agree on content. The practical consequence for a listener is that timing is the primary's timing, and the safest place to subscribe is the node where the deletion decision is made. In Redis Cluster the key lives on the shard that owns its slot, so the deletion happens there; keyspace notifications use ordinary non-sharded Pub/Sub, which propagates across the cluster bus, so a subscriber does not have to be attached to that specific shard - but it pays the corresponding bus traffic. ## How to answer the teammate Start by separating two questions they have merged: "is the key still readable?" (no, it is logically gone at the TTL) and "when do I get told?" (when it is physically deleted). Then decide whether the application actually needs punctuality. If it only needs the data to be unavailable at the deadline, nothing is broken and the notification lateness is irrelevant. If it needs an action to run at the deadline, keyspace notifications are the wrong mechanism entirely and the work belongs in a due-time queue the application polls or in a stream the workers consume. A useful diagnostic while investigating: compare the number of keys carrying TTLs against the observed lag, and check whether the affected keys are ever read - a key that is read at the deadline is deleted immediately and never shows the delay, which is why the problem so often looks intermittent.

  • Is a key readable between its TTL elapsing and its physical deletion?
    No. Every command that reaches the key checks its deadline first and treats it as missing, so reads return nil and EXISTS returns 0. The delay affects only when memory is reclaimed and when the notification is published, never the visible contents of the database.
  • Why is the delay usually invisible for hot keys and obvious for cold ones?
    A key that a client touches at or after its deadline is deleted right then, publishing the event immediately. A key nobody touches waits for the bounded periodic sampling to find it, and the wait grows with the number of TTL-bearing keys. That is why the same TTL looks punctual on busy keys and late on idle ones.
  • Where should a listener attach in a primary-replica setup?
    Where the deletion decision is made, which is the primary: replicas do not expire data on their own but apply the deletion the primary propagates, so the timing you observe is the primary's. Attaching a listener to a replica also means its view depends on replication lag on top of the sampling delay.

The library marks a book overdue the instant the date passes, but the notice only goes out when a clerk happens to pull that shelf or a reader asks for the book.

saying these in an interview costs you the question

  • Treating the expired event as a punctual timer callback for scheduling work
  • Believing an expired-but-not-yet-deleted key can still be read
  • Expecting the notification payload to include the value that was lost
  • Assuming expiry events arrive in TTL order across keys
  • Assuming a replica independently expires keys and fires its own timely events

context