In a Redis primary-replica setup, how does a key whose time-to-live has elapsed actually get removed on the replica, and what surprising behaviour does that cause when you read from or measure that replica?
answer
- Only the primary expires
- Replica waits for the synthesized DEL
- Replica reads mask, never delete
- Replica DBSIZE can exceed primary
- Promotion transfers expiry duty
basics
~20 sReplicas never expire keys on their own. They wait for the primary to send an explicit DEL/UNLINK when it reclaims the key. Meanwhile a replica read masks the key as missing, so its key count and memory can exceed the primary's.
solid answer
~50 sExpiration is decided **only by the primary**. When the primary reclaims a key — lazily on access or via the active cycle — it propagates an explicit DEL (UNLINK with `lazyfree-lazy-expire`) down the replication stream. The replica applies that delete like any other write. A replica therefore holds logically expired keys in its dataset until that command arrives. To keep clients correct, the replica's read path checks the deadline itself and reports such a key as missing, without deleting it. So reads are right, but *counters* are not: DBSIZE, INFO keyspace and used_memory on a replica can legitimately exceed the primary's, especially if the primary is idle and the sampler is slow to get around to those keys. The rationale is determinism: if replicas expired independently, the same read could differ across nodes and datasets would diverge. Writable replicas are the exception — they can expire keys they created themselves, which are wiped on the next full resync. After a failover the promoted replica takes over expiry duty.
go deeper
Recall the rule: replicas do not expire keys themselves; the primary sends a delete, and replica reads still never return expired values.
Explain the synthesized DEL/UNLINK in the replication stream and why replica key counts and memory can exceed the primary's.
Bring the operational angle: don't alert on count divergence, an idle primary delays reclamation everywhere, and expiry tuning is a primary-side change.
Frame it as determinism versus promptness in an asynchronously replicated store, including failover handover of expiry authority and the writable-replica exception.
## Who is allowed to decide a key is gone Redis replication is asynchronous and effect-based: the primary sends the *effects* of commands, not the commands' non-deterministic inputs. Expiry is exactly the kind of thing that must not be decided independently. Two nodes reading their own clocks, sampling their own random keys, and deleting at slightly different moments would drift apart, and a client reading a replica could see a different dataset than one reading the primary. So Redis makes the rule simple: **only the primary expires keys.** When the primary reclaims a key — through the lazy check on access or through the active sampling cycle — it synthesizes an explicit `DEL` (or `UNLINK` if `lazyfree-lazy-expire` is on) and writes it to the replication stream and the AOF. Replicas apply it as an ordinary write. The same synthesized delete is what makes AOF replay deterministic. ## Why replica reads are still correct Between the deadline and the arrival of that DEL, the replica physically holds a key that is logically dead. If it simply served it, replica reads would return values the primary considers gone. So the replica's lookup path performs the same deadline check and reports the key as **missing** — but does *not* delete it, because deleting would be an independent decision. The key stays in the dataset, invisible to readers, until the primary's DEL lands. ## The surprising consequences - **Key counts diverge.** `DBSIZE` and `INFO keyspace` on a replica can be higher than on the primary. Monitoring that alerts on "primary and replica key counts differ" will fire spuriously. - **Memory diverges.** `used_memory` on a replica can be higher for the same reason. If the primary is mostly idle, nothing triggers lazy expiry there, so reclamation happens only at the sampler's pace and the replica trails it. - **Scans can surface ghosts.** A keyspace scan on a replica may return names whose subsequent read returns nil. - **Expired events lag on replicas.** The replica's `expired` notification is tied to the arrival of the primary's delete, not to the deadline. - **A read-only replica cannot fix itself.** No amount of reading a replica reclaims memory there; the only thing that does is the primary reclaiming and propagating. ## Writable replicas and failover If `replica-read-only` is disabled, a replica can accept its own writes. Keys it created locally are tracked separately and *can* be expired by that replica, since no primary owns them. Those keys are not replicated anywhere and vanish on the next full resynchronization, which is exactly why writable replicas are a niche tool rather than a general pattern. After a failover, the promoted node becomes the authority: it now runs both lazy and active expiration for the whole dataset and propagates deletes to the remaining replicas. A backlog of logically expired keys inherited at promotion time will be cleaned up by its sampler, which can look like a burst of DEL traffic and a memory drop shortly after promotion. ## What to do with this operationally Alert on *primary* memory and `expired_keys` rate, not on replica-versus-primary key-count equality. If a replica's memory looks stubbornly high while the primary's is fine, suspect an idle primary plus a large volatile set rather than a replication bug — and remember that raising expiration aggressiveness is a primary-side change; replica-side tuning does nothing for it.
- Why not let replicas delete expired keys themselves to save memory?Because that decision depends on each node's clock and its own random sampling, so replicas would diverge from the primary and from each other, and a read served by different nodes could differ. Redis prefers a deterministic dataset driven by one authority, paying for it with some lagging memory on replicas.
- A monitoring check alerts because a replica reports more keys than its primary. Is this a bug?Usually not. The replica still holds keys whose deadline passed but for which the primary has not yet sent a DEL, typically because nothing accessed them and the active cycle has not sampled them yet. Compare expired_keys rate and primary memory instead, and only investigate if the gap keeps growing without bound.
The primary is the only person allowed to cross names off the guest list; replicas keep their copy intact and simply refuse anyone whose slot has passed until the official strike-through arrives.
saying these in an interview costs you the question
- Claiming each replica runs its own expiration cycle
- Claiming a replica can serve a value whose TTL elapsed
- Treating a replica/primary key-count mismatch as data loss or corruption
- Expecting reads against a replica to reclaim its memory
- Assuming a writable replica's locally expired keys survive a resync