An engineer proposes building a delayed-job scheduler by writing a Redis key with a TTL and running the job when its 'expired' keyspace notification arrives. What are the failure modes of that design, and what would you build instead?
answer
- Pub/Sub = fire-and-forget, at-most-once
- restart or deploy = silent permanent loss
- no ack, no redelivery, no backlog
- broadcast means every replica runs the job
- use a sorted set scored by due timestamp
basics
~20 sKeyspace notifications are fire-and-forget Pub/Sub: no persistence, no acknowledgement, no redelivery. A disconnected, slow, restarting or newly deployed consumer loses events permanently, and every subscriber gets a copy, so jobs are silently dropped or run twice. Use a sorted set of due timestamps polled atomically instead.
solid answer
~60 sThe design puts an at-most-once transport under an at-least-once requirement. Failure modes: - **No persistence or backlog.** Events are published to Pub/Sub and immediately forgotten. A consumer that is restarting, deploying or briefly disconnected misses every event in that window with no way to recover them. - **No acknowledgement or retry.** A consumer that receives an event and then crashes before finishing the job leaves no trace; Redis has no record that the job existed. - **Slow-consumer eviction.** A subscriber that cannot keep up hits the Pub/Sub client output buffer limit and is disconnected, losing everything silently. - **Fan-out duplication.** Every subscriber receives the message, so N replicas of the worker run the job N times unless you add your own locking. - **No payload.** The event carries only the key name; the job's data must live elsewhere. - **Timing is not punctual** - the event fires at deletion, which can lag the TTL. Build instead: a sorted set scored by due timestamp, polled with a ranged read plus atomic claim, or a durable stream with per-consumer acknowledgement.
code
text · 6 lines# schedule
HSET job:9 type email to [email protected]
ZADD jobs:due 1723640000 9
# worker loop: claim due jobs atomically
EVAL "local due = redis.call('ZRANGEBYSCORE', KEYS[1], '-inf', ARGV[1], 'LIMIT', 0, 10) for _, id in ipairs(due) do redis.call('ZREM', KEYS[1], id) end return due" 1 jobs:due 1723640000go deeper
Say clearly that keyspace notifications are fire-and-forget with no retry, so a restarted or disconnected worker loses jobs, and that job state should live in a data structure instead.
Enumerate the concrete failure modes - downtime, crash after receipt, slow-consumer disconnect, broadcast duplication, missing payload - and sketch the sorted-set alternative.
Design the replacement end to end: due-time index, atomic claim, in-flight reclaim after a timeout, plus operational visibility into overdue work; and name where notifications remain appropriate.
Frame it as matching delivery semantics to the requirement: at-most-once hints versus at-least-once work, the cost of building queue semantics in-house, and when to adopt a purpose-built job system instead.
## Why this design is so tempting It looks elegant: `SET job:9 payload PX 60000`, subscribe to `__keyevent@0__:expired`, and the notification is your timer. There is no polling loop and no extra data structure. The design is also easy to demo, because with one worker on a quiet instance it works perfectly - which is exactly why it reaches production before its flaws surface. ## The delivery model is the whole problem Keyspace notifications are delivered over Redis Pub/Sub, which is fire-and-forget with at-most-once semantics. The message is written to whichever subscribers are connected at that instant and then it is gone: there is no log, no offset, no acknowledgement, no redelivery, and no way to ask for what you missed. Every property a scheduler needs is absent. Concretely: **Consumer downtime loses jobs.** Any deploy, restart, crash, or network blip is a window in which every event published is lost permanently. Nothing in Redis knows a job was due, because from Redis's perspective the key simply vanished on schedule. **Crash after receipt loses the job.** The consumer got the message, started work, and died. There is no unacknowledged-message list for Pub/Sub, so nothing will hand the work to anyone else. **Slow consumers are disconnected.** Pub/Sub output for a subscriber is buffered, and a subscriber that does not drain fast enough hits the configured client output buffer limit for the pubsub class and is dropped by the server. The failure is silent from the publisher's side, and looks to the consumer like an ordinary disconnect - jobs disappear precisely when the system is busiest. **Every subscriber gets a copy.** Pub/Sub is broadcast, not a work queue. Run three worker replicas for availability and every job runs three times. Bolting a lock on top means each worker now needs an atomic claim, at which point you have built the harder half of a queue anyway and still have no redelivery. **No payload and no value.** The message carries the key name only, and the key's value has already been deleted, so the job's arguments must be stored under a second key or in another system - adding a second failure mode where the payload survives but the trigger is lost, or the reverse. **Timing is not exact.** The event is published when the key is physically removed, which may lag the TTL when the key is untouched and many keys carry TTLs. A scheduler built on it is not even punctual on the happy path. **Placement matters.** Expiry is driven by the primary in a replicated setup, so a listener attached elsewhere depends on propagation; and because notifications use ordinary non-sharded Pub/Sub, in Cluster they cross the cluster bus, which is bandwidth you are spending on a mechanism that still cannot guarantee delivery. ## What to build instead **Sorted set of due times.** Store each job as a member of a sorted set whose score is the Unix timestamp when it should run: `ZADD jobs:due <due-ts> <job-id>`, with the payload in a hash. Workers poll on a short interval, read the members whose score is at or below now, and *atomically claim* them so exactly one worker gets each - a small Lua script that reads and removes in one step, or a move into an in-flight set with a claim timestamp so abandoned jobs can be reclaimed after a timeout. The state is in the dataset, so it survives restarts, is covered by persistence and replication, and is inspectable: you can count overdue jobs, requeue, and audit. **A durable log with acknowledgements.** If jobs are due immediately rather than at a future time, an append-only stream with consumer groups gives per-consumer delivery, retained history, and explicit acknowledgement of each entry, so an unacknowledged job can be claimed by another worker. It solves the loss and duplication problems that Pub/Sub cannot; scheduling into the future still needs the due-time index above. **Or a real job system.** If retries, backoff, visibility timeouts, and dead-letter handling are requirements, that is a queue product's job, and rebuilding it on top of expiry events is a poor trade. ## The one-line rule Keyspace notifications are a *hint* that something happened, useful for cache invalidation and monitoring where a lost message costs a stale entry or a missing data point. They are not a delivery mechanism for work that must happen.
- What exactly happens to a notification if no client is subscribed at that instant?It is discarded. Redis publishes to the set of subscribers connected at that moment and keeps no history, offset, or backlog, so there is nothing to replay. From the server's point of view the delivery succeeded with zero recipients, and no error or metric marks the loss.
- Why does running several worker replicas make the design worse rather than more reliable?Pub/Sub is broadcast, so every connected subscriber receives every message and each replica executes the same job. Adding replicas for availability therefore multiplies duplicate executions instead of sharing the load. Fixing it requires an external atomic claim per job, which is most of the work of building a proper queue while still leaving the loss problem unsolved.
- If a sorted set of due timestamps requires polling, is that not worse than an event?Polling a sorted-set range on a short interval is cheap - one ranged read returning only members already due - and it buys durability, redelivery, and inspectability. The state lives in the dataset, so it survives restarts and is replicated and persisted like any other data. Predictable poll cost is a good trade for a scheduler that cannot silently drop work.
- Are keyspace notifications ever the right tool?Yes, where a lost message is tolerable and cheap to recover from: invalidating a local in-process cache entry, refreshing a dashboard, emitting metrics, or triggering a best-effort warm-up. The recovery path in those cases is a subsequent miss or a periodic refresh, so an at-most-once hint is sufficient.
It is a public address announcement, not a ticket queue: anyone in the room hears it, anyone who stepped out never will, and nobody can ask for a repeat.
saying these in an interview costs you the question
- Assuming Pub/Sub retains or replays missed messages
- Thinking a persistent Redis connection makes delivery reliable
- Adding worker replicas for reliability without noticing each one runs every job
- Expecting the expired event to carry the job payload
- Believing that enabling AOF or replication makes keyspace notifications durable