skip to content

For a send-later service whose delays range from seconds to a year, how do you choose between broker-native delayed delivery and a durable timer store?

level: principalimportance: should knowfreq 38%

answer

  1. horizon, precision, mutability, visibility
  2. delay caps differ by broker
  3. pending messages are opaque
  4. store as source of truth
  5. version check drops stale fires

basics

~20 s

Weigh maximum delay, precision, cancellation, listing needs and the cost of holding pending work. Broker delays suit short, fire-and-forget delays within the broker's limits; a durable store suits long horizons and edits. Many systems combine them in tiers.

solid answer

~50 s

Start from the requirements. Broker-native delayed delivery is simple — publish with a delay and the message appears later — but brokers differ: some cap the maximum delay, some offer only fixed delay levels, some have no per-message delay at all, and a pending message usually cannot be listed, edited or deleted on its own. A durable timer store (a due-time index or bucketed table) handles any horizon, supports cancel, reschedule and 'show my scheduled messages', and is the audit record, at the cost of running pollers. A common hybrid keeps the store as source of truth and has a mover hand jobs due within a short window to a broker delay or an in-memory timing wheel for precise firing. Cancellation then works by version check: the consumer compares the message's version with the store's current record and drops stale ones.

code

pseudocode · 8 lines
pseudocode
on_delayed_message(msg):
  job = store.get(msg.job_id)
  if job is null or job.status != 'handed_off':
    return                     # cancelled or already done
  if job.version != msg.version:
    return                     # stale: job was rescheduled
  if store.compare_and_set(job.job_id, job.version, status = 'running'):
    execute(job)

go deeper

for a junior

Remember the two options: ask a broker to deliver a message later, or store the job in a table and poll for due rows.

for a middle

Explain the typical broker limits — delay caps, fixed levels, no listing or deletion — and what a polled store gives in return.

for a senior

Describe the tiered hybrid: store as source of truth, a mover with a hand-off window, and version checks for cancel and reschedule.

for a principal

Decide from requirements and costs: horizon, precision, mutability, pending volume and team capacity, and say when broker-only is enough.

## The requirements that decide it A **send-later service** accepts work to run in the future — a reminder next week, a follow-up in 30 seconds, a renewal notice in a year. There are two broad ways to hold that work until it is due, and the choice follows from a handful of requirements: - **Horizon:** the longest delay you must support. - **Precision:** how close to the due time a job must fire. - **Mutability:** whether users cancel, edit or reschedule pending jobs. - **Visibility:** whether anyone needs to list, count or audit pending jobs. - **Volume:** how many jobs are pending at once, and what holding each one costs. - **Operational budget:** how much custom machinery the team can run. ## Broker-native delayed delivery Some **message brokers** let a producer attach a delay to a message; the broker holds it and makes it visible to consumers only when the delay expires. - **Strengths:** almost no custom code, delivery rides the broker's existing durability and consumer scaling, and precision is usually good for short delays. - **Limits that vary by broker:** some cap the maximum delay, sometimes at minutes or days; some support only a fixed set of delay levels; some have no per-message delay, so teams build it from per-delay queues or message expiry; how many pending delayed messages a broker handles well also varies. - **Opacity:** a delayed message is usually not queryable. You typically cannot list a user's scheduled messages, change a payload, or delete one message by ID before it fires. ## Durable timer store A **timer store** is a database table (or bucketed partitions) of jobs with a `due_at`, read by **pollers** that claim and dispatch due jobs. - **Any horizon:** a job due in a year is just a row. - **Mutable and queryable:** cancel is a status change, reschedule is an update, 'show my scheduled messages' is a query. - **Auditable:** the row records who scheduled what, when it fired and whether it was cancelled. - **Costs:** you run pollers, claim logic and partitioning, and precision is bounded by the poll interval. | Requirement | Broker-native delay | Durable timer store | |---|---|---| | Year-long horizon | often capped | yes | | Sub-second precision | often good | limited by poll interval | | Cancel or reschedule one job | usually not directly | yes | | List pending jobs | usually not | yes | | Custom code to run | little | pollers, claims, partitions | ## The tiered hybrid Many large systems use both, each for what it does best: 1. The **durable store** is the source of truth for every scheduled job, whatever its horizon. 2. A **mover** periodically selects jobs due within a **hand-off window** — say the next 15 minutes — and marks them `handed_off`. 3. It publishes each as a broker delayed message (or loads it into an in-memory **timing wheel**) with the remaining delay. 4. The broker or wheel fires it precisely; a consumer checks the store and executes it. 5. Jobs scheduled inside the window at creation time skip straight to step 3, but are still written to the store first. The window must be shorter than any broker delay cap and longer than the mover's worst-case lag, or jobs fire late. ## Cancelling and rescheduling Once a message sits in a broker, you may not be able to remove it. The usual answer is to make the store authoritative and the message a **hint**: ```json { "job_id": "j-81f2", "due_at": "2027-03-01T09:00:00Z", "status": "pending", "version": 3 } ``` - The published message carries `job_id` and `version`. - **Cancel:** set `status = 'cancelled'` in the store. When the message fires, the consumer reads the row, sees the status, and drops it. - **Reschedule:** update `due_at` and increment `version`, then publish a new message if the new time is inside the window. The old message still fires, carries version 2, sees version 3 in the store, and is dropped as **stale**. This tolerates duplicate publications from a mover that crashed mid-hand-off, too, as long as the consumer's check-and-execute step is itself safe to repeat. ## Cost and failure trade-offs - **Pending volume:** holding billions of long-horizon messages inside a broker may be expensive or unsupported; rows in a store are cheap. - **Blast radius:** a broker outage in the hybrid delays only the near-term window; the store still holds everything else. - **Stale-message waste:** heavy cancellation makes the consumer discard many messages; a narrower window reduces that. - **Complexity:** broker-only is simplest when every requirement fits inside its limits — choosing it is a legitimate answer for short, immutable delays.

  • How is a job cancelled after its message is already sitting in a broker delay?
    Treat the store as authoritative. Cancelling sets the row's status to cancelled; when the delayed message eventually fires, the consumer reads the row, sees it is no longer handed off, and discards the message. Nothing needs to be removed from the broker, which is useful because many brokers cannot delete one pending message by ID.
  • How wide should the hand-off window be?
    Wide enough to cover the mover's worst-case lag, so jobs are handed off before they are due, and narrow enough to stay under any broker delay cap. A wider window also means more messages waiting in the broker and more stale fires when users cancel or reschedule. Minutes, not hours, is a common choice.
  • When is broker-native delay alone the right answer?
    When every delay fits within the broker's limits, jobs are never edited or cancelled, nobody needs to list pending work, and the number of pending messages is modest. Short retry-style delays and fixed follow-ups often fit. Choosing the simpler design there is good judgment, not a shortcut.

saying these in an interview costs you the question

  • Broker delay handles any horizon; a year-long delay is just a parameter.
  • A delayed message in a broker can always be deleted by its ID.
  • Cancelling requires removing the message from the broker before it fires.
  • In the hybrid, consumers need no version check since the store is authoritative.
  • A durable timer store always fires more precisely than a broker delay.