skip to content

Every order must trigger an action exactly 30 minutes after it is placed, for millions of orders a month. How would you weigh EventBridge Scheduler one-time schedules against the other AWS ways of building that timer?

level: principalimportance: should knowfreq 34%

answer

  1. precision, cancellation, scale, cost
  2. live timers, not monthly volume
  3. delay queues stop well short of 30 minutes
  4. a wait state should carry real orchestration
  5. TTL is cleanup, never a deadline

basics

~20 s

Weigh precision, cost per timer, cancellability and scale ceiling. EventBridge Scheduler one-time schedules give exact firing and individual cancellation at the price of an API call and a quota slot per order; the alternatives each break on one of those axes.

solid answer

~60 s

I frame it on four axes. **Precision**: does the action need to land on the minute, or is "roughly half an hour" fine? **Cancellability**: can a timer be cancelled or rescheduled when the order changes, and how cheaply? **Scale**: what is the ceiling on live timers, and who cleans them up? **Cost and operational surface** per timer. Scheduler one-time schedules score well on the first three — an `at()` schedule fires precisely, is deleted by name when the order completes, and `ActionAfterCompletion=DELETE` cleans up the rest — but you pay a create call per order and consume an account quota, so I would check the schedules quota in Service Quotas against peak live orders. An SQS delay queue is cheaper but tops out at 15 minutes of delay and cannot cancel a specific message. A Step Functions workflow per order gives a cancellable wait plus the rest of the business process, at the price of an execution per order. A database column plus a periodic sweeper is the cheapest and the most work to own.

go deeper

for a junior

Know that EventBridge Scheduler can create a one-time schedule per entity and that SQS message delay is capped well below 30 minutes, so it is not a general timer.

for a middle

Compare the options concretely — Scheduler one-time schedules, a Step Functions wait, and a swept timer table — and explain how each handles cancellation and cleanup.

for a senior

Show operational judgment: convert volume into live timers, make schedule creation reliable on the write path, name schedules so they can be deleted, and make the handler idempotent because any timer can fire late or twice.

for a principal

Own the decision and its consequences — the coupling the write path now has to a scheduling service, the quota as a capacity plan, the burst behaviour of a cohort of timers, and when building a sweeper is genuinely cheaper than buying one.

## The shape of the problem "Do something per entity, later" is one of the most common designs interviewers probe, because every option is defensible and the reasoning is what is being graded. State the axes first, then place the options on them. **Precision** — must it fire at 30 minutes, or is a window acceptable? **Cancellation** — orders get paid, cancelled or amended, so most timers must be revocable by identity. **Scale ceiling** — how many timers are live at peak, and what quota or table growth does that imply? **Cleanup** — who removes a timer that has fired or been superseded? **Cost and ops surface** — per-timer charges, and how much code you own. ## Option 1: EventBridge Scheduler, one schedule per order Create an `at()` schedule when the order is placed, named deterministically from the order id, with `ActionAfterCompletion=DELETE` so it removes itself after firing. Cancellation is a `DeleteSchedule` call by that name; rescheduling is an update. Strengths: exact firing time; per-schedule IAM role, retry policy and dead-letter queue so a failed invocation is visible and redrivable; time-zone support if the deadline is expressed in local terms; no infrastructure of yours in the path. This is the option Scheduler was built for — its schedules-per-account quota is in the millions rather than the hundreds that constrained scheduled rules. Costs and cautions: an API call per order on the write path, so the order service now depends on Scheduler's availability — decide whether a failed `CreateSchedule` fails the order or is retried asynchronously. The quota is per account and region, so model *live* timers (30 minutes of order rate, not monthly volume) and check the current figure in Service Quotas rather than trusting a remembered number. And if you ever forget `ActionAfterCompletion`, completed schedules accumulate against that quota. ## Option 2: SQS delayed message Send a message with a delivery delay when the order is placed and let the consumer act when it arrives. Cheap, dead simple, no cleanup. It fails this requirement on two counts. SQS's per-message delay maxes out at 15 minutes, so 30 minutes needs chaining or a re-drive trick — a smell. And you cannot cancel a specific in-flight message; the consumer must re-check the order's current state and no-op if the action is no longer wanted. That state check is mandatory anyway in every option, but here it is the *only* defence, which means a stale timer always does a wasted read. ## Option 3: Step Functions workflow per order Start a Standard workflow at order creation with a wait state, then the action. Standard workflows can wait a very long time, and the workflow can be stopped by execution name when the order changes. This wins when the timer is one step of a larger process — wait, check, retry, escalate, compensate — because you get the whole orchestration, plus per-execution history for free auditing. It is heavier when the timer is *all* you need: one execution per order, priced per state transition, and an execution history to manage. The honest rule is that a wait state should ride along with real orchestration rather than exist for it. ## Option 4: your own timer table Store a due timestamp on the order (or in a timers table) and sweep it on a schedule — a single recurring job querying "due before now". Cheapest per timer by a wide margin, cancellation is a row update, and there is no external quota. You own everything else: the sweep interval sets your precision floor, the query needs an index that supports it, and the sweeper needs idempotency and leader-ish behaviour so two runs do not double-fire. DynamoDB's TTL is sometimes proposed here — do not use it as a timer. TTL deletion is best-effort and can lag well past the expiry time, so it is fine for reclaiming space and wrong for anything with a deadline. ## How I would actually decide If 30 minutes must be 30 minutes and the action is a single step: Scheduler one-time schedules, named by order id, self-deleting, with a dead-letter queue on the schedule. If the 30-minute wait is one beat of a longer saga: Step Functions. If the tolerance is loose, volume is extreme, and cost dominates: the sweeper, accepting that you now own a piece of infrastructure. Two things I would insist on regardless. **The handler must be idempotent and must re-read current state** — every timer mechanism can fire late, twice, or after the reason for it has evaporated, so the timer is a hint to check, never an instruction to act blindly. And if a large cohort of timers can land at the same instant — everything placed during a flash sale — the flat spike is a real risk; Scheduler's flexible time window spreads it when precision permits, and a sweeper naturally batches. ## The trap in the question The phrase "millions a month" invites a quota panic. Convert it: what matters is concurrent live timers, which for a 30-minute delay is thirty minutes of arrival rate, a much smaller number. Making that conversion out loud is most of the answer.

  • How do you cancel the timer when an order is paid early?
    Name the schedule deterministically from the order id and call `DeleteSchedule` on that name when the order changes state. Treat the delete as best-effort: the handler must still re-read the order and no-op if the action no longer applies, because a delete can lose a race with an invocation already in flight.
  • What if creating the schedule fails while placing the order?
    Decide deliberately. Either the order write and the timer creation are made atomic through an outbox — persist the intent, then a reliable consumer creates the schedule with retries — or the order fails loudly. What you must not do is swallow the error, because the silent failure mode is an order that never gets its 30-minute action and no signal that anything went wrong.
  • A flash sale puts 200,000 orders in one minute. What breaks?
    Two things: the burst of CreateSchedule calls on the write path, which you smooth with an outbox and retries, and 200,000 invocations landing 30 minutes later in the same minute. If the deadline tolerates slack, a flexible time window spreads the second spike; if it does not, the downstream must be provisioned or queued for that peak.
  • Why not use DynamoDB TTL as the timer?
    TTL deletion is best-effort background work and can lag well beyond the expiry timestamp, so an action driven by the TTL delete event may run much later than intended. It is the right tool for reclaiming storage on expired data and the wrong tool for anything with a deadline attached.

saying these in an interview costs you the question

  • Quotes monthly order volume as the number of live timers
  • Proposes an SQS delay queue for a 30-minute wait
  • Uses DynamoDB TTL as a precise timer
  • Assumes the timer fires exactly once, so skips idempotency
  • Never asks whether the timer must be cancellable

context