skip to content

In a reminder service where millions of users schedule jobs for exactly 09:00, how do you keep the top-of-hour spike from overwhelming dispatch?

level: seniorimportance: should knowfreq 40%

answer

  1. humans pick round times
  2. tolerance differs per job type
  3. offset from a hash, not random
  4. stage exact-time work early
  5. future load is already stored

basics

~20 s

Spread the load: add deterministic jitter at scheduling time where the product tolerates it, start pulling exact-time jobs slightly early, cap dispatch rate behind a queue, and scale workers ahead of time from the known due-time histogram.

solid answer

~50 s

People pick round times, so the due-time histogram has sharp spikes at `:00` and `:30`, many times the average rate. First decide the precision contract per job type. Where 'around 9' is fine — digests, marketing, background syncs — add jitter at write time, derived from a hash of the job ID so a reschedule or replay does not reshuffle it. For jobs that must fire on time, fetch due work a little early into a staging area and release it at the due second, or accept bounded lateness behind a rate limiter. Because future jobs are stored, the spike is known in advance: scale workers and downstream capacity before it arrives. Illustratively, 3,000,000 jobs at 09:00 with capacity for 20,000 per second take 150 seconds to drain; spread across five minutes they need only 10,000 per second.

code

pseudocode · 7 lines
pseudocode
schedule(job, requested_at, tolerance_seconds):
  if tolerance_seconds == 0:
    due_at = requested_at
  else:
    offset = hash(job.id) mod tolerance_seconds
    due_at = requested_at + offset
  store(job, requested_at, due_at)

go deeper

for a junior

Remember why the spike exists — people choose round times — and that jitter means spreading jobs that can tolerate a small shift.

for a middle

Explain deterministic jitter from a hash, why it is applied at write time, and how a rate-limited queue bounds lateness.

for a senior

Show per-job precision contracts, pre-fetching and pre-computing exact-time work, and scaling ahead from the stored due-time histogram.

for a principal

Frame it as a product and capacity decision: which jobs may move, how much lateness is promised, and what reserved capacity costs.

## Why round times spike In a **reminder** or **send-later service**, users and product defaults choose round times: 09:00, 18:00, midnight, the start of the month. The **due-time histogram** — how many jobs are due in each second — is therefore spiky. A service that averages a few hundred jobs per second may see millions due in the same second at 09:00:00. If every one is dispatched at once, the result is a burst that saturates workers, the database, and downstream providers such as an email or SMS provider, which often apply their own rate limits. ## Decide the precision contract first Not every job needs the same precision. Classify jobs by how much lateness or earliness the product tolerates: | Job type | Tolerance | Technique | |---|---|---| | User-set reminder ('at 9') | seconds, not early | pre-fetch, staged release, prioritise | | Daily digest email | minutes either way | write-time jitter | | Background sync or cleanup | tens of minutes | wide jitter, low priority | | Expiry or deadline | none early, small late | exact due time, reserved capacity | Store the tolerance with the job (for example `tolerance_seconds`) so every later stage can make the right decision. ## Jitter at write time **Jitter** is a deliberate offset added to a due time to spread load. - **Apply it when the job is scheduled**, not at dispatch, so the stored due time already reflects the spread and pollers need no special logic. - **Derive it deterministically**, for example `offset = hash(job_id) mod tolerance_seconds`. The same job gets the same offset every time it is written, rescheduled or replayed, and the hash spreads jobs roughly evenly. - **Respect the direction of the tolerance.** Adding a positive offset never fires a job early; a symmetric window is only for jobs where early is acceptable. - **Keep the user's requested time** in its own field so the UI still shows '09:00'. ## Exact-time jobs Jobs that must fire on the second cannot be moved, so you smooth the *work* around them instead: 1. **Pre-fetch.** Start reading jobs due at 09:00 a little before 09:00, into a staging area or an in-memory timing wheel. 2. **Pre-compute.** Render templates, resolve recipients and check preferences ahead of time, so the work left at 09:00 is just the send. 3. **Release on time.** At the due second, hand staged jobs to workers. 4. **Rate-limit the tail.** If capacity is still exceeded, a queue with a rate limiter (a token bucket) drains the excess with bounded lateness, highest-priority jobs first. ## Scale from the schedule, not the symptom A scheduling system has an advantage most services lack: **it knows its future load**. Querying pending jobs grouped by minute for the next few hours gives a forecast. Use it to add workers, warm connections and raise downstream quotas *before* 09:00. Scaling on queue depth alone reacts only after the spike has started, and new workers may arrive after the backlog is already late. - **Forecast per job type**, so you know how much of the peak is movable (jittered) and how much is fixed. - **Reserve capacity** for exact-time jobs and let tolerant jobs use whatever is left. - **Scale down afterwards** on the same forecast, since the spike ends as predictably as it starts. ## The arithmetic Assume, illustratively, 3,000,000 jobs due at 09:00:00 and dispatch capacity of 20,000 jobs per second: - Draining takes 3,000,000 / 20,000 = **150 seconds**, so the last user is about two and a half minutes late. - If those jobs tolerate a five-minute window and are jittered evenly across it, the required rate is 3,000,000 / 300 = **10,000 per second**, half the capacity, and nobody is late relative to their jittered time. - Holding exact-time jobs to 150 seconds of lateness may be acceptable for reminders and unacceptable for deadlines; that is why the contract comes first. ## Not the same as retry jitter Jitter also appears in **retry backoff**, where random offsets stop failed clients from retrying in lockstep. That is a separate concern. Here the jitter spreads *first* executions that users scheduled for the same moment, and it has to respect a per-job tolerance.

  • Why derive the jitter from the job ID rather than a fresh random number?
    A hash of the job ID gives the same offset every time that job is written, so rescheduling, replaying or re-importing it does not move its fire time around. It still spreads jobs roughly evenly across the window, and it makes the fire time reproducible when you debug a complaint.
  • How do you know how big the 09:00 spike will be?
    Ask the store: count pending jobs grouped by due minute for the next few hours. Because jobs are scheduled in advance, that histogram is a forecast, and it can drive pre-scaling of workers, connection pools and downstream quotas before the spike begins.
  • What if the email or SMS provider limits how fast you can send?
    Then that limit is the real capacity, whatever your worker count. Put a rate limiter sized to the provider's quota in front of the send step, order the backlog by priority and tolerance, and publish a lateness budget for exact-time jobs. Jitter for tolerant jobs keeps them out of the peak.

saying these in an interview costs you the question

  • Add random jitter to every job, including ones users expect on the minute.
  • Autoscaling on queue depth will absorb a top-of-hour spike in time.
  • Re-randomising a job's jitter on every reschedule is harmless.
  • Scheduling spikes are unpredictable, so only reactive scaling helps.
  • Dispatching everything at once is fine because a queue will buffer it.