A scheduled (cron-style) trigger fires a serverless function every 5 minutes to reconcile a data feed, and each run typically takes about 4 minutes. What can go wrong if the schedule doesn't account for overlapping executions or missed invocations, and how would you design around it?
answer
- scheduler doesn't check prior run state
- runtime close to interval -> overlap risk
- missed firing = gone, no persistence to catch up on
- lock/lease at run start prevents overlap
- watchdog alert on missing completion, not on error
basics
~20 sIf a run takes almost as long as the gap between runs, sometimes the next run starts before the last one finishes, and now two copies are working on the same data at once, which can cause conflicts or double work. You need a way to stop overlapping runs and a plan for what happens if a run is skipped entirely.
solid answer
~60 sScheduled triggers are push-based: the scheduler fires at each configured interval regardless of whether the previous invocation has finished, so if a 4-minute job runs on a 5-minute schedule, any slowdown (an unusually large data feed, a slow dependency) easily produces overlapping executions, where two invocations touch the same underlying dataset concurrently — leading to duplicate writes, race conditions, or corrupted reconciliation state. Separately, the scheduler and the platform aren't guaranteed to never skip a firing (a platform outage, a deployment window, or exceeded concurrency can cause a scheduled invocation to simply not happen), so the job needs to tolerate gaps, not just overlaps. The standard design is to make invocations idempotent and non-overlapping by construction: acquire a lock or lease (e.g., a conditional write to a database row) at the start of the run and have any invocation that can't acquire it exit immediately rather than proceed, combined with monitoring that alerts if the job hasn't successfully completed within some multiple of its expected interval so a missed run doesn't go unnoticed.
go deeper
Should recognize that if a job sometimes runs longer than the gap between scheduled runs, two copies could end up running at the same time, and that this could cause problems.
Should propose a basic mitigation like a lock/flag to prevent overlapping runs, and understand that a missed run doesn't automatically get retried the way a failed queue message would.
Should design a concrete locking/leasing mechanism with a clear no-op-on-contention behavior, and separate the overlap problem from the missed-invocation problem, proposing checkpoint-based catch-up and external monitoring for the latter.
Should reason about this as a broader distributed-systems coordination problem — evaluating trade-offs between locking approaches, designing idempotent/checkpoint-based jobs that tolerate both overlap-prevention no-ops and missed firings without data loss, and building observability that treats 'silence' as a first-class failure signal distinct from explicit errors.
## How a scheduled trigger fires A scheduled trigger, such as a cron() or rate()-style rule on AWS EventBridge Scheduler or an equivalent on other platforms, works by having the platform's own scheduling service push an invocation to the function at each computed point in time, entirely independent of whether any previous invocation is still running, has finished, or has failed. Mechanically, the scheduler evaluates the cron/rate expression, computes the next fire time, and at that instant issues an invoke call to the function — it doesn't check the function's current execution state first, because doing so isn't part of the scheduler's contract; it's a **pure time-based push**, structurally similar to an HTTP trigger in that it's push-based and typically invoked asynchronously, but driven by a clock instead of an external caller. ## Why the scheduler stays this simple This design exists because a scheduler's job is narrowly defined as 'reliably fire at these times,' which is a simpler and more composable contract than also trying to track and serialize against the state of whatever it's invoking — that concern is left to the function or an external coordination mechanism, by design, so the scheduler itself stays simple and horizontally scalable regardless of how long or short downstream invocations happen to run. ## Consequence one: overlapping executions The most direct consequence is the **overlapping-execution problem**: if a job's typical or worst-case runtime is close to or exceeds its scheduled interval, any slowdown — larger-than-usual input volume, a slow downstream dependency, a transient retry — causes the next scheduled firing to launch a second invocation while the first is still active. If both invocations touch the same mutable state (the same database rows, the same output file, the same external system), you get classic concurrency bugs: - **lost updates**, where one invocation's write clobbers the other's, - **duplicate side effects**, like double-sending a report or double-charging a reconciliation adjustment, - **corrupted aggregate state**, if both invocations read-modify-write the same counter without proper locking. This is functionally the same class of problem as any race condition between concurrent processes, except it's easy to overlook here because a schedule 'every 5 minutes' feels like it implies sequential, non-overlapping execution, when the platform makes no such guarantee whatsoever. ## Consequence two: missed invocations The less obvious but equally important consequence is **missed invocations**. Scheduled triggers are generally reliable but not infallible — these can all cause a scheduled firing to simply not happen: - a platform-side incident, - a deployment that temporarily removes or misconfigures the schedule, - an account-level concurrency limit being exhausted by unrelated functions, - (for some platforms) certain schedule-expression edge cases around daylight-saving transitions. Unlike a queue, where an unprocessed message just sits there waiting to be picked up later, a scheduled trigger firing has no persistence if it's missed — there is no message to catch up on, because the 'event' is nothing more than the clock reaching a certain time, and once that moment passes, it's gone. A reconciliation job that silently misses several consecutive firings can leave a data feed meaningfully out of sync for an extended period with no error, no alert, and no automatic backfill — the failure is invisible unless something is specifically watching for it. ## Guarding against overlap The standard architectural response to overlap is to make the job **self-serializing** rather than relying on the schedule's spacing to guarantee non-overlap. A common pattern is a lightweight distributed lock or lease acquired at the very start of the run — a conditional write to a database row with a TTL, an atomic put on a coordination service, or a platform-specific concurrency-limit-of-one on the function itself — where any invocation that fails to acquire the lock (because a prior run is still holding it) exits immediately as a safe no-op rather than proceeding to touch shared state. This converts 'two invocations racing on the same data' into 'one invocation runs, one invocation politely skips,' at the cost of occasionally skipping a scheduled run entirely when the previous one runs long, which is usually the correct trade-off for a reconciliation-style job where doing the same work twice concurrently is worse than doing it once, slightly late. ## Guarding against a missed run The standard response to missed invocations is twofold: 1. **Design the job to be self-catching-up where possible** — a reconciliation job that processes 'everything since the last successful high-water-mark checkpoint' rather than 'only what changed in exactly the last 5 minutes' naturally absorbs a missed firing on its next successful run, since it'll simply cover a larger window. 2. **Add independent monitoring** that alerts based on the absence of a successful completion signal within some multiple of the expected interval (for example, alerting if no successful run has been recorded in 20 minutes for a job expected every 5), rather than relying on the function's own error handling, since a missed invocation produces no error to catch — there's simply nothing that ran. ## A concrete scenario A concrete scenario: a billing-reconciliation function scheduled every 5 minutes reads pending transactions and updates a ledger; during a traffic spike, a run takes 7 minutes instead of the usual 4. Without a lock, the next scheduled firing starts 5 minutes in, sees an overlapping set of 'pending' transactions the first run hasn't finished marking as processed, and both invocations attempt to apply the same adjustments to the ledger — over-crediting affected accounts. With a lease-based lock, the second invocation instead detects the first is still holding the lock, exits as a no-op, and the ledger stays consistent, catching up fully on the next successful run two minutes later.
- Why is a function-level reserved-concurrency-of-one setting alone not always sufficient to prevent overlapping scheduled runs?Reserved concurrency of one will queue or throttle a second concurrent invocation attempt on the same function, which can prevent true simultaneous execution, but depending on the platform's exact throttling behavior it may cause the second invocation to be silently dropped/errored rather than gracefully deferred, and it doesn't help if the job is decomposed across multiple functions or steps where the concurrency limit doesn't span the whole workflow. An explicit lock plus a graceful no-op exit is more portable and predictable across those cases.
- How would you design a reconciliation job to be naturally resilient to occasionally missing a scheduled firing, without adding an explicit backfill mechanism?Instead of processing only 'records changed since the last fixed 5-minute boundary,' have the job track its own high-water-mark checkpoint of the last successfully processed point and always process everything from that checkpoint to now. A missed firing simply means the next successful run covers a wider time window than usual, catching up automatically without any special-cased backfill logic.
- What's the risk of using the function's own error-handling or a dead-letter queue to detect a missed scheduled invocation?A missed invocation never runs at all, so there's no error to throw, no failed execution to retry, and nothing for a dead-letter queue to catch — DLQs and error handlers only fire when an invocation happens and then fails, not when the scheduler fails to invoke in the first place. Detecting a miss requires independent, external monitoring based on the absence of an expected success signal, not on any mechanism inside the function's own execution path.
It's like a train that departs on a fixed timetable no matter whether the previous train has actually cleared the platform — if trains run late often enough, you eventually get two trains trying to use the same platform at once. And if a departure is simply cancelled with no announcement, nobody automatically runs a makeup train unless someone is watching the schedule board.
saying these in an interview costs you the question
- Assumes the platform automatically serializes scheduled invocations so overlap is impossible
- Believes a missed scheduled firing will automatically be made up later
- Thinks the function's own try/catch or a DLQ can detect a firing that never happened
- Proposes relying purely on choosing a longer interval as the complete fix for overlap risk
- Doesn't distinguish between preventing overlap and detecting missed runs as two separate problems