An hourly pipeline regularly takes 90 minutes to finish — what happens to the next scheduled run?
answer
- work is arriving faster than it drains
- either they overlap or they queue
- delay compounds every interval
- every run can succeed while data goes stale
basics
~20 sIt depends on the concurrency policy. Either two runs execute at once and contend for the same source and target, or the pipeline is capped at one active run and the queue backs up, so each run starts later than the last and lag grows without bound.
solid answer
~50 sRuntime exceeding the interval is a structural problem with only two possible behaviours, both bad if unplanned. **Overlap allowed:** the 10:00 run is still going when the 11:00 run starts, so two runs hit the same source at once, double the load, and race on the same target tables — safe only if each run writes a partition scoped to its own interval and never touches another's. **Overlap capped at one:** the next run queues, starts 30 minutes late, and the delay compounds every hour, so by the end of the day the pipeline is many hours behind and the freshness SLA is missed even though every run succeeded. Some schedulers instead drop missed intervals and only run the latest, which trades unbounded lag for silently unprocessed windows. Fixing it means changing the shape of the work — a longer interval, an incremental rather than full-scan read, or fanning the interval out into partition-scoped parallel tasks — not raising the concurrency limit and hoping.
code
text · 8 linescadence 60 min, runtime 90 min, one active run allowed
interval [09:00,10:00) start 10:00 end 11:30 lag 0 min
interval [10:00,11:00) start 11:30 end 13:00 lag 30 min
interval [11:00,12:00) start 13:00 end 14:30 lag 60 min
interval [12:00,13:00) start 14:30 end 16:00 lag 90 min
every run SUCCEEDS; freshness degrades all daygo deeper
Understand that a run taking longer than its schedule interval either overlaps with the next one or delays it, and that neither happens by accident.
Explain the compounding-delay arithmetic under a concurrency cap and name the specific hazards of overlap: source contention, target races and shared staging paths.
Diagnose it from duration-versus-cadence and lag metrics, and argue the real remedies — incremental reads, longer cadence, partition-scoped fan-out — rather than raising a limit.
Set the platform policy: which concurrency behaviour is the default, what freshness alerting is mandatory, and how teams are expected to size cadence against measured runtime growth.
## The invariant being broken A schedule only works if the *average* runtime is comfortably below the interval length. Once it is not, the pipeline has no steady state: work arrives faster than it is consumed. Every mitigation is a way of restoring that invariant, and no amount of configuration hides its absence. ## Behaviour one: overlapping runs If the pipeline permits several active runs, the run for `[10:00, 11:00)` is still executing when the run for `[11:00, 12:00)` starts. What that costs depends entirely on whether the work is partition-scoped: - **Source contention.** Two full-scan extracts against the same operational database at once double the load. If the source is the thing making the job slow, the overlap makes it slower, and runtime grows further — a feedback loop that ends with three or four concurrent runs and a source under real stress. - **Target races.** If both runs write to the same table without scoping, the outcome depends on ordering: duplicate rows from blind appends, or a run deleting rows its sibling just wrote. If each run instead overwrites only the partition named by its own interval, overlap is harmless by construction — the runs simply cannot touch each other's data. - **Shared external state.** Staging paths, temp tables, lock files and named checkpoints must be keyed by interval too. A fixed staging path is the single most common reason overlap corrupts output. So "can this pipeline overlap safely?" is really "is every piece of state this pipeline writes named after its interval?" ## Behaviour two: capped concurrency and a growing backlog Cap the pipeline at one active run and the arithmetic is unforgiving. Each 90-minute run in a 60-minute cadence adds 30 minutes of delay: ```text interval [09:00,10:00) starts 10:00 ends 11:30 interval [10:00,11:00) starts 11:30 ends 13:00 (30 min late) interval [11:00,12:00) starts 13:00 ends 14:30 (60 min late) interval [12:00,13:00) starts 14:30 ends 16:00 (90 min late) ``` Every run succeeds. Nothing alerts. By the end of a day the freshest available data is many hours old, and the only symptom is a stakeholder saying the dashboard looks stale. This is why **success-rate alerting is not enough**: the correct signal is lag — the gap between now and the most recent completed interval — or a coverage check that every interval up to a threshold has output. A third policy exists in some schedulers: drop missed intervals and only run the most recent one. That bounds the lag but silently skips windows, which is worse for a pipeline whose output must be complete and better for one that only ever publishes a current snapshot. Know which of the three your pipeline is configured for; the choice must match whether historical windows matter. ## Diagnosing it Plot run duration against the interval length on the same axis. A pipeline heading for trouble shows duration creeping toward the interval over weeks — usually because data volume grows while the job still does a full scan. The crossing point is predictable long in advance, which makes this one of the few capacity problems you can genuinely get ahead of. Also plot queue time separately from execution time: rising queue time with flat execution time means contention for workers, not a slow job, and the fix is capacity or priority rather than the pipeline itself. ## Fixing it properly 1. **Make the run faster.** Usually the job re-reads history it already processed. Filtering the source by the interval, adding a predicate the source can push down, or reading only new partitions often cuts runtime by an order of magnitude. 2. **Lengthen the interval.** If the business tolerates it, moving hourly to four-hourly restores headroom immediately and reduces per-run overhead. Match the cadence to the freshness the consumers actually need, not to the cadence someone picked at creation. 3. **Fan out within the interval.** Split one monolithic run into parallel partition-scoped tasks so wall-clock time falls while total work stays the same. 4. **Make overlap safe, then allow it.** If every write is partition-scoped and every staging path is interval-keyed, permitting two concurrent runs is a legitimate answer — it absorbs occasional slow runs without building a backlog. It is a bad answer when the source cannot take the extra load. 5. **Alert on lag.** Whatever you choose, add a check that the latest completed interval is within the freshness target. That is the alert that fires for this failure mode; the run-failure alert never will. ## How to answer Name both behaviours and say the choice is a policy, not a default you can ignore. Show the compounding-delay arithmetic. Then make the senior point: raising the concurrency limit is only safe if every write is keyed by the interval, and the real fix is restoring the invariant that runtime is comfortably below cadence.
- Which metric tells you this is happening before anyone complains?Run duration plotted against the interval length, and lag — the age of the most recently completed interval. Duration creeping toward cadence gives weeks of warning; lag is the alert that actually fires when the backlog starts. Success rate stays at 100% throughout and tells you nothing.
- When is simply allowing two concurrent runs a legitimate fix?When every write is scoped to the run's own interval — partition overwrite, interval-keyed staging paths, no shared temp tables — and the source can absorb double the read load. Under those conditions concurrent runs cannot interfere, and overlap absorbs occasional slow runs instead of accumulating a backlog.
- How does a policy of dropping missed intervals change the failure mode?It bounds lag at one interval by always running the latest window and abandoning the ones in between. Acceptable for a pipeline that publishes a current snapshot; unacceptable when historical windows must be complete, because the skipped intervals leave permanent holes that only a deliberate backfill fills.
It is a checkout queue where each customer takes 90 seconds and a new one arrives every 60. One till means the line grows all day; two tills clear it, but only if the two cashiers are not reaching into the same drawer.
saying these in an interview costs you the question
- Raises the concurrency limit without checking write scoping
- Says every run succeeded so the pipeline is healthy
- Assumes the scheduler skips the overrun interval automatically
- Uses a fixed staging path shared by all runs
- Alerts only on run failure, never on data freshness