A Prefect deployment's scheduled runs are piling up as Late — how do you diagnose it?
answer
- the run never started, so nothing failed
- follow the path from schedule to compute
- something has to be polling
- paused, mismatched, or full
- expect a stampede when it comes back
basics
~20 sLate means Prefect created the runs on schedule but nothing claimed them. Check that a healthy worker of the right type is polling that exact work pool and queue, that neither is paused, and that a pool, queue or global concurrency limit is not saturated.
solid answer
~50 s`Late` is an orchestration-side symptom, not a flow failure: the API created the run at its scheduled time and it is still `Scheduled` past a grace period. Work the path from run to compute. **Is a worker running?** `prefect work-pool ls` shows the pool's status and whether it has a recent worker heartbeat; a worker whose type does not match the pool's type will never take the work. **Is the pool or work queue paused?** Both pause independently, and a paused queue silently strands only its deployments. **Is the deployment pointed at the right pool and queue name?** A typo creates a queue nobody serves. **Is a concurrency limit saturated?** Pool, queue, global and tag limits all hold runs in `Scheduled`. **Is the worker healthy but failing to submit?** Its logs show image-pull, RBAC or credential errors. Once fixed, expect a thundering herd of the backlog; cap it with a pool concurrency limit before resuming.
code
bash · 9 lines# is the pool being served at all?
prefect work-pool ls
prefect work-queue ls --pool k8s-prod
# resume anything paused
prefect work-pool resume k8s-prod
# drain the backlog at a controlled rate before restarting workers
prefect work-pool set-concurrency-limit k8s-prod 5go deeper
Know that a Late run has not executed at all — the schedule fired but nothing picked it up — so the first check is whether a worker is running against that work pool.
Walk the chain end to end: worker alive and type-matched, pool and queue not paused, deployment routed to the right queue, concurrency limits not saturated, worker logs for submission errors.
Show the operational judgment: read worker logs rather than flow logs, decide whether the backlog should run at all, and cap concurrency before recovery so a queued fleet does not stampede the cluster or a source database.
Own prevention — redundant workers per production pool, automations alerting on Late and on stale worker heartbeats, per-environment pool isolation, and a written policy for what happens to a backlog after an outage.
## What Late actually means Prefect creates flow runs in advance from a deployment's schedule; each sits in `Scheduled` with a start time. When that time passes by more than a small grace period without the run being claimed and started, Prefect marks it `Late`. Nothing has crashed, nothing has been retried, and the flow's code has not executed at all. `Late` says: **the orchestration layer did its job, the execution layer never picked up the work.** Distinguishing that from `Failed` (the code ran and raised) or `Crashed` (the infrastructure died mid-run) is the first thing an interviewer listens for. ## Diagnose along the path **1. Is any worker polling that pool?** `prefect work-pool ls` reports the pool and whether it is ready; the UI shows last worker heartbeat. A stopped container, an evicted pod, an expired API key, a worker started against the wrong workspace or a wrong `PREFECT_API_URL` all present identically: pool exists, backlog grows, no poller. This is by far the most common cause. **2. Does the worker type match the pool type?** A `process` worker will not serve a `kubernetes` pool. The pool stays unserved even though a worker process is clearly running somewhere. **3. Is the pool or the work queue paused?** They pause independently. A paused *queue* is nastier because the pool still looks healthy and only the deployments routed to that queue go Late. **4. Does the deployment point at a pool and queue that exist?** Assigning a deployment to a queue name that no worker serves — a typo, or a queue created implicitly — produces a queue with a backlog and no consumer. **5. Is a concurrency limit saturated?** Pool limits, per-queue limits, global concurrency limits and tag-based limits all hold runs in `Scheduled`, which becomes `Late` once their start time passes. This is the benign case: the system is working as configured and you are simply oversubscribed. Look at how many runs are `Running` versus the limit before blaming infrastructure. **6. Is the schedule itself too tight?** A deployment scheduled every five minutes whose runs take eight will queue permanently. The fix is concurrency policy on the deployment, not more workers. **7. Is the worker polling but failing to submit?** The worker logs are decisive: `ImagePullBackOff` or an unauthorized registry, missing Kubernetes RBAC to create Jobs, no cluster quota, expired cloud credentials. Runs may bounce between states or stay pending while the worker retries. Always read worker logs, not just flow-run logs — the flow has no logs yet, because it never started. ## Fixing it without a stampede When a worker returns after an outage, the entire backlog becomes eligible at once. On a Kubernetes pool that can mean hundreds of Jobs submitted in seconds, blowing through cluster quota or hammering a source database. Two controls matter: - Set a **pool or queue concurrency limit** before restarting the worker, so the backlog drains at a controlled rate. - Decide whether the backlog is even wanted. Prefect has no Airflow-style `catchup` toggle; the runs already exist as scheduled objects. For a pipeline where only the latest state matters, delete or cancel the stale scheduled runs rather than executing a hundred redundant ones — and if the outage is long and planned, pause the deployment's schedule so the runs are never created. The judgment answer is that catching up is a *decision*: an incremental load that will simply reprocess to current state can safely skip the backlog, while an interval-partitioned job that writes one partition per run genuinely needs each run. ## Preventing recurrence Run more than one worker per production pool so a single crash is not an outage, and run workers under a supervisor or as a Kubernetes Deployment with restarts. Alert on the condition rather than eyeballing it: an automation that fires when a flow run enters `Late`, or a check on worker heartbeat age, catches a dead worker in minutes instead of at the morning stand-up. Separate pools per environment so a dev worker outage cannot look like a production one. And make late-ness visible per deployment, because a single deployment going Late while its pool-mates run fine points straight at a queue or concurrency issue rather than a dead worker. ## Version note Describes Prefect 3.x, where workers poll work pools. Under Prefect 2's original agent model the same symptom pointed at a stopped **agent** or a work queue no agent was serving.
- How is a Late Prefect flow run different from a Crashed one?`Late` means the run was never claimed — the schedule fired, no worker started it, and your code has not executed, so there are no flow logs. `Crashed` means execution began and the infrastructure died out from under it: the pod was evicted, the process was OOM-killed, the container vanished. Late points at the worker layer; Crashed points at the runtime.
- A worker is healthy and polling, but its deployment's runs still go Late. What now?Look at concurrency and routing. A pool, queue, global or tag concurrency limit that is saturated holds runs in Scheduled; so does a deployment assigned to a queue that worker does not serve, or a paused queue inside a healthy pool. Compare running-run count against the configured limits before touching infrastructure.
- A worker was down for six hours. Do you let the backlog run?It depends on whether runs are partition-scoped. If each run writes its own interval or partition, you want them all, drained under a concurrency limit. If the flow just reprocesses to current state, one run does the work of sixty — cancel the stale scheduled runs and trigger once. Decide before restarting the worker, not after.
- How would you get alerted on this automatically rather than discovering it later?Create an automation that triggers on a flow run entering the Late state and notifies your on-call channel, and separately monitor worker heartbeat age for each production pool. Alerting on lateness catches concurrency saturation and paused queues too, which a pure worker-liveness check would miss.
saying these in an interview costs you the question
- Says Late means the flow ran and failed, so restarts the code
- Debugs the flow's Python before checking whether a worker is polling
- Forgets a paused work queue inside a healthy work pool
- Ignores concurrency limits as a cause of runs waiting
- Restarts the worker and lets the whole backlog stampede at once