skip to content

Scheduling and Data Intervals

The difference between when a run executes and which slice of data it is responsible for. This is the concept interviewers use to catch cargo-culting: everyone can write a cron string, but far fewer can explain why a daily run starting at midnight processes yesterday.

on this pageshow

questions

6

What does the pipeline cron schedule `*/15 9-17 * * 1-5` actually fire on?

level: juniorimportance: must knowfreq 70%

answer

  1. five positional fields, smallest unit first
  2. the slash is a step, not an offset
  3. ranges include their upper bound
  4. nine to seventeen is nine hours, not eight

basics

~20 s

It fires every fifteen minutes, at :00, :15, :30 and :45, during hours 09 through 17 inclusive, Monday to Friday. That is 36 firings per weekday, the last one at 17:45, and none at weekends.

solid answer

~40 s

The five fields are minute, hour, day-of-month, month, day-of-week, in that order. `*/15` in the minute field means minutes 0, 15, 30 and 45; `9-17` is an **inclusive** hour range, so it covers 09:00 through 17:59, not up to 17:00; the two `*` fields put no restriction on day-of-month or month; and `1-5` is Monday through Friday. Combined: 36 runs on each weekday, first at 09:00 and last at 17:45. The two details candidates get wrong are that the hour range includes its upper bound, and that `*/15` steps from the start of the hour rather than from the moment the schedule was created. Also worth stating: a bare cron expression carries no timezone, so the effective wall-clock times depend on the timezone the scheduler is configured with.

code

text · 9 lines
text
*/15 9-17 * * 1-5

field 1  */15   -> minutes 0, 15, 30, 45
field 2  9-17   -> hours 09,10,11,12,13,14,15,16,17 (inclusive)
field 3  *      -> any day of month
field 4  *      -> any month
field 5  1-5    -> Monday..Friday

first firing 09:00, last firing 17:45, 36 per weekday

go deeper

for a junior

Be able to read a five-field expression aloud in order and state the first and last firing time without hesitation; this is a common screening question.

for a middle

Explain step and range semantics precisely, including why steps align to the clock, and name a cadence cron cannot express.

for a senior

Discuss the operational consequences: schedule alignment across many pipelines, timezone configuration, and when to reach for an interval or programmatic schedule instead.

for a principal

Set the convention for a platform — which schedule styles teams may use, which timezone is standard, and how staggering prevents hundreds of pipelines firing on the same round minute.

## Field layout The classic five-field expression is positional: ```text */15 9-17 * * 1-5 | | | | | | | | | +-- day of week (0-6, Sunday is usually 0) | | | +-------- month (1-12) | | +------------- day of month (1-31) | +------------------- hour (0-23) +--------------------------- minute (0-59) ``` Each field accepts a wildcard `*`, a single value (`30`), a list (`0,30`), an inclusive range (`9-17`), or a step (`*/15`, or `9-17/2` for every second hour in that range). A firing happens at every minute where **all** fields match. So `*/15 9-17 * * 1-5` matches minutes 0, 15, 30 and 45 of every hour from 09 to 17 inclusive, on Monday through Friday: nine hours times four firings equals 36 per weekday, first at 09:00, last at 17:45. ## The three details that trip people up **Ranges are inclusive at both ends.** `9-17` is nine hours, not eight. People expecting the last run at 17:00 are off by 45 minutes and four runs. **Steps are anchored to the field, not to the deploy.** `*/15` means "minutes divisible by 15", so the schedule is aligned to the clock. Deploying at 09:07 does not produce firings at :07, :22, :37. This is exactly what you want for data pipelines, because it makes the window boundaries land on round numbers. **Day-of-month and day-of-week interact oddly.** When *both* fields are restricted (neither is `*`), the traditional Unix behaviour is to treat them as OR rather than AND, so `0 0 1 * 1` fires on the 1st of the month *and* on every Monday. Implementations differ here; if you need "the first Monday", check your scheduler's documentation rather than assuming, and prefer expressing it in code. ## What cron cannot express Cron is a *pattern match on a calendar clock*, not an interval timer. That means several ordinary-sounding requirements are simply not expressible: - **Any cadence that does not divide its unit evenly.** "Every 90 minutes" has no cron expression, because the firing times drift across hours. `0 */6 * * *` works (00:00, 06:00, 12:00, 18:00) because six divides 24; `0 */7 * * *` does not do what people expect — it fires at 00, 07, 14 and 21, then restarts the pattern at midnight, leaving a three-hour gap. - **"N units after the previous run finished."** Cron does not know when the last run finished, or whether it is still running. - **Business-calendar rules** like "the last business day of the month" or "skip public holidays", which need real logic. This is why orchestrators usually offer a second style of schedule alongside cron — a fixed **interval** such as every 30 minutes or every 6 hours, measured from an anchor point — and, for anything genuinely irregular, a programmatic schedule you write yourself. ## Timezone A cron string on its own has no timezone; it inherits whatever the scheduler is configured to use. Two consequences matter. First, `0 9 * * 1-5` means 09:00 in that configured zone, which may not be any stakeholder's local morning. Second, if the configured zone observes daylight saving, some wall-clock times will be skipped or repeated twice a year, and a nominal 24-hour day becomes 23 or 25 hours. Defining schedules in UTC removes that class of problem at the cost of the schedule drifting relative to local business hours. ## Reading an unfamiliar expression under pressure Work left to right and say each field out loud: "minutes zero fifteen thirty forty-five, hours nine through seventeen, any day of month, any month, Monday through Friday." Then sanity-check the count — 4 per hour, 9 hours, 5 days a week — and state the first and last firing time. Naming the boundary cases is what separates a confident answer from a guess.

  • How would you schedule a pipeline to run every 90 minutes?
    Not with a five-field cron expression — 90 minutes does not divide an hour or a day, so no clock pattern matches it. Use the scheduler's interval-style schedule (every 90 minutes from an anchor), or enumerate the firing times explicitly across a 24-hour cycle if the tool only accepts cron.
  • Why is `0 */7 * * *` usually a mistake?
    Steps restart at the beginning of each field, so it fires at 00, 07, 14 and 21 hours and then jumps back to 00 — a three-hour gap once a day instead of an even seven-hour cadence. Any step that does not divide 24 produces this uneven wrap.
  • What is missing from a cron expression that a pipeline still needs?
    A timezone, and a definition of which data window each firing owns. Cron says only when to trigger; the orchestrator supplies the interval the run covers, and the configured timezone decides what the wall-clock times actually mean, including daylight-saving behaviour.

saying these in an interview costs you the question

  • Reads 9-17 as ending at 17:00 rather than 17:59
  • Thinks */15 counts from deployment time
  • Believes a cron expression carries its own timezone
  • Claims any cadence can be written as cron
  • Mixes up the day-of-month and day-of-week field positions

context

open as a page

In a scheduled data pipeline, why does the midnight run process yesterday's data rather than today's?

level: middleimportance: must knowfreq 78%

basics

~20 s

A scheduled run is named for the data interval it covers, not for the clock time it starts. A daily interval only closes at midnight, so the run firing then is responsible for the day that just ended.

open as a page

Why do pipeline data intervals use half-open bounds that include the start but exclude the end?

level: middleimportance: should knowfreq 50%

basics

~20 s

Half-open bounds make consecutive intervals tile the timeline with no gap and no overlap. A record whose timestamp lands exactly on a boundary belongs to exactly one run, so nothing is processed twice and nothing is skipped.

open as a page

An hourly pipeline regularly takes 90 minutes to finish — what happens to the next scheduled run?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It depends on the concurrency policy. Either two runs execute at once and contend for the same source and target, or the pipeline is capped at one active run and the queue backs up, so each run starts later than the last and lag grows without bound.

open as a page

How do you decide whether a downstream pipeline runs on a clock schedule or on upstream data readiness?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use a clock schedule only when the downstream can tolerate acting on whatever data exists at that moment. If correctness depends on the upstream being complete, trigger on an explicit readiness signal instead, because a guessed time offset fails silently when the upstream is late.

open as a page

What happens to a pipeline scheduled in a local timezone when that zone's daylight-saving change occurs?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

One local wall-clock hour disappears in spring and one repeats in autumn, so a schedule inside those hours may be skipped or fire twice, and that day's nominal daily interval is 23 or 25 hours instead of 24.

open as a page