skip to content

In Dataflow, what is the difference between draining and cancelling a streaming job?

level: middleimportance: must knowfreq 70%

answer

  1. one verb is polite, one is not
  2. what happens to data already inside the pipeline
  3. the watermark is pushed to the end of time
  4. one flushes open windows, one discards them
  5. Draining/Drained versus Cancelling/Cancelled

basics

~20 s

Draining a Dataflow streaming job stops it from reading new input, advances the watermark to infinity so every open window fires and in-flight data is written to sinks, then finishes. Cancelling stops immediately and discards buffered state and in-flight data.

solid answer

~50 s

Both stop a running Dataflow streaming job, but they differ in what happens to work already inside the pipeline. `gcloud dataflow jobs drain JOB_ID` closes the unbounded sources, stops pulling new messages, then advances the watermark to infinity so all open windows and timers fire; buffered elements are processed and written to the sinks, and the job ends in the **Drained** state. `gcloud dataflow jobs cancel JOB_ID` tears the workers down right away — in-flight elements and buffered window state are thrown away, and Pub/Sub messages that were read but not acknowledged simply become available again on the subscription. Drain is the clean-shutdown option when you are about to redeploy and can tolerate windows emitting partial results; cancel is for a job that is broken, wedged, or emitting bad data and that you want stopped now. Neither carries state into the next job.

code

bash · 5 lines
bash
# graceful: stop reading, fire open windows, write out, then finish
gcloud dataflow jobs drain JOB_ID --region=us-central1

# immediate: discard in-flight elements and buffered state
gcloud dataflow jobs cancel JOB_ID --region=us-central1

go deeper

for a junior

Recall that Dataflow streaming jobs are stopped with either drain or cancel, and that drain is the polite one that lets buffered work finish while cancel stops immediately.

for a middle

Explain the mechanism: drain closes the sources and advances the watermark to infinity so open windows and timers fire, while cancel discards in-flight elements and buffered state. Know the resulting job states.

for a senior

Be ready to choose under incident pressure and to name the consequences — partial final windows, a burst of output at the sink, unacknowledged Pub/Sub messages returning to the subscription, and duplicates when you restart after a cancel.

for a principal

Own the redeploy policy: when a team may drain-and-restart versus update in place, what the sink contract must be (idempotent or upsert-keyed) for either to be safe, and how that shapes window sizing and downstream SLAs.

## Two ways to stop a streaming job A Dataflow streaming job never finishes on its own — its sources are unbounded, so it runs until you stop it. Dataflow gives you two stop verbs, and interviewers ask about them because picking the wrong one at 2am either loses data or leaves a wedged job running. ```bash gcloud dataflow jobs drain JOB_ID --region=us-central1 gcloud dataflow jobs cancel JOB_ID --region=us-central1 ``` ## What drain does, step by step Drain is a *graceful* shutdown. When you issue it: 1. The job stops reading from its unbounded sources. A Pub/Sub source stops pulling; nothing new enters the pipeline. 2. Dataflow advances the pipeline's watermark to "infinity" (the end of time). Because window firing is driven by the watermark, this immediately makes every open window look complete. 3. All open windows fire and all pending timers run. Aggregations that were still accumulating emit whatever they have accumulated so far. 4. The resulting elements flow through the remaining transforms and are written to the sinks. 5. The job transitions **Draining → Drained** and the workers are released. The important consequence: buffered data is *not* lost — it is flushed. But it is flushed early, so a window that would normally have covered a full hour emits a partial result covering only the part of the hour that had arrived. Downstream consumers that assume one row per window per key can see a short, truncated final window. ## What cancel does Cancel is a hard stop. Dataflow shuts the workers down as fast as it can. Elements that were in flight, timers that had not fired, and accumulated window and per-key state are discarded. Nothing extra is written to the sinks; anything already written stays written. For a Pub/Sub source this is usually recoverable rather than catastrophic: messages the pipeline had read but not yet acknowledged remain unacknowledged, so they return to the subscription and a later job reads them again. That is also why cancel plus restart tends to produce duplicates at the sink for whatever was in flight — the pipeline's exactly-once guarantee is scoped to a single job's execution, not across a cancel. Cancel is the right call when the job is producing wrong output, when a bad deploy needs to be stopped before it writes more, or when a drain has itself stalled. A drain that is stuck can be escalated to a cancel. ## Where drain misbehaves Drain is not free and not always fast: - **Large state takes time.** Advancing the watermark to infinity fires *everything* at once. A job holding a lot of keyed or windowed state produces a burst of output at the end of a drain, which can hammer the sink or take a long time to complete. - **Global windows with data-driven triggers** may not produce a meaningful final firing; you get whatever the trigger emits at watermark-infinity. - **The final windows are partial.** If the downstream table is keyed by window and you re-run, you may get one partial row from the drained job and one full row from the replacement job for the same window — you need an idempotent or upsert-shaped sink to survive that. - Drain applies to streaming jobs; a batch job is stopped with cancel. ## Neither preserves state for the next job This is the point candidates most often miss. Drain flushes state *out*; cancel throws it away. In both cases the replacement job you start afterwards begins with empty state and reads whatever remains on the subscription. If you need the new job to *continue* from the old job's state — for example, to keep long-lived per-key accumulations or in-progress windows — the mechanism is an in-place update (submitting the replacement with `--update` against the running job's name), or restoring from a Dataflow snapshot. Drain-and-restart is the fallback you use when an update is not possible, and it is a deliberate trade: you accept partial final windows and re-processing of the source backlog in exchange for a clean, simple cutover. ## Picking one under pressure Ask two questions. First: *is the job currently doing damage?* If yes, cancel. Second: *do I care about the data currently buffered inside the pipeline?* If yes and the job is healthy, drain. If you care about the buffered **state** rather than the buffered data — you want the new job to inherit it — neither verb is your answer; update in place instead.

  • After draining a Dataflow streaming job, does the replacement job you launch inherit the drained job's state?
    No. Drain flushes state out of the pipeline and ends the job; a job you launch afterwards starts with empty state and reads whatever is still on the source. If you need state to carry over, submit the replacement with `--update` against the running job instead of draining, or restore from a Dataflow snapshot.
  • You drain a job with hourly fixed windows at 10:20. What lands in the sink for the 10:00 window?
    A partial result covering roughly 10:00–10:20, because drain advances the watermark to infinity and fires every open window immediately. If a replacement job then reprocesses part of that hour you can get two rows for the same window and key, so the sink should be idempotent or upsert-keyed by window and key.
  • A drain has been running for a long time and the job will not finish. What do you do?
    Check the job's state and per-stage metrics: a slow drain usually means a very large flush of buffered state, a slow or throttled sink absorbing the burst, or a stuck stage. You can wait it out if output still moves, but if nothing is progressing, escalate to `gcloud dataflow jobs cancel` and accept losing the in-flight data.

saying these in an interview costs you the question

  • Says drain preserves state for the next job
  • Thinks cancel flushes windows before stopping
  • Claims drained final windows contain complete window data
  • Believes cancel loses already-acknowledged Pub/Sub messages permanently
  • Uses drain to stop a job that is writing corrupt output

context