skip to content

For a Dataflow job other teams launch, how do classic and Flex templates differ?

level: seniorimportance: nice to knowfreq 30%

answer

  1. it comes down to when the graph is built
  2. one artifact is a frozen graph, one is an image
  3. a fixed graph forces indirection for parameters
  4. ValueProvider only matters for one of them
  5. the newer one recommended for new pipelines

basics

~20 s

A Dataflow classic template stages a pre-built job graph in Cloud Storage, so per-launch parameters must be read through ValueProvider and the graph shape is fixed. A Flex template packages the pipeline in a container image and builds the graph at launch, with real parameter values.

solid answer

~50 s

Both let someone launch a Dataflow job without your source code or a Python or Java toolchain, but they differ in **when the pipeline graph is constructed**. A **classic template** is produced by running your driver with `--template_location`, which stages a serialized job graph plus a metadata file in Cloud Storage; launching just instantiates that frozen graph. Because the graph was built before anyone supplied parameters, anything that varies per launch must be plumbed through `ValueProvider`, and transforms that do not accept a `ValueProvider` simply cannot be parameterized — nor can the graph's shape depend on a parameter. A **Flex template** packages the pipeline and its dependencies into a container image with a template spec in Cloud Storage; at launch the container runs your driver, so the graph is built with the real argument values. No `ValueProvider` plumbing, parameter-dependent graph shapes are fine, and dependencies are pinned in the image. Flex is the current recommendation for new pipelines.

code

bash · 12 lines
bash
# classic: stage a frozen job graph, then launch it
python pipeline.py \
  --runner=DataflowRunner \
  --project=my-project \
  --region=us-central1 \
  --temp_location=gs://my-bucket/tmp \
  --template_location=gs://my-bucket/templates/my-template

gcloud dataflow jobs run my-job \
  --gcs-location=gs://my-bucket/templates/my-template \
  --region=us-central1 \
  --parameters=inputPath=gs://in/2026-08-21/

go deeper

for a junior

Know that a Dataflow template lets someone launch a job with parameters without having the pipeline source or an SDK installed.

for a middle

Explain the timing difference — classic stages a frozen graph, Flex builds the graph at launch inside a container — and what that implies for how parameters reach the pipeline.

for a senior

Show the practical consequences: ValueProvider plumbing and its transform-support gaps, why a parameter cannot change a classic graph's shape, and how a container image fixes worker dependency drift.

for a principal

Own the release story: templates as versioned, immutable artifacts, who is allowed to launch them, how parameters are validated with a metadata file, and how a rollback works when a bad pipeline version ships.

## The problem templates solve Normally a Dataflow job is launched by running your Beam driver program, which needs the source, the SDK, and the right dependencies on the launching machine. That is fine for you and awkward for everyone else — an analyst, a scheduler, an orchestrator, or a service in another language. **Templates** decouple *building* a pipeline from *launching* it: you publish an artifact once, and anyone with permission launches jobs from it via the console, `gcloud`, or the REST API, passing parameters. Google publishes a set of ready-made templates for common moves, and you can build your own in either of two flavours. ## Classic templates: graph built at staging time You create one by running your driver with a template location instead of executing the job: ```bash python pipeline.py \ --runner=DataflowRunner \ --project=my-project \ --region=us-central1 \ --temp_location=gs://my-bucket/tmp \ --template_location=gs://my-bucket/templates/my-template ``` This serializes the constructed job graph to Cloud Storage. You normally also publish a metadata file next to it describing the accepted parameters so the console can render and validate them. Launching: ```bash gcloud dataflow jobs run my-job \ --gcs-location=gs://my-bucket/templates/my-template \ --region=us-central1 \ --parameters=inputPath=gs://in/2026-08-21/,outputTable=proj:ds.table ``` The critical consequence: **the graph already exists** when the launcher supplies parameters. A value read normally in your driver — `args.input_path` used to build a `ReadFromText` — was baked in at staging time. To vary it per launch you must declare it as a `ValueProvider` (a `RuntimeValueProvider` resolved on the worker) and the transform you pass it to must support `ValueProvider` arguments. Many do; not all do. And no amount of plumbing lets a *parameter decide the shape of the graph* — you cannot add a branch, choose a different sink type, or fan out over a parameter-supplied list of tables, because those decisions happen while the graph is being built and the graph is already frozen. ## Flex templates: graph built at launch time A Flex template inverts this. You package the pipeline into a **container image** and publish a small template spec JSON in Cloud Storage pointing at that image: ```bash gcloud dataflow flex-template build gs://my-bucket/templates/my-flex.json \ --image-gcr-path=gcr.io/my-project/my-pipeline:1.4.0 \ --sdk-language=PYTHON \ --metadata-file=metadata.json gcloud dataflow flex-template run my-job \ --template-file-gcs-location=gs://my-bucket/templates/my-flex.json \ --region=us-central1 \ --parameters=input_path=gs://in/2026-08-21/,output_table=proj:ds.table ``` At launch, Dataflow starts a launcher that runs your program **inside that container**, with the supplied parameters as ordinary arguments. Your driver then constructs the graph the way it always does. So: - **No `ValueProvider` plumbing.** Parameters are plain values in normal code. - **Parameter-dependent graph shape works.** A parameter can select a sink, add a branch, or drive a loop that builds N transforms. - **Dependencies are pinned in the image**, which removes the whole class of "the worker environment does not match my dev environment" problems, and gives you an immutable, versioned artifact you can tag alongside your release. - **Any Beam pipeline can be templated**, including those using transforms that never supported `ValueProvider`. The cost is a container build and registry in your release pipeline, and a launch that has to start that container before the job graph exists — so a Flex launch has a startup step a classic launch does not. ## Choosing For new work, Flex is the current recommendation, and the reasons are the ones above: fewer restrictions, better dependency hygiene, an immutable versioned artifact. Classic templates remain relevant mainly because a lot of existing tooling and many Google-provided templates use them, and because their launch path is very simple. Operationally, both give you the same organizational win: the people who run the job do not need the code. A scheduler triggers a launch with today's parameters, an analyst can launch a backfill from the console, and the pipeline version in production is an artifact you can point at, rather than whatever was on someone's laptop. ## What to say "The difference is when the graph is built. Classic stages a frozen graph in Cloud Storage from a `--template_location` run, so per-launch parameters need `ValueProvider` and the graph shape cannot depend on a parameter. Flex ships a container image plus a spec, and the container builds the graph at launch with real values — no `ValueProvider`, parameter-driven graph shapes are fine, and dependencies are pinned in the image. I'd pick Flex for anything new."

  • Why does a classic Dataflow template need ValueProvider at all?
    Because the graph is serialized to Cloud Storage before any launcher supplies parameters. A plain value read in the driver is baked in at staging time. A `ValueProvider` is an indirection whose value is resolved at run time on the worker, so the frozen graph can still consume a per-launch argument — provided the transform you pass it to accepts a ValueProvider.
  • What can a Flex template do that a classic template cannot?
    Let a parameter change the shape of the graph — select a different sink, add a branch, or build N transforms from a supplied list — because the driver runs at launch with real arguments. It also templates pipelines using transforms that never supported ValueProvider, and pins dependencies in an immutable, versioned container image.
  • What does templating buy an organization beyond convenience?
    It separates who builds a pipeline from who runs it. Schedulers, analysts and other services can launch jobs with parameters without the source, an SDK, or a matching environment, and the version running in production is a published artifact you can point at and roll back, rather than whatever was on a developer's machine.

saying these in an interview costs you the question

  • Thinks both template types build the graph at launch
  • Says a classic template can branch on a launch parameter
  • Claims a Flex template still needs ValueProvider
  • Treats the template as the running job itself
  • Assumes any pipeline works as a classic template unchanged

context