In Great Expectations, what does a Checkpoint add beyond an Expectation Suite?
answer
- the suite is a spec, something must run it
- binds rules to a specific slice of data
- one invocation, several suite-batch pairings
- it also fires follow-up actions
- reporting is not the same as enforcing
basics
~20 sA Checkpoint is the runnable unit: it binds an Expectation Suite to a specific batch of data, executes the validation, and fires configured actions such as storing the result and rebuilding Data Docs. The suite alone only declares rules.
solid answer
~50 sAn Expectation Suite is a static declaration - rules with no data attached and no way to run itself. A **Checkpoint** is what you invoke from a pipeline task: it pairs one or more suites with the batches they should be validated against, runs the validation, and then executes an ordered list of **actions** on the outcome - persisting the validation result to a store, rebuilding **Data Docs**, sending a notification. It returns a result object with a top-level `success` covering every validation it ran. That is why a Checkpoint is the object your orchestrator or CI step references by name: it is reproducible, parameterizable (you can pass in the batch for today's partition at run time), and the single place where "which rules, on which data, and what happens next" is written down. Note that actions run whether or not validation succeeded - a Checkpoint is a gate only if your code reads `success`.
code
yaml · 15 linesname: orders_gate
config_version: 1
class_name: Checkpoint
validations:
- batch_request:
datasource_name: warehouse
data_asset_name: orders
expectation_suite_name: orders_raw
action_list:
- name: store_validation_result
action:
class_name: StoreValidationResultAction
- name: update_data_docs
action:
class_name: UpdateDataDocsActiongo deeper
Remember the one-line distinction: the suite says what the rules are, the Checkpoint is the thing a pipeline actually invokes to run them against real data.
Explain the binding of suite to batch, that a single Checkpoint can run several suite-batch pairs, and that its action list persists the result and rebuilds Data Docs after the run finishes.
Demonstrate placement judgment - post-load versus pre-publish versus CI - and be explicit that a Checkpoint reports while your task enforces. Mention parameterizing the batch so one named gate covers every partition.
Own the standard for how gates are named, owned and rolled out across pipelines, including which failures block a publish and which merely annotate, and how you avoid a proliferation of checkpoints nobody maintains.
## The separation the tool is making Great Expectations deliberately splits *what the rules are* from *when and where they run*. An Expectation Suite is a serializable list of assertions - `expect_column_values_to_not_be_null`, `expect_table_row_count_to_be_between`, and so on - with no data bound to it and no execution semantics. A **Checkpoint** is the executable object that closes that gap. ## What a Checkpoint holds Conceptually a Checkpoint names three things: 1. **What to validate** - a batch: a table, a partition of a table, a landed file, or a query result. In GX 1.x this is expressed through a batch definition attached to a data asset; in 0.18 it was a batch request inside the checkpoint config. 2. **Which suite(s) to apply** - one Checkpoint can run several suite-to-batch pairings in a single invocation, which is how a team validates the five tables produced by one job in one step. 3. **What to do with the outcome** - an ordered list of actions. The canonical ones store the validation result in the results store and rebuild Data Docs; others send notifications to chat or email. ## Why the run-time binding matters Pipelines validate *today's* data, not a fixed table. A Checkpoint is parameterizable: the orchestrator passes the batch identifier - the partition date, the file path, a runtime dataframe - so the same named Checkpoint validates a different slice on every run while the suite stays constant. That is what makes a checkpoint name safe to hardcode in a DAG task or a CI job: the reference is stable, the data is not. It also gives you a stable identity for results. Every run of `orders_gate` writes a validation result tagged with that checkpoint and a run identifier, so Data Docs shows a history for that gate rather than a pile of anonymous validations. When someone asks "has this check been failing all week?", the checkpoint identity is what makes the question answerable. ## Actions, and the trap in them The action list executes **after** validation, regardless of the outcome. Storing the result and updating Data Docs happen on failure as well as success - that is the point, since a failure is exactly the run you want documented. The trap is concluding that a Checkpoint therefore *enforces* anything. It does not. The invocation returns a result object whose top-level `success` is false if any validation inside it failed, and unless your task reads that flag and raises, the pipeline sails on with bad data downstream and a very tidy red entry in Data Docs that nobody opened. A related subtlety: an expectation that cannot even be evaluated - a column that no longer exists in the source, for instance - is captured as an exception recorded in the result rather than crashing the process. That is helpful for getting a complete report, and it means your gate logic should look at both failures and recorded exceptions, not just assume a returned result means the checks all ran cleanly. ## Where checkpoints get invoked Three common placements, in rising order of usefulness: - **Post-load**, on the landing table right after ingestion, so a bad extract is caught before transformation consumes it. - **Pre-publish**, on the model or mart before it becomes visible to consumers, so the blast radius of a failure is a stale table rather than a wrong one. - **In CI**, against a sample or a staging environment, so a change to a transformation is checked before it ships. Orchestrators generally call a checkpoint either through plain Python in a task or through a community operator that fails the task on a false `success`. ## What to say in an interview The crisp formulation is: the suite is the specification, the Checkpoint is the invocation - it binds rules to a batch, produces a durable, identifiable result, and runs the follow-up actions. Then add the operational caveat immediately, because it is the half candidates forget: a checkpoint reports, your code enforces.
- Do a Checkpoint's actions still run when validation fails?Yes - the action list executes after validation regardless of the outcome, which is deliberate: a failed run is exactly the one you want stored and rendered into Data Docs. Actions are reporting, not control flow. If you want the pipeline to stop, your task must inspect the returned result's success flag and raise.
- How does the same Checkpoint validate a different partition on every run?The batch is parameterized at run time. The orchestrator passes the identifier for today's slice - a partition date, a file path, or a runtime dataframe - while the suite and the checkpoint name stay constant. That is what lets a DAG task hardcode the checkpoint reference and still validate fresh data each execution.
- Why does having a stable checkpoint name matter for Data Docs?Every run writes a validation result tagged with the checkpoint identity and a run id, so Data Docs can show a history for that specific gate instead of a heap of unrelated validations. That history is what lets you answer whether a check has been flapping for a week or only failed today.
saying these in an interview costs you the question
- Saying a Checkpoint blocks the pipeline by itself
- Thinking actions only fire when validation succeeds
- Believing one Checkpoint can only run a single suite
- Hardcoding a fixed batch instead of parameterizing per run
- Treating the Data Docs rebuild as an alerting mechanism